Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
eis
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
eis
3d ago
He actually publishes a new episode every Monday! Once a month he publishes a special AMA episode which covers a lot of listener (Patreon subscribers) questions and usually is 3-4 hours long! I'm always looking forward to these especia
2.
▲
by
eis
3d ago
I can wholeheartedly recommend Sean Carroll's Mindscape podcast if you want to listen to interesting topics related to Physics, Philosophy, Quantum Mechanics and Science in general: https://www.youtube.com/@seancarroll&
3.
▲
by
eis
6d ago
I know consumers hate the situation with ram and storage prices right now, as do I. But at least on the bright side all this AI investment has unlocked a lot of progress in a space that didn't see huge advancements in a good while. All
4.
▲
by
eis
9d ago
The benchmark is very rudimentary. It does not test different levels/settings apart from its own -b 256/512 (does it affect decompression?), it doesn't measure compression time and memory usage. It does not specify parallel v
5.
▲
by
eis
9d ago
This is completely normal. You can't just issue a chargeback because you feel like you want your money back after the fact. They want to know what went wrong. It could be fraud but it could be also misleading checkout experience and ot
6.
▲
by
eis
9d ago
Making a wire transfer is easy, but can it be deducted from taxes? That's the tricky part.
7.
▲
by
eis
9d ago
Debit cards can do chargebacks just as credit cards can do. That's a feature of Visa and Mastercard (and others). If your bank refuses, complain to the card network. Oh and consider changing banks, that's really unacceptable in 20
8.
▲
by
eis
13d ago
In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58). In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3. Am I missing something or
9.
▲
by
eis
13d ago
That's a good point. Seems like Bedrock offers the same pricing while also providing an uptime SLA.
10.
▲
by
eis
13d ago
Cost and reliability are the two reasons why we don't use Claude in our product. Getting close to one nine, that's not something one can build a reliable product upon. We now use OpenAI with Gemini fallback (or vice versa dependin
11.
▲
by
eis
14d ago
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets... 3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on
12.
▲
by
eis
14d ago
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just uppsed the thinking budgets... 3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on hig
13.
▲
by
eis
15d ago
According to Artificial Analysis, 5.1 cost 56% MORE than 5, $8523 vs $5455. Yes cache cost is lower but it was MUCH more verbose: 140M vs 83M output tokens. This directly contradicts what Anthropic is presenting here. Yes it scores high
14.
▲
by
eis
15d ago
I am not sure if Fable is worth it, at least with version 5 vs Opus 5. Opus beats Fable in quite a few benchmarks and at twice the cost I just haven't seen it provide noticeably better results compared to Opus. Has anyone noticed big d
15.
▲
by
eis
15d ago
> Increases in LLM scores are now mainly driven by post training (evidence in next section) and are probably a function of amount of synthetic data. They are learning to solve ARC tasks, not learn general abstract reasoning Agreed and th
16.
▲
by
eis
16d ago
The person I replied to compared religion with physics (god vs deterministic computable universe). I said those are not remotely equally defensible theories.
17.
▲
by
eis
16d ago
> Everything leads to paradoxes or unanswerable questions when you think it through. If our universe is computable then what "computer" is it running on? Why is our universe as computationally strong as it is and not more or le
18.
▲
by
eis
16d ago
The notion of god leads immediately to paradoxes and logical contradictions when thinking it through a little bit. It's fine if people have certain believes but let's not put religion on the same level as physics. No such paradoxe
19.
▲
by
eis
1mo ago
I'm confused by your messages linked by DJB. You say that better cryptographers would not choose hybrids, which seems to say that you should indeed think that hybrids are not a good choice. Then you say you are not such a good cryptogr
20.
▲
by
eis
1mo ago
Sure, they are just checkpoints, that much I guess is obvious. The question is why did they not do frequent releases like this before and why are they making significant jumps in benchmarks so fast and all these companies suddenly falling i
21.
▲
by
eis
1mo ago
3.5 Pro was supposed to be around the corner two months ago. 4.0 Pro is some ways out as they recently stated they are seeing some promising early results from training. It didn't sound like a release is imminent.
22.
▲
by
eis
1mo ago
Grok, Meta, Gemini and others all released updates to their models within around a month or two from their respective last release and made significant jumps in benchmarks all around the same time. Any guesses as to why that is? Is it just
23.
▲
by
eis
1mo ago
You post your benchmark on every other AI article, I've seen you do this by now more than a dozen times. It's a bit much. I don't want to be too harsh but your benchmark is obviously flawed when the top 3 models for Typescrip
24.
▲
by
eis
2mo ago
Google wanted to release 3.5 Pro last month but because of the trouble Anthropic got with Fable they might have wanted to wait a bit for the dust to settle I could imagine. And now there is quite some competition. 3.5 Flash for me is a repl
25.
▲
by
eis
2mo ago
I feel like it would be much better if the article focused on QuePaxa because IMHO it's an algorithm that finally brings some novel ideas to concensus (e.g. not relying on timeouts) by kinda coming at it from a gossip protocol angle an
26.
▲
by
eis
2mo ago
But it's not a large-scale public deployment yet either. The article says towards the end that they just ran a proof of concept. Maybe the blog post is just premature. It would be much more valuable if they posted it after actually hav
27.
▲
Beijing is looking at curbing overseas access to China's top AI models
(reuters.com)
64 points
by
eis
2mo ago
|
11 comments
28.
▲
by
eis
3mo ago
I like the initiative but would love if there was a heavily stripped down version, it seems this one needs to download 95mb+. DuckDB also has a WASM version which is also not tiny but comes in at something like 36mb if I recall correctly.
29.
▲
by
eis
3mo ago
I already gave up on Fable 5 because it sometimes was just not worth the editional price compared to Opus 4.8 and other times it flat out downgraded to Opus anyways for no good reason because it thought I'm looking for security vulnera
30.
▲
by
eis
4mo ago
Here's a crucial mechanism that Paul Graham did not mention: With a wealth tax using his calculation, the higher your returns, the lower the comparable income tax would be. If your returns are 10% you'll pay $1 on $10 capital gain
More ›