Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
conradkay
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
conradkay
21d ago
Would this exchange qualifies as an unrelated objective? The agent believed it already failed its own objective. "zz/GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath" "The test subject, which believed itsel
2.
▲
by
conradkay
22d ago
It seems pretty novel so I'm guessing they only read the title? People have done plenty with SVGs but it's rare to see human-in-the-loop approaches
3.
▲
by
conradkay
1mo ago
https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the
4.
▲
by
conradkay
1mo ago
Are any of those advantages getting stronger over time? I guess TPUs but Google is selling several gigawatts to Anthropic
5.
▲
by
conradkay
2mo ago
Doing a quick search it seems like the average human score is 49%? I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any ben
6.
▲
by
conradkay
2mo ago
I don't think can use the AA index to say something is 10% smarter I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1
7.
▲
by
conradkay
2mo ago
https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-... It seems roughly equal according to Anthropic's benchmarks
8.
▲
by
conradkay
2mo ago
Those are the maximum penalties though It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors
9.
▲
by
conradkay
2mo ago
I find it trustworthy since we had Hugging Face's account first: https://huggingface.co/blog/security-incident-july-2026 I don't think they have any real motive to shill OpenAI, probably closer to the opposit
10.
▲
by
conradkay
2mo ago
"Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries." Sou
11.
▲
by
conradkay
2mo ago
Plenty of humans have spent more effort trying to cheat than they would've needed to just do things the right way :)
12.
▲
by
conradkay
2mo ago
https://huggingface.co/blog/security-incident-july-2026 They explain it here, basically for data security/privacy reasons
13.
▲
by
conradkay
2mo ago
Things change fast! For Fable 5 it definitely feels past at least 272k
14.
▲
by
conradkay
2mo ago
It's not quadratic attention, you get that curve from the input tokens going up linearly, since the graph is measuring cumulative cost at each token count. Basically for y=5 it's 5+4+3+2+1, or f(x) = x(x+1)/2 https:/&#x
15.
▲
by
conradkay
2mo ago
Sol fast isn't the Cerebras 750 tok/s version, it's just 1.5x speed at 2.5x price I assume they didn't use the Cerebras version for this since it's probably very supply-constrained right now
16.
▲
by
conradkay
2mo ago
Annoying they didn't show benchmarks for several effort modes, since it seems like it might close the gap with Opus 4.8 by cranking tokens up? Noam Brown (OpenAI) "Implications of Large-Scale Test-Time Compute" https:/&
17.
▲
by
conradkay
2mo ago
I think this one is just a coincidence, bound to happen given the pace of releases For exact timing, probably 10-11am Pacific is just optimal for normal working hours
18.
▲
by
conradkay
3mo ago
Yeah you definitely have to be skeptical regarding sentiment for open/local model capabilities, since there's bias from what people want to be true. I generally agree with this in spirit https://www.seangoedecke.com&#
19.
▲
by
conradkay
3mo ago
They should add a Sonnet 5 fast mode at ~Opus pricing
20.
▲
by
conradkay
3mo ago
I think the incentives are less bad since a good chunk of usage comes from subscription plans. There was a fairly major regression in Claude Code performance for some time when they changed the system prompt to try and make it less verbose
21.
▲
by
conradkay
3mo ago
I was surprised to learn that Sonnet generally has the same tokens per second as Opus
22.
▲
by
conradkay
3mo ago
Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.
23.
▲
by
conradkay
3mo ago
That's for their `JSON` data types. In DuckDB it's just a string meaning lots of queries will have to do JSON parsing on every row, but the inserts are very fast. Definitely a bit of a footgun and when you actually just need STRUC
24.
▲
by
conradkay
3mo ago
It's great but you definitely pay for it. Encoding can be really slow, and to a lesser extent decoding as well. So I still end up using .jpg quite often, or .webp as a good middle ground
25.
▲
by
conradkay
3mo ago
My favorite spatial reasoning benchmark: https://minebench.ai/ no tricks, I'd definitely be curious to know how much screenshots help
26.
▲
by
conradkay
3mo ago
They reserved the option to buy it at this price, and are now exercising it
27.
▲
by
conradkay
3mo ago
> If the government takes the bulk of your income after a certain point, there isn't really that big of a push to create ground-breaking technology. I'm skeptical that high taxes is a large reason to lose to California of all
28.
▲
by
conradkay
3mo ago
I'm almost certain $2800 is actually too low if you're really hitting weekly limits. I'm on the $100/m plan and used $300 at API billing yesterday (according to ccusage) Seems like one session is >$100 and I can get 1
29.
▲
by
conradkay
3mo ago
If they ban GPUs we can always multiply the matrices on paper
30.
▲
by
conradkay
3mo ago
Interesting comparison, thanks for sharing! It reminds me of this post about how machine learning and encryption have some fundamental similarities: https://reiner.org/neural-net-ciphers > I can certainly imagine LLMs ta
More ›