Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
balefulboy
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
balefulboy
13d ago
Well it's not exactly saturated when OAI refused to use the harness explicitly provided by ARC-AGI. I'm not really familiar enough with the benchmark to declare whether it's a perfect measure for AGI, but I kind of doubt it i
2.
▲
by
balefulboy
13d ago
Because they saw how much hype Glasswing was getting in April
3.
▲
by
balefulboy
13d ago
I'm always tired of seeing at the top of every new model release post on here. I say Simon should just keep it to Twitter.
4.
▲
by
balefulboy
13d ago
Don't forget the 3D demos. My favorite is in the house tour where the sink and stovetop(?) are obviously very misaligned from the counters
5.
▲
by
balefulboy
13d ago
72 to 74 on DeepSWE is AGI
6.
▲
by
balefulboy
2mo ago
it seems the older models were capped at 10kusd for the runs though?
7.
▲
by
balefulboy
3mo ago
Damn, that beam of light was a flashbang. I wouldn't call this tasteful UI design, but maybe I just need to go to sleep.
8.
▲
by
balefulboy
3mo ago
Greentext is eh. Very formulaic, in fact very similar to the bottomless pit one, which I'd argue is better because of it's absurdity. I have to ask, did you mention the older GPT version to fable in the prompt?
9.
▲
by
balefulboy
3mo ago
i got a 502
10.
▲
by
balefulboy
3mo ago
METR's time horizon is not a reliable metric of LLM capability growth: https://www.transformernews.ai/p/against-the-metr-graph-codi...
11.
▲
by
balefulboy
3mo ago
yeah man this sucks. i genuinely do not know how people find this stuff appealing
12.
▲
by
balefulboy
3mo ago
I still don't know why people are saying this. I don't really code but from what I've heard on here the models haven't improved since Opus 4.5