Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
irthomasthomas
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
irthomasthomas
4d ago
> For instance the task naming in the task file starts with an optimistic 1, 2, 3, 5, 5a but then eventually gets to 8a, 8a1, and then ends up with 8b2c2b3 and “8b2c2b2b checkpoint1”. The code that it produced got ever more wild. I don’
2.
▲
by
irthomasthomas
4d ago
Astra scores the same on DeepSWE 1.1 (~75%) as Gemini Flash 3.8 and Deeepseek Flash 4.1 So general coding ability has plateaued, for now. Also consider the context windows. 1M token models where a breakthrough two years ago. Today they are
3.
▲
by
irthomasthomas
6d ago
Is there a reason they scoped that so narrowly to Buckmaster/codex/2 months two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem w
4.
▲
by
irthomasthomas
6d ago
Not on par, but in the same league. Astra is way ahead on visual tasks, but scores the same as gemini and deepseek on DeepSWE.
5.
▲
by
irthomasthomas
6d ago
Openai said that a new model became available to them during this. But that could mean anything from a big new base model to a LoRA, fine-tuned on a few dozen prompts...
6.
▲
by
irthomasthomas
6d ago
hmm I'm hoping there is a bug on their API because my first impression is not good. I asked it to return bash code between <bash></bash> tags. It is failing frequently and writing it's own tool calling format inst
7.
▲
by
irthomasthomas
6d ago
Quite a flex calling their GPT-6 competitor "Flash"! But it is faster than their last flash model due to a combination of architectural innovations including engrams and a new encoder/decoder design that uses 8B parameters f
8.
▲
by
irthomasthomas
7d ago
Has the method for extracting the COT been blocked, now? Otherwise why could we not generate some fresh samples?
9.
▲
by
irthomasthomas
7d ago
Chutes.ai models are served from a Trusted Execution Environment, so the GPU owners can't see your prompts.
10.
▲
by
irthomasthomas
7d ago
And deliberate or not it is still plagiarism by the sound of it.
11.
▲
by
irthomasthomas
8d ago
The researcher told them it was an independent effort, and they still pushed ahead with it.
12.
▲
by
irthomasthomas
8d ago
This should make an excellent choice for arbiter in llm-consortium, mercury-2 was pretty good. One of the main drawbacks of the multi-model system is the added latency of the llm judge, but having a model run at 1100tps goes a long a way to
13.
▲
by
irthomasthomas
8d ago
If the goal was not to scoop them, why did openai put a massive team on this, working weekends, only after they heard rumors of the solution?
14.
▲
by
irthomasthomas
8d ago
Or they trained a LoRA on the victims chats in order to launder their plagiarism.
15.
▲
by
irthomasthomas
8d ago
It can still be academic plagiarism even if they ticked the box to allow training on their prompts.
16.
▲
by
irthomasthomas
8d ago
Doesn't that count as plagiarism?
17.
▲
by
irthomasthomas
8d ago
"When a further trained version of our internal model became available over the course of the effort, we updated our agents to that model." woah, this gives some credit to the rumor that openai finetuned a model over the course of
18.
▲
by
irthomasthomas
8d ago
They where working on the problem for a year using codex.
19.
▲
by
irthomasthomas
8d ago
Attack is the best form of defence
20.
▲
by
irthomasthomas
8d ago
Is it opensource?
21.
▲
by
irthomasthomas
11d ago
You need to literally review the review with another llm pass to push back on the first. Ask it to do something like reassess the severity claims and only surface real P0 to P2 issues.
22.
▲
by
irthomasthomas
12d ago
I don't think so: https://news.ycombinator.com/item?id=49538217
23.
▲
by
irthomasthomas
12d ago
Technically on par with gemini flash 3.8, but I give Astra more points for style, and for breaking from the pack by not adding the headgear, and the fish in a basket.
24.
▲
by
irthomasthomas
12d ago
Which is why torrenting is the dumbest way to do this, since every downloader is also an uploader. That is what gets torrenters locked up. It is possible to leach only, but its not the default mode, and if this guy was broadcasting his IP a
25.
▲
by
irthomasthomas
12d ago
Because of the ongoing training costs. They are certainly making a healthy profit margin on inference.
26.
▲
Nvidia's AVO harness boosts Opus 5 to 100% on ARC-AGI 3
(developer.nvidia.com)
1 points
by
irthomasthomas
12d ago
|
0 comments
27.
▲
by
irthomasthomas
13d ago
A model that can't beat gemini flash 3.8 on deepSWE is not AGI. I would not be surprised if ARC skills don't carry over to real tasks. In that case, training for ARC could even hurt real world performance. I have't looked in
28.
▲
by
irthomasthomas
13d ago
ARC does not test for intelligence, only for the lack of it. A model that scores high MAY be AGI, while one that scores poorly cannot be AGI. That is all this test can tell us.
29.
▲
by
irthomasthomas
13d ago
Without prompt caching this becomes more expensive than fable 5.1 after turn 50, assuming you start with 40k tokens and add 2k per turn.
30.
▲
by
irthomasthomas
13d ago
It's going to cost a fortune in opencode without prompt caching.
More ›