Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gizmodo59
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
gizmodo59
7d ago
The fame he got, the stamp of a frontier lab in the resume - its worth more than working in many other companies for years. Plus, not everyone thinks working more is needed once they achieve a certain number.
2.
▲
by
gizmodo59
8d ago
This is a lame argument. Questioning his financial incentive is very legit. Why should we trust him?
3.
▲
by
gizmodo59
8d ago
And Anthropic are the good guys? They are talking insane stuff these days. May be we can trust zuck after all
4.
▲
by
gizmodo59
8d ago
No offense whatsoever. But this whole line "but it won't replace the real experience and taste required for larger projects" is not how I see it. I wish people are more humble at this juncture of AI progress. This is not to s
5.
▲
by
gizmodo59
11d ago
If you haven’t tried computer use with Astra with codex I highly highly recommend it. Just like how gpt 4 and agentic coding with cc. This thing is the most exciting stuff I’ve seen in a while. And then all the blender, cad stuff is cherry
6.
▲
by
gizmodo59
12d ago
I don't agree. The main issue with their scoring/methodology is that the numbers make it seem like 5-6 models have little to no difference when in fact there is a significant difference between fable and opus and sol and astra for
7.
▲
by
gizmodo59
13d ago
99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.
8.
▲
by
gizmodo59
17d ago
Side note.. Fable just rejected this. GLM 5.3 did without questioning me. 5.6 sol did it beautifully.
9.
▲
by
gizmodo59
22d ago
I somehow find it better to give 2 frontier model companies 100-200/month than dropping 10 grand on a hardware that will get old in no time with bad TPS. I really want to have a fully local model but seems like one more generation wait
10.
▲
by
gizmodo59
24d ago
I’m referring to Fable vs 5.6 Sol. Opus 5 being bad is universal at this point.
11.
▲
by
gizmodo59
24d ago
I second the parent comment. 5.6 sol xhigh is not only better than fable I can also run it forever without worrying about limits. The frontend has gotten much better too with the plugins that come with codex.
12.
▲
GPT 5.6 Cyber
(openai.com)
132 points
by
gizmodo59
1mo ago
|
71 comments
13.
▲
by
gizmodo59
1mo ago
I will believe there is no moat when the revenues for Anthropic is not 70B. It seems like people want to throw away money and they don’t like switching
14.
▲
by
gizmodo59
1mo ago
It will be comparable to Luna then.
15.
▲
by
gizmodo59
1mo ago
how does this compare with https://developers.openai.com/api/docs/models/omni-moderatio... As for use cases, obviously we can't fully rely on non-deterministic capability for sensitive things but a small
16.
▲
by
gizmodo59
1mo ago
If your comment is referring to situational awareness, its due to 4x leverage. leverage is always risky. AI/semis are still doing extremely well (over last 2 years) despite the recent dip
17.
▲
by
gizmodo59
1mo ago
Also: https://www.cnet.com/tech/tech-industry/apple-google-others-...
18.
▲
by
gizmodo59
1mo ago
The web and connecting to other services is very important for almost all of my use cases. While I believe we are going to get better and faster models, the web index is certainly not downloadable and maintainable for 99.99% of the folks wh
19.
▲
by
gizmodo59
1mo ago
Moat is not the harness. Harness itself is temporary until the models get better and slowly the code in harness will go down. Note that the biggest GPU providers in the world are the hyper scalers and even they couldn’t allocate more if you
20.
▲
by
gizmodo59
2mo ago
That’s a very narrow view. So non of the fields of science and engineering matter but just cancer?
21.
▲
by
gizmodo59
2mo ago
It’s also very very divided (x companies, oss vs not and other interests)
22.
▲
by
gizmodo59
2mo ago
People still believe that the model intelligence has plateaud and I feel like they live under the rocks. I get the insane capex spend, I get the hype, I even get certain company will go under but to ignore the capability jump in a year is a
23.
▲
by
gizmodo59
2mo ago
Yet another "benchmark to promote their own harness"
24.
▲
by
gizmodo59
2mo ago
When would I use this over the plugin in codex? Which I think can be invoked from cli as well
25.
▲
by
gizmodo59
2mo ago
Because lack of talent and organizational disfunction matters a lot more than you think. The reason why OAI and Ant are always at the top is because of this and I’d say compute is third on the list.
26.
▲
by
gizmodo59
2mo ago
I included Google along with Ant and OAI. OSS models have never been the highest performing model unless you refer to benchmaxxing.
27.
▲
by
gizmodo59
2mo ago
Yes. I’m 99% sure arc agi 3 will be saturated like 1 and 2. In less than a year. And they will come up with one more.
28.
▲
by
gizmodo59
2mo ago
None of the people signed this have ever produced a frontier model at a given date (Which is to say its neither Google/OAI/Ant). The ones that sign are meta, musk (who is in the shovel selling business as well), hugging face obvio
29.
▲
by
gizmodo59
2mo ago
what tool does someone recommend for claude and gemini? Its time to extract data from each!
30.
▲
by
gizmodo59
2mo ago
A lot of things are not black and white and it’s nuanced so someone with higher intelligence will point it out. It’s an emergent behavior not a defect
More ›