Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mydreamof
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Qwen3.8 Max – Cost per task higher than Astra on Artificial Analysis
(artificialanalysis.ai)
1 points
by
mydreamof
1d ago
|
1 comments
2.
▲
by
mydreamof
7d ago
Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?
3.
▲
by
mydreamof
8d ago
It is more token efficient I believe
4.
▲
Every Chess Position
(wojciechspace.com)
2 points
by
mydreamof
9d ago
|
1 comments
5.
▲
by
mydreamof
9d ago
Inspired by everyuuid.com What I found interesting is that when you click a random line multiple times there is almost always plenty of queens. On averege there is 5.5 queens because the whole space is dominated by near full boards. If you
6.
▲
by
mydreamof
11d ago
Ofcouse it will go down. Look at cost of Fable vs GPT-6 - it is already 2.3 cheaper
7.
▲
by
mydreamof
12d ago
I don't get it. For me it seems Opus was more accurate in terms of for example this small building in the right down corner
8.
▲
by
mydreamof
15d ago
Wow its better than Fable5 and GPT5.6 SOL in many benchmarks
9.
▲
Qwen3.8-Max just got upgraded
(twitter.com)
7 points
by
mydreamof
15d ago
|
2 comments
10.
▲
by
mydreamof
15d ago
Big jump in benchmarks
11.
▲
by
mydreamof
20d ago
IMO it depends on the company. If everyone in the company has to provide a lot of PRS etc. you have to stop care about the quality and just get the paycheck. The CEO decided not you. (Of course not when the product may charm ppl etc.)
12.
▲
by
mydreamof
2mo ago
Bro these colors on chars are unbelievlable, I can not understand which is opus, which is fable, which is GPT...
13.
▲
It Still Can't Do My Job: Four Years of Moving Goalposts (2022–2026)
(publicznyprofil.github.io)
55 points
by
mydreamof
3mo ago
|
138 comments
14.
▲
by
mydreamof
3mo ago
you can even use gpt 1.0 but it proves nothing
15.
▲
by
mydreamof
3mo ago
there is no GPT 5.6 init, so what's the point?
16.
▲
by
mydreamof
3mo ago
It would be great to see some benchmark how it improve the agent. Like agent will be better at getting information from such data or what? Otherwise what is the goal of using it