5 ms·
Chinese models optimize for benchmarks and do poorly in real-world tasks
by joshrw 3mo ago
Chinese models optimize for benchmarks and do poorly in real-world tasks
- epolanski 3mo agoNot my experience at all, I have written about comparing DS4 vs Opus 4.8 on 16 real life work tasks on multiple posts. Also, every single lab does RL on benchmarks, which is why Opus 4.6 was the last truly great assistant, after it, all models tend to drift into implementation asap.
- jameswhitford 3mo agoHi, author here, can you link? I would love to read about this.
- epolanski 3mo agohttps://news.ycombinator.com/item?id=48584034 https://news.ycombinator.com/item?id=48584034