9 ms·
Beware of the benchmarks listed. SciCode and EnterpriseOps for instance: https://shukla.io/blog/2026-08/gym.html https://shukla.io/blog/2026-08/gym.html
by BinRoo 1mo ago
Beware of the benchmarks listed. SciCode and EnterpriseOps for instance: https://shukla.io/blog/2026-08/gym.html https://shukla.io/blog/2026-08/gym.html
- Onavo 1mo agoThe Chinese models also like to cut corners on stuff like science. Their scores on stuff like biotech and scientific knowledge is far from ChatGPT unfortunately. (Claude is pretty good but it just refuses all prompts).