7 ms·
GPQA Diamond: gpt-oss-120b: 80.1%, Qwen3-235B-A22B-Thinking-2507: 81.1% Humanity’s Last Exam: gpt-oss-120b (tools): 19.0%, gpt-oss-120b (no tools): 14.9%, Qwen
by Leary 1y ago
GPQA Diamond: gpt-oss-120b: 80.1%, Qwen3-235B-A22B-Thinking-2507: 81.1%
Humanity’s Last Exam: gpt-oss-120b (tools): 19.0%, gpt-oss-120b (no tools): 14.9%, Qwen3-235B-A22B-Thinking-2507: 18.2%
- amarcheschi 1y agoGlm 4.5 seems on par as well
- thegeomaster 1y agoGLM-4.5 seems to outperform it on TauBench, too. And it's suspicious OAI is not sharing numbers for quite a few useful benchmarks (nothing related to coding, for example). One positive thing I see is the number of parameters and size --- it will provide more economical inference than current open source SOTA.
- jasonjmcghee 1y agoWow - I will give it a try then. I'm cynical about OpenAI minmaxing benchmarks, but still trying to be optimistic as this in 8bit is such a nice fit for apple silicon
- modeless 1y agoEven better, it's 4 bit
- lcnPylGDnU4H9OF 1y agoWas the Qwen model using tools for Humanity's Last Exam?