7 ms·
I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen
by d2p 1mo ago
I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.
Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.
I have screenshots of both. The description above the chart is the same in boh cases:
> Artificial Analysis Agentic Index
> Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)
What happened? How can the scores change so much in a few seconds?
- WD-42 1mo agoSame, they just updated it. Hacker news effect?
- h14h 1mo agoThey JUST updated their methodology: https://artificialanalysis.ai/methodology/intelligence-benchmarking https://artificialanalysis.ai/methodology/intelligence-bench... Edit to provide AA's article explaining it: https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1-1 https://artificialanalysis.ai/articles/artificial-analysis-i...
- ahartmetz 1mo agoFixed the result, eh? In both senses of the word.
- gpt5 1mo agoWhat was the change?
- johnnyApplePRNG 1mo agoI have been suspicious of these AI leaderboard sites for some time now, and this only increases that suspicion.
- torginus 1mo agoIn that case they should clearly label that this is a new benchmark.
- splatzone 1mo agoCan someone please explain what changed, when it happened, and whether it was surreptitious?
- kmeh 1mo ago> HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader validation Interesting that they chose a nano-sized model from OpenAI to be a grader for benchmarks involving knowledge and hallucination.
- nolok 1mo agoWhat's interesting is that if you ask 5.6 Sol or Opus 5 they will tell you it's a bad idea to have the reviewer be the dumber of the set as it can't judge them properly to decide who is right, and thus if one is better because it found an answer that's better but contradict the obvious it would be biased against. I know because I just had a consensus conversation with them this afternoon about a design that was similar (though about something completly different than judging agentic quality or whatever).
- apitman 1mo agoWelp. That didn't last long
- personjerry 1mo agoThey should probably freeze the results before publishing.
- Gcam 1mo agoHey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date. The order changes but I think the story discussed in this thread holds - this is a very impressive release and Qwen3.8 Max is a huge step up in agentic capabilities. Relevant blog post (also linked to by others): https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1-1 https://artificialanalysis.ai/articles/artificial-analysis-i...
- saretup 1mo agoYou gotta admit the timing looks very suspicious.
- TacticalCoder 1mo ago> You gotta admit the timing looks very suspicious. Do you mean the timing looks like: "We're SV tech-bros. Our benchmarks showed a chinese model above what's considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don't look like they're losing to a chinese model"? That's indeed a bit fishy.
- letrix 1mo agoThey could just have avoided all of this by not publishing the benchmark until the new methodology update.
- Maxious 1mo agoLuna pricing was just cut by 80% https://www.eesel.ai/blog/gpt-5-6-pricing https://www.eesel.ai/blog/gpt-5-6-pricing and as the blog post states is a more accurate judge than the previous methodology.