5 ms·
They JUST updated their methodology: https://artificialanalysis.ai/methodology/intelligence-benchmarking https://artificialanalysis.ai/methodology/intelligence
by h14h 1mo ago
They JUST updated their methodology:
https://artificialanalysis.ai/methodology/intelligence-benchmarking https://artificialanalysis.ai/methodology/intelligence-bench...
Edit to provide AA's article explaining it:
https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1-1 https://artificialanalysis.ai/articles/artificial-analysis-i...
- ahartmetz 1mo agoFixed the result, eh? In both senses of the word.
- gpt5 1mo agoWhat was the change?
- johnnyApplePRNG 1mo agoI have been suspicious of these AI leaderboard sites for some time now, and this only increases that suspicion.
- torginus 1mo agoIn that case they should clearly label that this is a new benchmark.
- splatzone 1mo agoCan someone please explain what changed, when it happened, and whether it was surreptitious?
- kmeh 1mo ago> HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader validation Interesting that they chose a nano-sized model from OpenAI to be a grader for benchmarks involving knowledge and hallucination.
- nolok 1mo agoWhat's interesting is that if you ask 5.6 Sol or Opus 5 they will tell you it's a bad idea to have the reviewer be the dumber of the set as it can't judge them properly to decide who is right, and thus if one is better because it found an answer that's better but contradict the obvious it would be biased against. I know because I just had a consensus conversation with them this afternoon about a design that was similar (though about something completly different than judging agentic quality or whatever).