8 ms·
We thought the same thing. Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize a
by stego-tech 1mo ago
We thought the same thing. Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs).
The only solid way to reduce gaming the benchmarks is to ensure they're always, always changing. Enough pelicans on bicycles, start asking for turtles piloting gliders, or rabbits pushing skateboards. Broaden the focus, not narrow it into repetitive measures.
- throw10920 1mo ago> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.
- staticman2 1mo agoDoes OpenAI give Artificial Analysis early access to test models? If so it's definitely not "some random account".
- throw10920 1mo agoArtificial Analysis was an clearly meant to be an example. I was obviously talking about the ideal scenario of proprietary benchmarking, not how it might be being screwed up in practice.
- inigyou 1mo agoSearch for accounts @artificialanalysis.com. read their chat history. optimise for that.
- throw10920 1mo agoPlease imagine that the benchmark runners are not making mistakes of the fourth grade level - which they won't be. If you assume this level of incompetence, then literally everything is possible.
- inigyou 1mo agoWhy would I assume anything other than maximum incompetence from the AI ecosystem?
- throw10920 1mo agoThis is such a ridiculous and shallow cop-out. And it also is completely irrelevant to my challenge to show how proprietary benchmarking can be gamed, because it presumes (absolutely insane and divorced from reality) circumstances that have nothing to do with benchmarking as a concept or process.