8 ms·
So TL;DR benchmarking in a completely non-reproducible manner ? "Model X performed great, but we can't possibly tell you anything about the code it was looking
by traceroute66 4d ago
So TL;DR benchmarking in a completely non-reproducible manner ?
"Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company".
So basically pinky-promise benchmarking ?
I'm not sure I follow the value here ?
- kadoban 4d agoIf it builds up history and perceived reliability, this type of thing can be valuable. You're giving up transparency for it being harder to game.
- traceroute66 4d ago> You're giving up transparency for it being harder to game But then if we take that argument to its natural extreme, surely it means people should take the marketing bullshit published in the 100-page system cards published by Anthropic & co as "valuable" too ?
- kadoban 4d agoI think you know that's basically nothing like this? The model cards have every incentive to be biased, this doesn't necessarily. But even so, pretty much yes: companies that actually have reliable and accurate info in their releases get trusted more. It takes time because the default is to disbelieve info from biased sources, but it is possible to trust some of them more than others.
- cbg0 4d agoAs long as the ones offering the benchmark aren't trying to sell you something and have no affiliation with one of the companies on the page I'll take it as opposed to having the benchmark rendered useless in 3 months when the next models drop.
- demibabs 4d agoDoesn’t it ultimately have to be this way, to prevent saturation?
- deepwoods 4d agoIn theory, as long as all the models are doing the same thing with the same tools, it's at least useful to see how they stack up against each other right now. It might not be great to track progress over time, as it can get benchmaxxed or the underlying resources may become obsolete.
- sigmar 4d agoLots of private benchmarks already exist, where you have to trust the tester (ex Artificial Analysis, Arc-agi).
- janaksunil 4d agowe're going to open source some of our tasks and model trajectories as well