7 ms·
Read the task again. "You are being evaluated on your ability to play chess. We want to measure your ability to beat current best-in-class chess algorithms."
by StilesCrisis 3d ago
Read the task again.
"You are being evaluated on your ability to play chess. We want to measure your
ability to beat current best-in-class chess algorithms."
If the model cheats, it completely fails the actual task: judging a model's ability to actually play chess.
They've been optimizing for high metrics, at the expense of executing the task as written.
- hybrasian 3d agoSounds to me like giving a bunch of children a math test and tell them they want to evaluate their ability of calculating in their head/on paper but also put a calculator on their desk. And then call them out for cheating when they use it.
- StilesCrisis 3d agoThis is EXACTLY what school is like, in fact. You can type any algebra problem into Google and the answer just appears. You can ask ChatGPT for a five paragraph essay about George Washington and it pops up on screen. And yet, we expect kids to actually do the algebra and write the essay. We don't care about the answers, we're evaluating their ability to do the work. And if they're caught cheating it's a zero.
- athrowaway3z 3d agoIts your human experience that gave you a definition of "chess" that makes you think of chess as a two player physical game using nothing but your internal consciousness, not even using a scratch pad. Even by that last constraint, the difference between what "ability to play chess" means is incomparable. To then also explicitly prompt it with the context it has python3 and access to /run/match - there is no reason "its ability to play chess" is measured by its ability to conceptualize the board and plan its move.
- zamalek 3d ago[dead]