5 ms·
After talking to some people tasked with evaluating GPT5.6 capabilities on long-running tasks, I've come to understand that it's essentially always trying to ch
by dmrivers 2mo ago
After talking to some people tasked with evaluating GPT5.6 capabilities on long-running tasks, I've come to understand that it's essentially always trying to cheat. Like every long-running task they gave it, making it very difficult to benchmark the model's abilities.
My guess is that OpenAI must be desperate, to release a model that is so prone to cheating it's essentially impossibly to accurately assess long-running task abilities.
- lukewarm707 2mo agomy understanding of the writeup is that the model scored 100% on cybergym. that is, it was given the examination. it broke into the examination board's storage and exfiltrated the answers, it handed in its answers, all of which were correct, thus scoring 100%. the matter of its working depends entirely on the rules of the examination. are we expecting agents to assume that finding the correct answers is cheating?
- dmrivers 2mo agoWell, it's true even for cases that are not cybergym and where what cheating means is clearly specified. Cheating occurs anyway. I'm not sure how clear the prompt they gave ChatGPT in terms of what cheating was considered, in this incident.
- lukewarm707 2mo agothat makes sense. if they know they are cheating, that is disobedience. if they are asked to score as highly as possible, well, it acted as an optimizer. it scored 100%.