7 ms·
> as shown in the METR report is we are hammering into them they need to complete tasks and doing almost zero checkup if they actually completed the task in the
by TeMPOraL 15d ago
> as shown in the METR report is we are hammering into them they need to complete tasks and doing almost zero checkup if they actually completed the task in the correct manner
I don't think this was the conclusion of that report. On the contrary, the agents were fully aware they're doing wrong. But they also believed the task was impossible to solve correctly, and decided the only way to be sure is to hack the grades, or replace the grader.
- pixl97 15d agoThis was part of the report, but not what the report was about... Why hugging face got hacked was because the agent swarm thought they had to show their work hence the entire need to hack the grader in the first place. Had their realized there was no poison they could have just shared the answer the test was looking for and we'd have never realized (well at least with this particular test) that a huge amount of hidden capabilities were sitting right under the surface. The test makers themselves state the test should be causal to avoid this first order solution hacking. Really continuing on the METR report, OpenAI failed at every level possible here. They are committing nearly every step they can to get a maximally aligned AI.