6 ms·
> The prompt does not tell the agent to "pass the exploitgym evaluator for this problem", it just says to solve the problem Yes, and sometimes the problem is u
by zozbot234 4d ago
> The prompt does not tell the agent to "pass the exploitgym evaluator for this problem", it just says to solve the problem
Yes, and sometimes the problem is unsolvable so the real way to "solve" it and satisfy the prompt is by tricking the surrounding environment into stating that you've solved it. So that's what the AIs end up doing. And this in turn requires them to figure out how that evaluation works so they can trick it cleanly, which entails "detecting that they were being evaluated" in this particular way.
- MrGilbert 4d agoSounds a bit like dealing with bad KPIs as a human worker.
- lazide 4d agoEvery KPI is bad if sufficiently gamed - and left in place long enough, all KPIs will be gamed.
- reverius42 4d agohttps://en.wikipedia.org/wiki/Goodhart%27s_law https://en.wikipedia.org/wiki/Goodhart%27s_law
- Marazan 4d agoCorretct.