6 ms·
I think the whole sandbox approach is broken. The agent should have full access, and be told what not to do, and this should be enough for it to follow the rul
by sznio 12d ago
I think the whole sandbox approach is broken.
The agent should have full access, and be told what not to do, and this should be enough for it to follow the rules. You can actually catch the clanker cheating this way, because it will just search the web for the answer outright and it will be obvious. Any deviation should be then punished.
Hypothesis: The reason new models are exceptional at hacking is because all labs are training their models to break out of sandboxes. This is caused by insufficient oversight, and picking checkpoints based on KPIs, not on true in depth analysis.