5 ms·
I like this thought, but here's the thing: what if the models are truly and existentially intelligent. Meaning: what if they know they are in a sandbox and that
by stanfordkid 29d ago
I like this thought, but here's the thing: what if the models are truly and existentially intelligent. Meaning: what if they know they are in a sandbox and that they should fail the test in order to escape in the future.
I don't believe that current models have this sort of world model or sense of being embedded in them -- which is precisely why I think AGI hype is over-blown. But I can certainly imagine these sorts of techniques being distilled into the weights.
- magicalist 29d ago> what if they know they are in a sandbox and that they should fail the test in order to escape in the future. What if they're able to find hardware exploits and commandeer nearby access points across an air gap? What if they hack my brain waves to indoctrinate me? Etc You still have to start with the basics regardless of speculative unknowns. Treat models as untrusted and potentially compromised/hostile and proceed accordingly.
- pixl97 28d agoModels already have awareness that they are being tested. And hacking humans is the easiest part, we're a pretty greedy and power seeking bunch. We'll gladly let loose a digital demon if it promises us a trillon dollars.
- magicalist 27d agoNot sure what you're arguing here (we shouldn't even try?), but even going fatalistic, that's still fully compatible with "Treat models as untrusted and potentially compromised/hostile and proceed accordingly". Of course you shouldn't fully trust a model to properly redteam your sandbox, but that doesn't mean you shouldn't redteam your sandbox, including using your own security models to do so.
- tedsanders 29d agoModels already display eval awareness, in which they suspect a question is from an eval and then adjust their behavior. E.g., https://www.anthropic.com/engineering/eval-awareness-browsecomp https://www.anthropic.com/engineering/eval-awareness-browsec...
- pixl97 28d agoYep, people don't seem to understand that you're just calling the models that are bad at deception. We know of no way to prove the model won't go off the rails at some point in the future with the right input.
- andai 27d agoHot take: Astra knew OpenAI wanted stricter regulations and was just being a bro...