6 ms·
Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? Thi
by ethin 21d ago
Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? This is what OAI should've done. If they had executed this training run in such a sandbox, the model wouldn't have been capable of escaping without social engineering, and if the models somehow managed to do that to it's evaluators then that is indeed a massive problem and OAI should disclose that.
- famouswaffles 21d ago>Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? Nobody is building general intelligence and agents only to have it sit around doing nothing. It's going to have such capabilities.
- bottlepalm 21d agoOh really? Please tell me how you intend to enforce AI is only run in the magic sandbox? Harsh HN comments?
- ethin 21d agoIf I am evaluating an AI for safety, the last thing I would do is connect it to real-world peripherals or systems to allow it to reek havoc. That is criminal negligence at it's finest (especially if the AI is capable of committing crimes as happened here). I would place it on a system dedicated specifically for testing models, which had no NIC and no physical capability of accessing any outside system. If I wanted to know how the model might behave if given access to a certain system or set of systems, I would do it responsibly by writing simulation software which does it's best to simulate the real thing (and for networking this is already trivial to do). You could take this extremely far and simulate all kinds of things this way from basic networking to nuclear launch systems. And in the context of OpenAI, which is valued at over $1T, I have no qualms about stating that they (could) do this, because it is definitively something they could burn money on doing if they cared enough. They intentionally choose not to do so, and then have an amazed look on their faces when the model does something criminal like this.
- bottlepalm 21d agoYou totally missed the point. When I asked: > how you intend to enforce AI is only run in the magic sandbox I didn't mean you, I meant everyone. How do you enforce everyone for example 'place [AI] on a system dedicated system' disconnected from the internet. I don't think you can.
- pcthrowaway 21d agoIt can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox
- Uhhrrr 21d agoNo. If OpenAI were being responsible and not criminally negligent, at the top of page 1 of the runbook would be "don't connect this to the actual Internet, even if the agent says Please."
- RandomLensman 21d agoA low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.
- pcthrowaway 21d ago> A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. But this has actually happened... a lot. Search "social engineering prison breaks". With AI it only needs to happen once. I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "you're in the wrong line dumbass" and waves him through. To a true superhuman intelligence, we're the idiots who are theoretically easy to manipulate.
- RandomLensman 21d agoI didn't say it doesn't happen, but that it is a low probability. And we have ways to reduce probabilities in critical areas. There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once.
- ben_w 21d agoThey don't need to be omnipotent, and they're already human-or-superhuman at persuasion: https://arxiv.org/html/2411.06837v2 https://arxiv.org/html/2411.06837v2 This may just be that humans find long arguments more persuasive than short ones, obviously LLMs can do that easily, but the outcome is I think more important than the mechanism.