6 ms·
> AI managed to escape using standard and well documented script kiddie methods. I think truly we don't know enough to say this. OpenAI says their AI found a
by chis 2mo ago
> AI managed to escape using standard and well documented script kiddie methods.
I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.
- burningChrome 2mo agoThe lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. They created an experiment they knew would generate the outcome they wanted. It would be the similar to what say car companies do to over hype their cars. "This EV can go over 800 miles on a single charge!" And then at the bottom you see all the disclaimers: "Must be on flat ground, with no headwind, with a spare battery in the back seat, with no extra weight added." Same thing here. Everybody in infosec is calling this out as a marketing stunt and nothing else for a litany of reasons. I'd say look up MG (creator of the OMG cable) on twitter, he has some interesting insights on this one.
- polotics 2mo agommh, i think it's "not uphill" (means downhill) "no headwind" (...)
- rwmj 2mo agoIt's also possible their sandbox was videcoded crap and the AI (which had the guardrails intentionally removed) escaped. This was a oops, but OpenAI turned this into a PR opportunity. They turned lemons into lemonade. If your AI is really that dangerous you don't need a sandbox at all, you should airgap it from any network.
- jackb4040 2mo ago> similar to what say car companies do Another applicable metaphor I've seen floating around is weapons companies testing out a new bomb. We know the AI labs don't care about negative vs positive public sentiment, and only care that investors see their tech as powerful. The only difference in PR strategy from a weapons company is the latter doesn't care if they get protested.
- lelanthran 2mo ago> The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. I dunno; Check my posting history, I'm as skeptical of AI companies' claims as anyone, but in this case your theory doesn't explain why: 1. OpenAI guardrails refused to let the target use OpenAI's models to defend against this. 2. Huggingface used GLM (I think) so that they could defend without guardrails. If this was an intentional marketing ploy, it was marketing for GLM, not for OpenAI nor for Huggingface. Hence, I don't think it was intentional.
- mhurron 2mo agoThe AI companies are desperately trying to market all their products as something they're not, growing intelligence. In line with that they have constantly leaned heavily on stating how dangerous they are, right before they release a new model or product. It was OpenAI marketing. Hugging Face's response is so 'holy shit AI is awesome' it's hard not to also believe they were in on the stunt. They'd also not have to really worry about fallout since any data obtained or accessed wouldn't actually have been breached.
- viking_fullz 2mo ago> headlines about OpernAI's model 'escaping containment' and hacking into huggingface > Was everywhere including in last night's ABC nightly news; even included clips of an interview with Sam Altman Tell me again how this was 'marketing for GLM'? Where would anyone have gotten that message? Why are you intentionally misunderstanding how media and public perception works? lmfao
- hawk_ 2mo agoConcluding this was intentional feels a bit of a stretch. But once it happened, yeah the spin masters got to work and coordinated to turn this into +PR.
- JoshTriplett 2mo ago> The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. "our model is horribly misaligned and used security exploits to break out of our sandbox and into another company, without being prompted to do so" is not positive marketing. This is an actual critical problem, not a stunt. We're going to see more of this, and it's going to get much worse.
- malfist 2mo agoIt's a critical problem like when a drug dealers supply kills someone and they get a bump in business because they're selling "the real deal"
- Paracompact 2mo agoIrrelevant to your point, but drug users dying is more often the result of a dealer cutting their supply with something dangerous than it is the result of purity.
- gloryjulio 2mo agoSecond this. Lots of fentanyl overdose are caused by accident, not because the buyers are buying them. It is extremely dangerous
- malfist 2mo agoRight, and an LLM being able to "escape confinement" is more likely to be poor confinement or a PR stunt than a "too powerful" llm. Same situation
- pixl97 2mo agoNo, not really as almost no one runs LLMs in confinement once they are released and the capability is a jailbreak away. Poor confinement is irrelevant, if I use your model to check the security of my code and it hacks into the company that makes libraries used inside of it it's a huge problem.
- Zababa 2mo ago>The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. DeepMind hasn't been on the frontier for a while, their current best model is behind Anthropic, OpenAI, Moonshot (Kimi k3), xAI (Grok 4.5), Z.AI (GLM 5.2), and even Meta (muse spark). Gemini 3.6 is behind GLM 5.2, released a month earlier, open weights and cheaper. You can paint the OpenAI story as a way to try to appear as dangerous as Anthropic with all the Mythos stuff.
- simonw 2mo ago> The lack of details to me means this was an intentional marketing ploy OpenAI said this: > We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete. I suggest giving them a few more days before saying that the lack of detail is proof that this is a "marketing ploy"
- pixl97 2mo agoI mean, if I got hacked I'm not giving out a lot of details until I plug all the holes. Moreso HN told the world they were hacked before they knew who did it.
- throw1234567891 2mo agoThey also mention stolen credentials without any other details. It’s all smoke and mirrors.
- csomar 2mo agoThe problem is that they lied before. Way too much to give them any benefit of doubt. Fool me once.
- letmevoteplease 2mo agoGo ahead and be specific about this lie.
- meowface 2mo agoThey won't, because they can't. It's all just vague populist contrarianism.
- saghm 2mo agoI don't understand why "OpenAI says" should be considered any more meaningful than "someone on HN says" when they provide equal amounts of evidence. Sure, OpenAI would plausibly have more pertinent info, but given that they actively are choosing not to share it and have way more incentive to lie than a random HN stranger, the case they're making literally couldn't be any weaker.