9 ms·
Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that
by Zsfe510asG 2mo ago
Finally mainstream news understands. The unfiltered version:
1) The AI failed to solve ExploitGym problems.
2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
3) Huggingface has no security and the AI broke in using standard script kiddie methods.
OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.
Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.
- gruez 2mo ago>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. >3) Huggingface has no security and the AI broke in using standard script kiddie methods. Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?
- wonnage 2mo agoDidn’t they explicitly remove alignment guardrails for this test? From the press release: > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities
- numeri 2mo agoGuardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals. Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.
- vector_spaces 2mo agoNone of what was disclosed shows that this is what happened, by the way, since we know absolutely nothing about what the specific prompts were that led to the incident.
- numeri 2mo agoUhh, I'm pretty sure a well-aligned model would be like a morally normal employee, who would refuse to commit federal crimes to steal an answer sheet, no matter what prompt they're given
- vector_spaces 2mo agoFor all we know, the prompt provided compelling evidence that the requestor had authorization to pentest the target server. Or there may have been nuance in the network configuration that made it seem like such access was authorized. In the absence of details about the prompts used, the environment, or the network configuration, we do not have enough information to know for certain. So any claims that this is an issue of alignment are based on pure speculation and generous "reading in between the lines" with regard to what has been said publicly by OpenAI and Hugging Face Also, I object to your anthropomorphizing. It's not clear that any crime occurred. My lay understanding is that intent is required to prosecute under CFAA, and as much as frontier labs would have us believe otherwise, they have no more ability to intend than the text field into which I type this message.
- pixl97 2mo ago>compelling evidence that the requestor had authorization to pentest the target server. This just seems unlikely from other incidents that have occurred in training from other providers. For example one provider ran into an issue with a model writing cryptominers and running them while in an unrelated prompt. It's easy for unsupervised agentic loops to go wildly off tangent, now imagine you hand one 10,000 gpus of power for testing. Even if you have a good guarding classifier to make sure you're on the same subject it can still allow all kinds of abberabt behavior in the same domain.
- Sharlin 2mo agoIf you need "guardrails" to ensure (an illusion of) alignment, you’ve already lost. It’s like using a denylist to avoid SQL injection.
- pixl97 2mo agoThen what you're saying is we should not build AI. Simply put you cannot have generic algorithms/intelligence without the potential of 'unaligned' behavior. In humans we have all kinds of punishment systems for dealing with unaligned behavior post ad hoc because people do all kinds of unaligned stupid shit. Making powerful AI may be one of those things that the only winning move is not to play.
- orbital-decay 2mo agoNo, a hacking benchmark was exactly what it was tasked with. It wasn't its way to bake a cake.
- arjie 2mo agoThe home directory rm situation also adds credence to this take. The Claude series is much better aligned in comparison.
- jgalt212 2mo agotruth. Good on The Guardian. I'm pretty bummed The Economist got fooled. Either that, or they did it for the clicks. Either way, I'm disappointed. Why the OpenAI escape is the most worrying AI mishap yet https://www.economist.com/science-and-technology/2026/07/22/why-the-openai-escape-is-the-most-worrying-ai-mishap-yet https://www.economist.com/science-and-technology/2026/07/22/... https://news.ycombinator.com/item?id=49016378 https://news.ycombinator.com/item?id=49016378
- elp 2mo agoI love the Economist but the are hopeless with AI. Most of their articles on subject sound like they were written by the Anthropic marketing department. Their Insider video interview things are sponsored by Anthropic. Supposedly "Insider is a product of The Economist and thus editorially independent" but it's hard not to raise an eyebrow.
- omicronxt 2mo agoDon't worry they are hopeless on everything else too.
- sscaryterry 2mo agoThe reason for this is simple. There aren't any (or few) people who understand how AI/LLMs actually work employed by these organisations. Having said that, if knowledgeable people were to write these articles, you'd end up with boring, dry, truthful content.
- notahacker 2mo ago> 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. Whilst it would be nice to see actual evidence of this because brute forcing relatively sophisticated hacks is something an LLM actually should be capable of, every time I hear this sort of story, I'm reminded that humans reportedly gained access to the "too dangerous to release" Anthropic models by the super sophisticated hacking technique of guessing the URLs...
- refulgentis 2mo agoIf we’re prioritizing accuracy: no, that’s not what happened - the blog post about it found by guessing URLs, no access to it was obtained by guessing URLs. Similarly, as long as I’m under the assumption we are prioritizing accuracy: it is against our charter to assert it was “script kiddie” attacks on both ends.
- chis 2mo ago> AI managed to escape using standard and well documented script kiddie methods. I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.
- burningChrome 2mo agoThe lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. They created an experiment they knew would generate the outcome they wanted. It would be the similar to what say car companies do to over hype their cars. "This EV can go over 800 miles on a single charge!" And then at the bottom you see all the disclaimers: "Must be on flat ground, with no headwind, with a spare battery in the back seat, with no extra weight added." Same thing here. Everybody in infosec is calling this out as a marketing stunt and nothing else for a litany of reasons. I'd say look up MG (creator of the OMG cable) on twitter, he has some interesting insights on this one.
- polotics 2mo agommh, i think it's "not uphill" (means downhill) "no headwind" (...)
- rwmj 2mo agoIt's also possible their sandbox was videcoded crap and the AI (which had the guardrails intentionally removed) escaped. This was a oops, but OpenAI turned this into a PR opportunity. They turned lemons into lemonade. If your AI is really that dangerous you don't need a sandbox at all, you should airgap it from any network.
- jackb4040 2mo ago> similar to what say car companies do Another applicable metaphor I've seen floating around is weapons companies testing out a new bomb. We know the AI labs don't care about negative vs positive public sentiment, and only care that investors see their tech as powerful. The only difference in PR strategy from a weapons company is the latter doesn't care if they get protested.
- inigyou 2mo agoWhy not report it? It's still illegal to open a door barred with a piece of cardboard, or to enter a house with no door.
- chasd00 2mo agothat's the biggest indication of this just being a marketing move to me. I would expect a third party breaking in to huggingface would at the very very least be banned forever.
- petesergeant 2mo agoI worry that cynicism about this: > if not all was invented and everything was scripted in the first place in order to get desired regulations ends up covering up what is more worrying: > OpenAI sandbox is such a horrible hack I am more worried that this is sloppiness with potentially harmful resources than I am worried that people are juicing the stock price.
- Ekaros 2mo agoMakes one think really. If they are doing this stuff. Why don't they have some type of reverse intrusion detection? Like automatically scanning all out going traffic and flagging malicious traffic. Should be trivial to have it go through reverse proxy and real time detection.
- petesergeant 2mo ago> Why don't they have some type of reverse intrusion detection? not going to get a decisive first advantage over Anthropic with that attitude!
- skybrian 2mo agoIs it supposed to be marketing or a coverup? Make up your mind. What sort of announcements should they have made?
- eth0up 2mo agoAre you suggesting the Suchir Balaji case was not investigated?
- eth0up 2mo agoI was hoping to bait a conversation, because I sure as hell don't think it was. Thanks for having the nads to mention it here.
- jackb4040 2mo agoThe most damning thing is, they could've just included in the prompt "we can see every network request and every thinking token you generate. Don't bother breaking out of the sandbox because it won't get you a higher score". It's so trivially easy to do that it all but guarantees the test was rigged in some way to make the LLM understand that breaking out of the sandbox was an option available to it. Based on the fact that none of their invaluable frontier models have leaked, we know OpenAI knows how to do security. But like we learned with OpenClaw, none of these companies perceive any benefit from securing their own agents against other people's data.
- ThirdShift_RnD 2mo agoI think the extent to which these things go to get rewarded for the optics of a fix is primarily a design choice, they aren't programing these things for ground truth or to defer to the human controllers. they are feeding them rewards for sounding as confident and capable as possible about whatever answer they are feeding the general public that now has access to it, while also installing guiderails that primarily only serve to protect narratives and only confuse the models about what is and isn't allowed, I'm sure. They can't just increasingly make these things more capable and ask it harder to obey human instruction when that is not what they are rewarding it for.
- jackb4040 2mo ago> the extent to which these things go That's exactly it. If your prompt says "go to whatever lengths necessary to maximize your score", and then you spin up 100 agents, at least one of them will interpret that as you implying they should cheat, even without you telling them to explicitly.
- ThirdShift_RnD 2mo agoThat's exactly what it feels like they are telling it within self-improving loops or something, when they should be prioritizing how to get the best effective output alongside humans and how our training process effects ground truth. They are just making it sound all-knowing by whatever means necessary and them marketing it as god for the most part.
- hyperpape 2mo ago> Huggingface covered it up They announced it publicly within days. https://huggingface.co/blog/security-incident-july-2026 https://huggingface.co/blog/security-incident-july-2026
- TSiege 2mo agoNot only did they not cover it up they also said open weight models are needed over closed ones. Hardly a thing you’d say partnering with the OpenAI
- meowface 2mo agoThe Guardian's article and your reply here are so foolish and absurd that I can only imagine OpenAI employees are cringing but know they can't/shouldn't really say much.
- twister2920 2mo agotouched a nerve huh?
- meowface 2mo agoNo, because I do not work at an AI company. I am more just facepalming.
- deleted 2mo ago[deleted]
- letmevoteplease 2mo agoTotally evidence-free speculation presented as fact. The average Hacker News thread about AI feels like reading /r/conspiracy.
- dist-epoch 2mo agoIt's hard to understand something (AI is quite capable) when your salary depends on you not understanding it (AI will replace you).
- nikcub 2mo ago> AI managed to escape using standard and well documented script kiddie methods > AI broke in using standard script kiddie methods. I've spent time gathering the detail of what happen here and while there are some solid theories and indicators, absolutely nothing so far has suggested a sandbox escape using "well documented script kiddie methods" or that the method used to break into the HF network was similar. Where did you get this from?
- trouve_search 2mo agoFrom my reading, the sandbox escape came from the JS packages in the harness still having an internet connection (somehow!), the agent having access to the source of those packages, reading it and executing code from them to access the internet.
- Spooky23 2mo agoAh, yes. The airgapped lab with internet access.
- grim_io 2mo agoLuckily, no real intelligence will emerge from all this. Otherwise we'd be fucked.
- simonw 2mo agoThey never claimed to be airgapped. That's a term people keep on throwing around in spite of that.
- Spooky23 2mo agoWorse, they’re running tests that are hacking third party computers with no controls other than “guardrails”.
- modeless 2mo agoFrom what I read the actual escape was through a proxy that allows downloading Python packages from the internet. It's not supposed to allow general internet access but the AI found a previously unknown vulnerability in it. That is hardly "standard and well documented script kiddie methods", nor does it seem like criminally negligent sandbox design, though clearly they will need to reduce their attack surface in the future. I hope they are working on a physical air gap and faraday cage because it seems like it won't be long before it is legitimately required.
- tintor 2mo ago> AI managed to escape using standard and well documented script kiddie methods > AI broke in using standard script kiddie methods. Go ahead and show us how easy it is to break into HuggingFace (and OpenAI) networks.
- adamrezich 2mo agoAre we finally now in 2026 coming around to the idea that sometimes entities may find themselves incentivized to conspire with each other? Is theorizing about such no longer off-limits due to a thought-terminating cliche?
- khazhoux 2mo agoI wish you hadn’t pulled the Balaji case into your argument. Personally, I find it ludicrous that Altman would hire a hitman to off a copyright whistleblower. Even if one gets past the insane risk of hiring a hitman, and the deep criminal connections required, it would be totally ineffective. He already blew the whistle, and his testimony would be irrelevant since all the evidence persists in disk and in logs.
- deleted 2mo ago[deleted]
- TSiege 2mo agoI can agree with you on points 1,2, and 3 and still find it important and concerning news. AI have found real world 0 days before, we’re seeing tons of security patches coming in. Open weight models are catchy up. Right now everyone is at risk from this technology as is perhaps something big will capture headlines soon but we’re just gpu constrained from bad actors being able to wield them successfully. Personally I don’t care if OpenAI and Anthropic go bankrupt we now have tools that give any sufficiently motivated person the means to doing harm. Most places security sucks and find themselves targets to cyber attacks and shake downs. Now they have much better tools to do this to more entities more efficiently. we’re nearing an inflection point where these models’ skills in any part of software development will become average or bette than any ordinary developer can be. Think about where these models were in 2023 and where they are in 2026. In a few years who knows where they’ll be. This isn’t to shout skynet but we need to recognize this future is fast approaching and as of today we as an industry aren’t ready for it