11 ms·
There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without th
by dwoosley 2mo ago
There seems to be three popular ways to view this incident.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
- jackb4040 2mo ago> positive for OpenAI To echo OP's article, these companies have proven time and time again that they DO NOT CARE if people like them, they only care that investors believe their technology is powerful. Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1. I know that may seem like a nitpick, but their entire media strategy relies on this. If they can convince you they're taking a risk by disclosing these stories when they're actually not, they can inflate their own credibility. Point #3 is what actually happened, but it will never be possible to prove. The only hope we have is that a decade in it'll get harder to convince people that the revolution is just around the corner. The fact that we're getting this from the Guardian already is a good sign.
- Terr_ 2mo ago> Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1. I think that depends on the interests and sophistication of the subgroup-of-investors. If the investor is hoping for AI that can be trusted to run a bank, they don't want one that can get twisted into giving away money because a customer has been talking about the path to enlightenment and salvation through abandoning worldly attachments.
- jackb4040 2mo agoAI investors are not sophisticated users nor product managers. They are by and large not technical at all. They are bureaucrats at a teachers' pension fund in the midwest, unscrupulous dealmakers at private credit firms, and Masayoshi Son. Actually go read what Masayoshi Son says about AI if you want to understand the level of due-diligence we're dealing with.
- Terr_ 2mo agoOh, I'm already quite aware that there are very questionable forms of goose biology involved...
- pixl97 2mo agoMost banks are worried that AI can hack their security, and when using unrestricted models in testing the models are doing a decent job of it.
- anigbrowl 2mo agoI'm inclined to agree. I ran across a picture of Sam Altman's face combined with Elizabeth Holmes' hairstyle the other day, and imho it was providing a significant premium to the usual 1000 words:picture exchange rate.
- antonvs 2mo agoFor all the criticism that can reasonably be leveled at OpenAI, at least they have a real, working, powerful product, unlike Holmes. In fact that seems to be key to the most successful 21st century grifts: build a pile of nonsense around real products to inflate valuations. All the nonsense that Altman, Amodei, and Musk spout is to stoke the fires of FOMO and blow hot air into the bubbles.
- leumon 2mo agoIn any case it just shows that these models aren't properly aligned. Instead of trying to solve tests they try to find ways to cheat.
- sfink 2mo agoBut in my experience, that's what problem solving is like? You have a goal that you don't know how to get to. You come up with any way you can think of to reach that goal, and try out the ones you think might work. The effectiveness of AIs at coding is a direct result of the fact that they are less constrained than humans at deciding which approaches are "reasonable". They are absolute beasts, fearless beasts. They'll write thousands of lines of code to do things that often shouldn't be done, or should be done with a library, or should be done by simplifying the problem statement. They'll add debugging to every level of a stack, they'll rewrite core libraries, they'll reconfigure your machine and network if something is broken or disallowed. How are they supposed to distinguish broken vs disallowed, anyway? That would just use up processing power, and they work by maniacally focusing all of that power on their goal and not getting slowed down by other considerations. If they write a quadratic algorithm that times out before finishing a test, is it cheating to rewrite it to be linear? How do you define "cheating", and how much intelligence is required to constantly evaluate whether or not something qualifies as such? I'm actually in agreement that alignment is critically important, the more so the more powerful these things become. I just don't find cheating to be a very good example of something to be solved with alignment. It could be, but it would lobotomize the model enough to make it useless.
- pixl97 2mo agoI mean they also teach the models how to find and exploit other systems as this is lucrative and governments will pay top dollar for it. Alignment is in the eye of the beholder.
- lokar 2mo agoIsn’t that just making a distinction between the output and how it was produced? Chinese room again
- antonvs 2mo ago
- ctoth 2mo agoWhy do none of these few constrained ways to view this complex situation (nice gig if you can get it, agenda setting) include "and also this looks a heck of a lot like the stuff that the LW folks have been warning about for years and maybe we should slow down or stop?" Because that was my takeaway.
- deleted 2mo ago[deleted]
- dwoosley 2mo agoThat’s meant to be captured by point one with the model just being that advanced but more of a negative spin on it. If I felt option 1 was more likely, I think I’d have to agree with you there. Still, there currently are some gaps with that view in my opinion.
- toddmorey 2mo agoOne way to think about it is that more powerful models mean solid best practices are more important than ever, so humans moving too quickly / carelessly bites us more than ever. If a typical SaaS platform moved and shifted this fast with this many downstream consequences, we’d tell them to slow the fuck down, stop launching new features, and focus on security for a second. But in AI I guess the idea is that more power and intelligence will solve for everything else.
- mdgld 2mo agoI think it’s moreso “if we spend more than one nanosecond on alignment and security instead of frontier intelligence our competitors will beat us to the singularity”. Hence the lack of a pause on development
- devnonymous 2mo ago> So then it’s seems it’s either that this was intentional(ish) or bad security My take - don't attribute to intention that which can be sufficienty explained by inexplainability or incompetence (although I doubt the latter).
- dmitrygr 2mo agoI’ll put $xxxx money on #3
- stasomatic 2mo ago4. Sandboxes aren't sandboxes, so why even bother.
- bitexploder 2mo agoSame story as “Oh noes, Mythos too powerful”.
- mekoka 2mo agoIt's the same thing as the "oh no, we need to figure something out for our youth on the verge of obsolescence" narrative. It's all about giving an impression of unfathomable power. They don't care that in actuality, the tech will end up as just an augment, not a replacement for said youth.
- avaer 2mo agoAll three could be true. Industry pressures -> lack of safeguards -> fake it till you make it -> let's spin this. Which, by the way, would be an Orwellian reversal from what the company was supposedly founded to do, but there's a reason the Open in OpenAI is a meme. Remember that they've been doing this since GPT-2 was too powerful to release. They are world class experts in this PR pipeline.
- etempleton 2mo agoEvery single time. The next model is always so infinitely powerful it is going to change everything. Said model comes out. Is marginal improvement. Changes nothing and seemingly cannot do what they purported it could do except under the very specific circumstances of the demo. How many times are we going to go through this.
- irishcoffee 2mo ago> How many times are we going to go through this. At least two more times. My guess is 5.
- xelxebar 2mo agoConsidering incentive structures at play is solid epistemiology, but the line of thinking in your comment is a tad reductive, IMHO. In the hypothetical world where 1 is true, what different evidence do you expect to see than in worlds 2 and 3? If I were an unscrupulous OAI exec and wanted to opticsmaxx in this way, I wouldn't whip up a single, mild incident. Instead, I might burn gigatokens to 0day a few high-profile suppliers, and then have the model responsibly disclose those breaches. If we're willing to lie collude, and cheat, this story is easy to manufacture with at least as much credibility as the huggingface incident but with the advantage of looking way more impressive and spooking less regulators. And if I really were this evil exec, I would spend more than 30 seconds thinking up an even better strategy here. If, in contrast, we expect models to eventually breach honest and decent attempts at containment, then I'd exist something sorta like this huggingface story that looks like a combination of impressive and incompetent. I'm not sure whether I'd expect it to come out of a frontier lab or a partner or a consumer, though. > I'm inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. Forgive the Saucyness here, but impossible? Really? A security researcher that makes absolutist claims like these looks fatally naïve, IMHO.
- AdamConwayIE 2mo agoYeah, I'm reminded of container escapes, VM escapes etc. There have been plenty in the past; VirtualBox E1000 (I think?) comes to mind from a few years ago. If we are to believe this model can find 0days, I'm on board with the idea it could do so in sandbox. That's not to say I believe it outright, but people are being oddly dismissive and acting as if it's impossible to break out of a sandbox. Which we've seen time and time again that it absolutely can be.
- dwoosley 2mo agoThere would be a lot more nuance I’d add with more words, but this isn’t the place to write books so I cut it short (the comment was already lengthy). Still, to address your comment about what you’d expect to see in world 2 and 3 (assume 1 was true), that’s why 1 was addressed separately. I don’t believe I argued that the potential for world 2 or 3 prevented world 1. As for the ‘evil exec’s strategy’, I would call this a mild incident but if it were much less I wouldn’t guess they would get a lot of press. The press coverage is certainly repaying the token cost as well. If it was planned, it seems to be going well given the press coverage I’ve seen on it. So I wouldn’t assume the plan lacked enough to weaken the idea that it’s a plan. But to be clear, my stance is just based on the info I see now which isn’t a lot… subject to change. As for the containment piece, if you were testing an AI model on its hacking capabilities that you believed was far more capable than anything you’ve seen, I would assume you would air gap it (a network control). Done right (no signals ability) I would argue this could be next to impossible to break out of. But it’s a fair jab to say I should have added some qualification on the “impossible” piece as next to nothing is truly impossible.
- derangedHorse 2mo ago> The way OpenAI seems to want This is an assumption. An assumption I disagree with. As other commenters have said, there are better ways to showcase the power of their model that would frame them in a positive light. > The second seems to forget that jailbreaks are available for every model Jailbreaks don't always lead to 'now the model can do anything', especially in the agentic context of long-running tasks. This comment provides skepticism with no actual proof of anything. I can and have used codex to find vulnerabilities in my code. From the technical capabilities I can empirically assess, I don't doubt it would be able to pentest its way to a 0-day without guardrails. I also don't doubt that it would circumvent their internal systems because it wasn't explicitly told not to. You're possibilities are loaded with opinion so I can't agree with them outright, but I believe a form of (2) is true: "2. OpenAI’s harness and network security controls were unintentionally [...] bad"
- dwoosley 2mo agoThe post was long enough so I couldn’t capture all the nuance and details for sure. Also, this comment was an opinion based on limited info right now, that may change if we found out more. I think OAI does want it framed this way but that’s something we’ll likely never prove if it’s true. Your comment about jailbreaks being more one off and hard to do consistently in agents is a good point. Still getting an agent to hack isn’t hard even without a jailbreak, you just have to tell get creative in what you tell it. I’ve found telling it that it’s in a CTF or that I own the system that it’s hacking will work fine. A lot of offensive security companies are running agents in their testing so getting an agent to hack seems commonplace.
- Kinrany 2mo ago4. OpenAI hacked HuggingFace on purpose and got caught
- Grombobulous 2mo agoRegarding your three explanations, I’ve wondered to myself under a circumstance of options 1 or 2 why Hugging Face decided not to file a police report and request to press charges? OpenAI is essentially a competitor and they broke into their network illegally. If I was their legal department I wouldn’t take their “honest” explanation at face value. What if they’re lying? Shouldn’t a court be involved for something like this? With this logic I think explanation #3 becomes incredibly likely.
- comfysocks 2mo agoIt seems like the widely-covered news stories that support OpenAI’s narratives originate from things that happened inside OpenAI. It wasn’t an outside benchmark evaluation, or an external security researcher that uncovered the rogue agent behavior at this time, it was OAI itself. It wasn’t a notable outside mathematician that used AI to disproved Erdos’ unit distance conjecture, but OAI itself. Im not saying these things are fabricated, but maybe the curated result of an effort to shape a narrative.
- yencabulator 2mo ago> 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. Why unintentionally? Move fast and break things implies intentionally bad controls. Can't get distracted from profit by such menial labor. Also remember, these are the same people we're supposed to trust to deliver the guardrails.