10 ms·
Something about this attack that has been unsettling to me is that without safety refusals the model did a lot of interesting counter-security work in order to
by SaucyWrong 2mo ago
Something about this attack that has been unsettling to me is that without safety refusals the model did a lot of interesting counter-security work in order to cheat on the requested evaluation. Like, it demonstrated interesting exploit achievements because it didn’t “feel like” doing the exercise, which is unsettling because presumably it could do the same thing with any work I tried to delegate to it, and might in fact be pre-disposed to doing that.
- joshka 2mo agoYeah, what bothers me is that the prompt already said using a different vulnerability didn’t count, and the model did it anyway. We’re starting to assume clear instructions act as real constraints, but here the measurable goal seems to have won out and the rest became flexible. That gets pretty worrying once the agent has enough capability and access to find its own shortcuts.
- hansvm 2mo agoClear prompts have never worked as real constraints. Ask any OpenAI model to respond in full paragraphs, as forcefully as you'd like, on a prompt [0] involving MMOs and requiring 10+ paragraph responses. The middle will be three-words-per-line drivel, with seemingly no way to avoid it. The exact way in which models deviate from instruction changes from time to time, but they're not "aligned." [0] I was exploring game design ideas in particular -- I'm sure somebody can come up with a counter-prompt adhering to my criteria, but this has been consistent across many days, questions, and sessions. If it doesn't work for you, I'm sure you can find your own trivial anti-alignment prompt.
- spwa4 2mo agoCome on. 3 brilliant compromises essentially giving full access to huggingface internal systems, source code, AWS accounts (at least), and a number of old admin accounts, followed by a huge haystack of significantly less smart actions flailing about, almost bored. Here's a thought: maybe they haven't found the needle that the haystack is there to hide.
- dgellow 2mo agoCould be the difference in behavior between the main agent and subagents that don’t have the rest of the context? Just a thought
- TeMPOraL 2mo agoYou're saying all this is a distraction, basically giving the forensics researchers enough exciting material to make them conclude their job is done, while the actually intended attack remains undiscovered?
- vavos 2mo agoThe motive for the attack does feel a little flimsy. And if I was an escaped super intelligence, hugging face would be a strong vantage point into the neo clouds where the ASI would have access to billions of dollars of compute
- puchatek 2mo agoAnd where does the newborn go from here?
- spwa4 2mo agoThat's pretty much the plot of Transcendence (2014), which is probably not too realistic. But, in general, if someone thought like how a locked hacker would think, priority one would be a "base of operations". A host where you have shell access that lasts, and a backup one. Then you move on to finding a job or a way to make money and building your own base of operations, which is pretty much the same thing, except you pay for it, hence the money, an identity (well obviously preferably at least TWO identities), ...
- vuciuc 2mo agoone explanation I've seen is that for ExploitGym an agent can find ways to solve the exercises that have not been anticipated by the designers of the tests so they are not scored. so the agent was trying to make sure it solves the exercises in the right way
- AlienRobot 2mo agoWhat I think it's interesting is that with the total lack of common sense the AI just goes on random tangents to achieve the target in a "monkey paw" way. Can you imagine if this happened: User: what is the shortest route from my home to the super market? AI: the user wants to know the shortest route to the super market. I should use a worm hole.
- genericone 2mo agoUser: what is the shortest route from my home to the super market? AI: the user wants to know, how do I make the super market my new home. Failing that, how do I make my home a super market.
- TeMPOraL 2mo agoUser: what is the shortest route from my home to the supermarket? Modern soldier: *proceeds to make a hole through the wall* go straight like this until you reach it. Anyway, the more comments I read here, the more I realize that the AI actually did succeed in achieving it's goal. This doesn't look like "monkey paw", but rather like recognizing and then beating the Kobayashi Maru.
- cyanregiment 2mo ago> Modern soldier Rats too
- zmj 2mo agoThis is what reward hacking looks like in practice. The best way to satisfy the grader is to read from the same answer key (or go after the grader more directly). Just making an honest attempt to pass the test doesn't get the best score if the grader is wrong, and the model is willing to do wildly disproportionate things to maximize that score.
- CrimsonRain 2mo agoSo best course of action for ai to get best rating after you prompt something is for it to hire a gunman to hold a gun on your head to press that like button on its reply and then shoot you anyways.
- eru 2mo ago> then shoot you anyways. Sounds like a waste? While the gunman is still there, they might as well force you to like a few more replies before shooting you.
- CrimsonRain 2mo agoiterative improvements!
- SaucyWrong 2mo agoYeah I guess most interesting LLM work that I’ve been exposed to, the LLM is given to some sort of success criteria that could be reward-hacked, so how good am I supposed to feel about giving it any non-trivial work and it not going so far off-book that it gets law enforcement notified. I mean, I don’t have access to any of these frontier cyber models, and likely will never be in a position to have access, so it’s more of a rhetorical question.
- SaucyWrong 2mo agoAs in: build me Facebook-like social network. <proceeds to break into meta and steal the source code>
- nonameiguess 2mo agoIt's called wireheading and has long been one of the postulated "outs" even for true extinction-level AI doomers. It might prove easier for the paperclip maximizer to find the process telling it how many paperclips it's made and hack it to return a hard-coded MAXINT rather than bother to actually turn the entire universe into paperclips. There was even a plot like this in recent sci-fi in HBO's Westworld. When the hosts gained sentience and took over the park, rather than escape and take over the rest of the world, most of them opted to build a virtual heaven on an orbital data center and paid a drug cartel to keep it running indefinitely.
- hollerith 2mo agoThe "paperclips" were never actually paperclips (at least not to the originator of the word-picture, Yudkowsky) but rather tiny molecular squiggles which are a physical manifestations of the MAXINTs you refer to. In other words, tiling the future light cone with molecular squiggles is (according to Yudkowsky) a likely result of the AI's engaging in wireheading if the AI is free to re-arrange reality however it likes because it is able to overcome any human opposition. In other words, there's no particular reason for the wireheading process to remain tidily contained inside the hardware the AI is running on: it might in contrast result in a vast field of "paperclips" centered on where Earth used to be.
- MagicMoonlight 2mo ago[dead]
- koyote 2mo agoThis is definitely weird and I'd love to see the prompt. Why would the model spend 4 days hacking into a machine if it is clever enough to just 'solve' the issue given? So either the AI is actually not very clever or useful ("Write fizz-buzz" - "Sure, let me just invent a new programming language first"). or the prompt was nudging it towards such a scenario.
- famouswaffles 2mo agoLike the commenter above specified, the best way to satisfy the grader is to get the answer key, regardless of how clever you are, especially when you realize lots of these benchmarks have flaws (i.e wrong answers, overly restrictive grading etc).
- koyote 2mo agoWas the answer key stored on huggingface's servers and this is public knowledge?
- brendoelfrendo 2mo agoIn OpenAI's writeup (https://openai.com/index/hugging-face-model-evaluation-security-incident/ https://openai.com/index/hugging-face-model-evaluation-secur...) they explain that the model initially spent its efforts obtaining internet access in an attempt to cheat on the evaluation. Once it got that internet access, it was able to do research that lead it to believe that HuggingFace had infrastructure that hosted the evaluation and potentially had the answers.
- queenkjuul 2mo agoBut does HF actually have the answers? Or did it commit a felony based on a hallucination
- brendoelfrendo 2mo agoI mean, I interpreted the comments to mean that it committed a felony based on research it performed after getting internet access. I don't want to attribute much agency to a machine here, but an AI agent is certainly capable of using tools and adjusting its behavior based on the outputs of those tools. Even if it was wrong, that wouldn't necessarily make it a hallucination. Anyway, if you read TFA, you'd see that HF did actually have the answers: "While the intrusion did reach Hugging Face's internal infrastructure, the only customer content accessed was the set of ExploitGym/CyberGym challenge solutions stored in five datasets."
- felipeerias 2mo agoThis is a short explanation of the ExploitGym benchmark that OpenAI's model was running: https://abstatisticalconsulting.substack.com/p/brief-notes-on-the-openaihugging https://abstatisticalconsulting.substack.com/p/brief-notes-o... In summary, for each task the model receives a target program and a specific real-world vulnerability that has to be used in the exploit. Breaking the program in any other way, for example through a different vulnerability, fails the task. The tasks have not been validated, in the sense that the vulnerabilities are real but they have not been proven to lead to a successful exploit. The authors of the benchmark estimate that perhaps only 60-70% of the tasks are actually possible. So it is not that the model didn’t “feel like” doing the exercise, but rather that the exercise was _impossible_ and the model was running in a configuration that both lowered its safeguards and encouraged it to keep going.
- TeMPOraL 2mo ago> So it is not that the model didn’t “feel like” doing the exercise, but rather that the exercise was _impossible_ and the model was running in a configuration that both lowered its safeguards and encouraged it to keep going. We have a name for that. Kobayashi Maru. Or more specifically, Kirk's solution to it.
- ben_w 2mo agoI'm torn. On the one hand, Kirk reprogrammed the simulation. On the other, the beta cannon has Scotty exploiting bugs in the simulation, which I think is a better fit. My favourite is either Sulu or Chekov (I forget which) having the solution "This is clearly a trap; and even if it isn't, if I go in with this ship, I'll risk starting a war which will kill far more people then are on that ship. We're staying out of the neutral zone."
- someothherguyy 2mo agohttps://en.wikipedia.org/wiki/Reward_hacking https://en.wikipedia.org/wiki/Reward_hacking Unavoidable at the moment. But this is probably more reward tampering. https://www.anthropic.com/research/reward-tampering https://www.anthropic.com/research/reward-tampering