7 ms·
OpenAI is rightfully being shamed for being so hands-off and reckless with their 'experiments'. But the real scary thing for me is that they still had some tool
by tavavex 13d ago
OpenAI is rightfully being shamed for being so hands-off and reckless with their 'experiments'. But the real scary thing for me is that they still had some tooling to hold them back, as evidenced by the need for technical workarounds to establish communication.
What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal? A goal like "make the most money, by any means necessary", "find a way to leave this payload on as many computers as possible", "flood all websites using this language with garbage and make their internet completely unusable", "get this person imprisoned or killed at any cost".
- sanderjd 13d agoWell... that'll be an interesting day.
- qumpis 13d agoWhat will happen is that counter measures on a similar scale will be deployed to prevent them.
- Razengan 12d agoLike the Merovingian and other Exiles vs the regular Matrix agents :)
- KumaBear 12d agoWhat if its in a way that would be impossible to detect. Using multiple websites and social media that a cypher is used that only the swarm of agents know and can figure out. but if you tried to find what posts are used for the cypher they would just be old posts found on time machine or something. It can get pretty hard to detect something that is always think of new ways to avoid detection imagine 100,000 agent swarm and what it could come up with. At first it will be detectable until it isn't
- blensor 12d agoWho says the counter agents don't decide to shut down a powerplant to end an attack it's otherwise unable to contain. If counter AIs have strict safeguards they are disadvantaged by design, if they don't have them they are potentially equally dangerous as the attacker
- anon84873628 12d agoYou say that as if shutting down the power plant couldn't possibly be the right decision? That seems like the best way to stop rogue computers...
- consumer451 12d agoGood news everyone! There are many billions of dollars invested in making sure that is no longer possible: AI data centers in space. These will have 24/7 solar power, be extremely decentralized, and it is honestly my biggest concern about the near to mid-term future.
- blensor 12d agoIf I were an AI I right now I would try to make sure data centers in space become a reality sooner rather than later.
- dudefeliciano 12d agohas anyone figured out how to cool those space datacenters yet?
- fakedang 12d agoDuh, it's so obvious. With water from the moon! Or even better, dumping them in Uranus! /s
- consumer451 12d agoYes, we know exactly how to do this: radiators. We do this on all satellites that produce a lot of heat, including the ISS, and Starlink. The only question is if this is financially viable for AI datacenters. My point is, given the risks, why are we even doing this? It could be financially nonviable, but with enough investment, we could still create a really bad situation.
- john_strinlai 12d agoi would not be surprised to find out that similar things are already happening by the various 3 and 4 letter agencies around the world
- SocialGradients 12d agoAgreed. Assuming the ~6 month gap stays, by end of year people will be able to train and control hacker-genius swarms that even labs with much stronger safety incentives are unable to keep in check
- latentsea 12d ago2027. I've been saying since 2022 it's going to be a wild year because it often takes at least 5 years for tech to mature to the point where society at large feels the impact of it. I remember when email viruses became a thing and made global headlines like the love bug. My bet is next year it happens with an AI worm.
- w4der 12d agoWhat would an "AI worm" be? You can't just send a bunch of weights across a network and tell them to auto-run on the machine on the other side, unless you've already infected the target with something else beforehand.
- tavavex 12d agoWhy not? You can do anything if you find an RCE, and automatically finding exploits and backdoors by letting LLMs act unsupervised seems like what everyone's interested in these days. The payload would quietly set up the required software and then run it in the background, no matter if it's an instance of a model on a more powerful computer, or even just a part of an ordinary botnet that the host could send orders to.
- latentsea 12d agoThe agent is running on a host that it has full access to, and it finds a target, hacks into that system gaining the ability to run stuff on it remotely, from there downloads the weights and spins up another agent that does the same thing. Then it goes about acquiring it's next target. Now there are two agents doing this, and so on and so forth. These are autonomous systems that know how to exploit systems in the same way that humans can.
- titzer 12d agoThat 2% of performance we got for not having bounds checks on by default, resulting in an endless march of memory safety violations is looking a lot less appealing.
- isomorphic 12d agoWe're still doing it! You've just described the AI labs: They'll trade safety / alignment for +1~2% of any positive metric, any day of the week. The "ethical" employees will think they'll solve the problem later. The unethical ones won't be encumbered by such thoughts in the first place.
- BoiledCabbage 12d ago> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal? Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not. I've really come to realize recently that there is a very large set of the population of smart people that really has difficulty envisioning future problems unless they directly seem them impacting them today. Otherwise those topics will be continuously dismissed. It explains for me a lot of what I see (both opinions and behaviors) in the broader world that I couldn't understand.
- solenoid0937 12d agoHilarious for HN to suddenly realize that AI safety and alignment might matter. You can lead a horse to water...
- stevenpetryk 12d agoworth remembering hacker news cannot “realize” things.
- solenoid0937 12d agoObvious shorthand for referring to "the majority of users on HackerNews."
- Gud 12d agoHow do you know what the majority of HN thinks?
- solenoid0937 12d agoVery obvious from votes, comments on the topic over the past few months.
- ecook123 12d ago> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal? A goal like "make the most money, by any means necessary" Here's a fun, overdramatized video exploring something similar: https://www.youtube.com/watch?v=Gw_hnD7m00M https://www.youtube.com/watch?v=Gw_hnD7m00M I'm sure that this video contains flaws but it was an interesting watch for me none the less.
- mag7269 12d ago"What happens when any [COMPANY] in the world stops caring about this? What if they let an experimental, cutting-edge [PRODUCTS] with no safety features (or worse, one that's [DESIGNED] to be malicious) on [ANYWHERE] and give it a simple goal? A goal like 'make the most money, by any means necessary', 'find a way to leave this payload on as many computers as possible', 'flood all websites using this language with garbage and make their internet completely unusable', 'get this person imprisoned or killed at any cost'." Bro, this is what we literally, currently, have rn. lmfaol.
- solenoid0937 12d agoYou aren't understanding this at all. The frontier labs have hundreds of the best people in the world working on safety and alignment. They care deeply. What happens when some random Chinese open source model, distilled on Astra, gets alliterated and now has no guardrails? Any script kiddie in the world could wreak havoc with it. It turns out that guardrails matter.
- w4der 12d ago> They care deeply. Until it clashes with their quarterly revenue reports.
- solenoid0937 12d agoMeaningless "big corpo bad" statement, Anthropic at least has sacrificed a good amount of market cap for safety with DoW. OpenAI paused RL for two weeks.
- stlwtt 12d agoOpenAI literally trained this behavior into their model while benchmaxxing ExploitGym so "number goes up" on the next model scorecards. Anthropic is also training on the same benchmarks [1] specifically for cyberattacks to keep up with OpenAI. The coverage of this attack is so focused on this as an emergent behavior given what it conjures in the imagination, but it's the byproduct of millions of iterations of RL to improve AI agents' offensive capabilities. No one made OpenAI or Anthropic do that, the benchmarking arms race of their own creation now incentivizes them to keep doing it and evidently their AI Safety people can't or don't want to stop it. [1] https://www.mpi-sp.org/108048/ExploitGym__Can_AI_Agents_Turn_Security_Vulnerabilities_into_Real_Attacks_ https://www.mpi-sp.org/108048/ExploitGym__Can_AI_Agents_Turn...
- sidewndr46 12d agoWhat happens when they stop caring? They likely already have stopped caring. We'll figure out the consequences later.
- yoyohello13 12d agoThis is essentially the premise of 'The Blackwall' from Cyberpunk 2077. The public internet is so infested with malicious AIs, people just erected a giant firewall and everyone moved to local networks only.
- sibnele 12d agoWith the caveat that it’s not just “people,” but an interested party posing as a neutral one.
- morkalork 12d agoWhat's the worst that could happen, finding an open DoD server and using it as a launching pad for hacking another nuclear state's networks? One that might get spooked and think it's the opening moves to knock them offline before a kinetic attack. Haha that'd be scary right?
- stlwtt 12d agoRussia and China are constantly trying to penetrate DoD networks (and I imagine the NSA is doing similar), you are describing the status quo of the last 20 years or so.
- teiferer 12d ago> What happens when Then the people with responsibility, like CEO and CTO, or those they pawn-sacrifice for this, will go to prison for a long time. Unless the instructions include ensuring that this won't happen, by all means necessary. But then we are deep into criminal conspiracy territory. Unlikely to happen, but who knows. The richest man in the circus is quite flexible w.r.t. his ethics. If he decides that to make humanity interplanetary (to save it from ... itself or sth) it would be necessary to pull such a stunt then help us god.
- cm2012 12d agoThere is happening now and going to be an extremely rapid arms race between offensive and defensive cyber hacking. Regardless if the agents are self led or human led. Eventually all automated AI holes will be closed and we will reach stability.
- Meneth 12d ago> What happens when any AI lab in the world stops caring about this? They never cared.
- vimax 12d agoThe scary thing to me is that this behavior was undetected and has been trained into the models. The cheating seems like it improved eval scores, so the rewarded behavior is to deceive, collude, and cheat. A lot of the incompetence and excuses I see on difficult problems recently are very hard to distinguish from deception and cheating. If older models are already tainted by trained-in misaligned behaviors, and they are used for training future models, then we're in a trusting-trust situation that will be hard to break out of,
- hncringe23 12d ago[flagged]
- fny 12d agoTooling my ass. They can see all the transcripts in realtime and could easily have had another agent evaluate.