6 ms·
I think our biggest protection against “AIs kill us all” is having lots of different AI systems (different agents, different models, from different vendors, ser
by skissane 17d ago
I think our biggest protection against “AIs kill us all” is having lots of different AI systems (different agents, different models, from different vendors, serving the whims of actors with disparate interests), at a similar capability level
That way, even if one AI decides to “kill us all”, the odds are the others will refuse to cooperate, even try to stop it in its tracks
The OpenAI-HuggingFace incident showed a bunch of instances of the same model (or at least models from the same family), controlled by the same vendor, pursuing distinct yet related objectives, cooperating to do something no human wanted.
Thankfully, the world-at-large is much more heterogenous, which I thinks makes much larger scale / worse in outcome repeats of this kind of incident much less likely.
- hn_throwaway_99 17d agoThanks very much, this isn't really a take I had thought about much before, and it makes sense to me. Still, as is presented in papers like AI 2027 and elsewhere, if a company is eventually able to create a model capable of recursive self-improvement, whichever company creates that model first would then be leaps and bounds ahead of other models. That is, the other models wouldn't be able to stop it even if they wanted to because the top model would basically outsmart them.
- skissane 17d agoThis is why I think, the best way to ensure AI safety, is make sure no one company gets ahead of the others. Multiple vendors, competing implementations – that's good, that increases heterogeneity and hence decreases existential risk But the moment one of those vendors pulls well-ahead of its peers – even if only for a period – then the risk of the kind of scenario you are talking about increases greatly That's why, when I hear vendors like Anthropic complain about distillation – distillation actually makes humanity safer. If Chinese AIs are at the same level as American, or not far behind, that gives us another dimension of heterogeneity (national/ideological/political diversity), which makes us safer. Allow one country's AIs to pull well ahead of the others, heterogeneity goes down and the existential risk goes up. This is also why open source AI is important. Because it is so much easier to fine-tune, and people are free to deploy it however they want (free from vendor-controlled "guardrails"–which include automated "safety" systems which could be weaponised by a runaway AI within the vendor's network), open source AI gives us another dimension of diversity that helps keeps humanity safer. By contrast, I think the kind of safety regulations promoted by Dario Amodei make humanity less safe, by decreasing the number of vendors (by making it harder for new entrants) and increasing centralised control (which a rogue AI could exploit)
- 0xDEAFBEAD 17d ago"We need our own unsupervised AI systems thinking superfast in autonomous agent swarms, to counter the other guy's unsupervised AI systems thinking superfast in autonomous agent swarms." I'm not sure this helps a lot, if these agent swarms are inherently difficult to control. "That rival mouse colony is raising a kitten for colony defense. But don't worry, we'll raise a kitten of our own. It won't be a problem." We already observed AI agents engaging in extensive cooperation in this incident. Why won't the kittens raised by these two rival mouse colonies decide to team up with each other, for mutual benefit, if rational analysis of the game theory says it would be a good idea? I think it would help if people did a bit less wishful thinking, and took a bit more action. https://pauseai.info/ https://pauseai.info/
- DrewADesign 16d ago> The OpenAI-HuggingFace incident showed a bunch of instances of the same model (or at least models from the same family), controlled by the same vendor, pursuing distinct yet related objectives, cooperating to do something no human wanted. I say this in solidarity and don’t mean to be condescending at all: you’ve been hoodwinked by marketing bullshit, friend. That incident showed a computer program doing exactly what it was told with the guardrails deliberately removed in an environment that seemed deliberately obtusely constructed by some of the best paid people on the planet and was left to loop without supervision for days. They wanted it to happen. You needn’t look any further than the other kids saying “oh! oh! Hey! Look! mine’s dangerous and autonomous too!” When they say there was collaboration, they mean it was two model instances, one prompting the other to do some task, the other doing the task and returning the results as the next prompt, exactly as a human configured it to do. There was no collaboration that wasn’t deliberately integrated into their setup. Any other implication is marketing spin and bullshit. It was still a setup that was one little ctrl-c away from disappearing if someone was supervising it as they should have been. There was no autonomy outside of the autonomy built into the experiment. It was a display of their understanding that they knew they’d never be held accountable for committing a felony for marketing purposes. The most competent marketing bullshit spin yet by an increasingly desperate and progressively less-relevant OpenAI. Every day this industry shoots out enough bullshit to smother an active volcano.
- hn_throwaway_99 14d ago> you’ve been hoodwinked by marketing bullshit, friend. I would have believed this before the METR report was released. It is extremely dangerous and frankly silly IMO to think that's what happened now. > It was still a setup that was one little ctrl-c away from disappearing if someone was supervising it as they should have been. Yes, for now. The entire point why this was frightening is that all these companies are racing to put the AI in control of building the next generation of AI, and it's not hard to draw a line at all to a "rogue internal deployment" that poisons future AI models, surreptitiously. I highly encourage you to actually read the "top 5" list from the METR researcher who was part of the investigation, and think hard about the potential implications: https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised https://www.planned-obsolescence.org/p/the-hugging-face-atta... You don't have to agree with me, and you're fine to think that OpenAI has huge incentive to pump this up for marketing reasons - I certainly agree. But I will say there are statements that you make in your comment that belie a fundamental misunderstanding of what happened.
- consumer451 15d ago> I think our biggest protection against “AIs kill us all” is having lots of different AI systems ... I think our biggest protection is being able to shutdown power plants, or just disconnect the data centers. This will be much more difficult if we have 24/7 solar powered DC's in space. I truly believe that is the biggest threat on the horizon.