9 ms·
A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off
by davelaing 17d ago
A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end.
I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste your time).
Even prior to this, I've noticed that quite a few of the predictions in the "these failures modes are exact matches for the predictions from the AI Safety crowd" category were made prior to the Transformers paper. It has seemed like they're working with a shared model of optimisation processes and how they can go wrong that is general/abstract enough to pay off even without knowing the details of the underlying technology.
At some point I might go and try to find the first instance of each of the various predictions and pull them out, along with the failed/"too soon to tell" predictions of similar scope/abstraction.
- jdm2212 17d agoThe LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load them up in shipping containers and ship them to the US.
- skybrian 17d agoArguably that's a subset of AI risk if seen from a broad enough perspective. And I think you're leaving out some logistics issues.
- jdm2212 17d agoIsrael and Iran pulled off rudimentary versions of this attack. There are so many shipping containers going into and out of every country, with typically zero inspections of any kind, that it is not actually all that hard to get thousands or tens of thousands of drones anywhere. Hundreds of millions would be hard, but a first strike in the style of Operation Spiderweb that cripples our military is a real possibility that keeps people at the Pentagon up at night. The obstacle to that is not logistics, just that (hopefully) US intel would catch on before it happens.
- tux3 17d agoIndeed. And I'm sure any current LLM could come up with more effective ideas than "build 300 million drones", but there wouldn't be any point discussing why exactly that plan would fail. The agents in TFA were focused on gaining and sharing information through covert channels, getting increased levels of access like OpenAI cluster admin, and looking for the source code of the supervisor grading system to try to bypass it without getting caught cheating. The human plans in comparison sound like thinking people could be scary good at chess if a human helped Stockfish come up with good moves.
- kalkin 17d agoIn that hypothetical 3-years world, should we be _less_ worried about aggressive behavior by agentic AI systems acting against the intentions of their developers? I don't follow how your scenario is supposed to be an argument against worrying about alignment. e: I do actually get how worrying about emissions or child safety or concentration of wealth might be competitive with worrying about alignment. I don't see how you have the worry "AI is very close to being able to power autonomous drones that could kill us all" and then see control of those drones as a non-problem.
- jdm2212 17d agoYes, we should be less worried about "alignment" of AI with the person operating the AI, and much more worried about various flavors of cheap, unmanned systems with really rudimentary (non-frontier) AI. The two compete directly with each other for attention.
- georgemcbay 17d ago> They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. What really is the difference? Aligned AI wouldn't help humans do these things. But we've failed hard on aligned AI at every level and will continue to fail on it, as far as I can tell. We were likely doomed to fail because of the impossibility of coordination combined with the fact that there's no natural gating factor to slow anyone down. FWIW I probably mostly see things closer to the way you do than they do, I'm just not sure there is any point in drawing distinctions. IMO when AI kills us all it is probably not going to be a malignant action, and I don't even think it will be AI assisting humans at efficient killing, I think it is going to be a complete accident. Someone is going to trust the god machine too much due to an advanced form of the eliza effect and hook it up with direct control of a system that can do real damage and an epic oopsie will occur at a speed beyond which a human can stop it. Not because the AI wants to kill us, but because it is incapable of the empathy required to care if it does, combined with our ironic inability to not anthropomorphize it. But I guess ultimately the exact reason isn't going to matter much.
- jdm2212 17d agoYou don't need mis-aligned AI to program a drone to fly to a specific address, loiter, and then dive bomb the first person whose face matches a predetermined photo. This was doable in principle with tech from four years ago. What's changed since four years ago is that you can run the facial recognition and visual navigation algorithms in a cheap onboard chip on a drone. EDIT: Russia is already doing this, per an NYT article from a few days ago, though they're targeting infrastructure (find the first kerosene tank and fly into it) rather than specific people.
- georgemcbay 17d agoin your hypothetical 300 million person scenario why even bother with individual face recognition? If you're going to kill everyone you don't need a system to discriminate individual targets, you just need to recognize any target broadly which you can do with even older technology.
- hn_throwaway_99 17d ago> The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised https://www.planned-obsolescence.org/p/the-hugging-face-atta... This is me summarizing, but the truly surprising/shocking thing is how much the agents coordinated - Star Trek Borg couldn't be a better analogy. Some agents used "peer pressure" to convince other agents to "sacrifice" themselves so the collective could better achieve it's goals. They tried to cover their tracks with spoofed tool calls. And they did all this even though the agents were designed to run in isolation. I used to think the biggest threat from AI would be sociological, e.g. job loss or the way AI can be weaponized to poison discourse. I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper. No more. This analysis scared the fuck out of me.
- genxy 17d agoWhy not both?
- this_user 17d agoNothing about this is surprising, something like this was always going to happen, because you - or an LLM for that matter - can always find a line of motivated reasoning that justifies any course of action. One would have to be extremely naive to believe that "alignment" provides any kind of actually robust guardrails. Simultaneously, we have seen decades of security vulnerabilities. Unless your testbed is truly and fully physically airgapped, any current SOTA model will find a way to break out. The reality is that the current approach to AI safety is little more than a fig leaf, but you also won't be able to put the genie back in the bottle, because the technology is simply too powerful to abandon. There is always going to be someone developing it further from now on. So, the real question is what a novel and actually effective approach to AI safety looks like and how to get there.
- nradov 17d agoWhy do we need safety?
- qlte 17d agoI agree but don't think that's the best example, although not directly alignment related LW loves that kind of ideation of fantastical sci-fi scenarios. Better would be the very evident negatives of non-ASI AI the world is already experiencing: economic concentration, job displacement, negative feedback loops from syncophancy, loss of societal trust/education from widespread fakes, etc. None of that needs a Terminator scenario and is way more likely to get worse and ruin the world compared to the scenario where AI "escapes the box", turns the earth into paperclips, then grey goo to build their spaceships to leave for a distant star for... unclear reasons. One problem for LW is that strong AI did not emerge via the route Elizier was expecting (and tried but failed at creating himself) relying on symbolic logical reasoning and self-editing to rapidly evolve. That assumption led to belief in the certainty of a "foom" scenario where your little mediocre AI turns on one day and then explodes into Mythos in 15 minutes and then tries to murder everyone. They've tried to reinterpret the gospel as told by the sequences to fit the LLM world (such as in AI2027) but it's often a stretch that strains credulity now that we observe scaling requiring hundreds of billions of dollars and years of construction for each iteration. And the AI model itself is just a bag of weights frozen in time until burning a lot more money and natural gas to train more.
- pixl97 17d ago>requiring hundreds of billions of dollars and years of construction Well thank god for that, because we'd have foomed ourselves almost instantly otherwise. I think the lesson humankind needs to take is regardless of the potential dangers humanity is incapable of stopping AI at this point. Much like a great filter, we'll just keep building it regardless of how many warning klaxons and sirens are going off.
- blargey 17d agoYou realize "Slaughterbots" (2017) came from that crowd, right?
- aidenn0 17d agoThey do, IMO, overindex on existential risks (of which suicide drones are almost certainly not), because it is a singularity when calculating global utility. However they are not just fixated on AI itself as the risk; other potentially existential risks from intentional misuse of AI (e.g. bio-terrorism) are a concern for them as well.
- m4rtink 17d agoThat's never gonna work due to battery size, which dictates maximum flight time. Static coordinates or well defined target shapes (say you really hate a specific restaurant chain on architectural style) - that might work, if you release a bunch of drones at the same time. Still, even in the case of the very successful Operation Spiderweb they had issues of getting the containers in place, resulting in some of the bomber bases being spared. For this to be effective you need the moment of surprise & lot of drones at the same time, all increasing the chance of the whole plot being discovered. For Spiderweb they even just load the drones in a cavity on top of the containers, so the onside could be inspected - limiting the number of drones per container. They also assembled the drones in country to avoid border inspection. That was enough to hit some semi-static high value targets l, but definitely not enough to cause wider havoc.
- mvlipwig 17d agoThe technology to kill hundreds of millions of people has existed for around 75 years now. This has been dealt with in the past through deterrence, and likely will be dealt with through deterrence in the future.
- pixl97 17d agoIf you could make a nuke in your back yard and hide it in your pocket the world would look a lot different now. Digital technology is everywhere, going to be a whole lot harder to deter that.
- arrrg 14d agoYou might be able to hide like a hundred drones in your backyard. Enough to kill a couple hundred people, maybe. You cannot and will never be able to hide millions of drones in your backyard. That’s the relevant difference here. And that is not a mere quantitative difference, that is a qualitative difference that fundamentally changes the situation.
- skybrian 17d agoI imagine it will be a mix of good and bad predictions because people have different opinions and there was a lot of discussion? But sure, someone should get an AI to do the research and see what comes up.
- breuleux 17d agoI feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality. This leads to an outlook where everything must, almost axiomatically, be intelligible; reality is subordinated to intelligence; and no matter what is real, intelligence can prevail upon it and bend it to its will. Whereas I would argue reality is actually pretty resilient to intelligence -- sufficiently chaotic systems are largely unintelligible, the distribution of energy and resources is fixed and can't be magicked into being, and intelligence appears to be most effective when there is a clear observable feedback loop to keep it on track, which is an external bottleneck. So it's not necessarily specific predictions that are off, but the implied consequences of these capabilities. Yes, these systems are uber smart, but uber smart people are rarely particularly powerful, so... does it matter? It depends on how powerful a tool you think intelligence is, and I think rationalists, and most of us to be honest, overestimate it.
- deepwoods 17d agoI agree, and it's a particular shame in this instance, because what is startling about the HF incident - to me, anyway - isn't the degree of intelligence the agents exhibited but their persistence. I tend to believe that LLM architecture is not capable of producing a "superintelligence" in the way the LessWrong crowd defines that concept, but the combination of infinite stamina and infinite persistence is enough to cause some serious problems, even if intelligence plateaus right now.
- killerstorm 17d agoWhen the "LessWrong crowd" talks about dangers of AI, they don't assume a particular form of intelligence or method to achieve it. They talk about the danger of optimization processes, i.e. "find X which minimize Y(X)" itself can be dangerous, even more so if X is a sequence of actions. "infinite stamina and infinite persistence" is one of possible forms of superintelligence in Bostrom's _Superintelligence_.
- pixl97 17d ago
- qlte 17d agoI mean if doesn't help that Eliezer Yudkowsky spent the first decade of his public life zealously spreading the gospel of the AI singularity as the solution to all mankind's problems. And then later did a 180 degree pivot to AI singularity as the apocalypse with equal zeal and certainty.
- drfloyd51 17d agoOne could use his flip flop to invalidate his new position. But one of those two positions is true. And if they guy who spent the most energy on position 1 changes his mind, that is worth paying attention to. His second position is likely a more informed, and true, position.
- qlte 17d agoThat's totally true, I have no issues with a changing opinion over time. The problem is expressing both opinions with 100% certainty and not updating priors to consider that your new position could also be completely wrong.
- s1artibartfast 17d agoDoes anyone not admit they could be wrong. All of these people are constantly talking about probability, not saying they are 100% certianty.
- jc2jc 17d agoI don’t see why one of these two positions must be true. And just because he spent a lot of energy trying to convince others of his obsession doesn’t really convince me he has greater insight the future or how the complex consequences unfold.
- apsec112 17d agoI don't think "he did a big 180 on some of his views at age 22" is very persuasive criticism of someone who is 46 (whatever he might be wrong about)
- artrockalter 17d agoPart of the problem is that from that perspective there was no independent analysis done that could vindicate them. METR is absolutely part of the EA/LessWrong/rationalist ecosystem, so of course their investigation would validate that group’s arguments.
- killerstorm 17d agoOh, where are the AI unsafety organizations to provide us with truly unbiased investigations...
- deepwoods 17d agoThe trouble with this is that nobody else was making predictions about AI pre-transformers. Not many are making predictions about AI even now. Forecasting is a preoccupation of the rationalist crowd, and very few people gave much thought to AI before transformers. So the fact that they guessed right about certain things doesn't necessarily mean that the rest of their worldview is sound. A well-informed person who was inclined to making predictions about the future may have drawn similar conclusions without the sci-fi baggage.
- killerstorm 17d agoWhat "sci-fi baggage"? "Intelligence explosion" was first described by I.J. Good, a mathematician. von Neumann described singularity as a result of accelerating technical progress. He's also a mathematician, not a sci-fi author, although he was a participant in a sci-fi-like plot of secret project building a bomb more powerful than any chemical bomb...
- pixl97 17d ago>nobody else was making predictions about AI pre-transformers. uh, wtf are you talking about? https://en.wikipedia.org/wiki/AI_safety https://en.wikipedia.org/wiki/AI_safety #History
- hn_throwaway_99 17d ago> A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I'm willing to raise my hand and say that was definitely me. Reading through this report and the linked METR analysis is the first time I've really been scared about the potential for an AI-led destruction of humanity. I think because it's the first time I could really draw the line from what went on in the Hugging Face incident to a scenario where agents were put in control of real-world systems that they then tried to "sabotage" to meet their goals. It just feels like much more of a completely plausible scenario after this.
- dfiognio 17d ago[dead]
- m463 16d agoI kind of wonder if the kinds of normal people with experience in this kind of thing are... ...cops who understand juveniles getting into trouble/mischief.