11 ms·
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
- Willish42 15d agoMore than other AI moments in the last few years, this feels to me like an event in tech that will be seen retroactively as an important watershed moment. I really appreciated the negative framing and critical tone of this author's overview. Having also read through the METR report, I fee like frontier labs' pattern of getting PR about how impressive their models are "behind the scenes" has poisoned their ability to take security and reliability seriously. Sure these are new failure modes and the agents operate at a scale that's difficult to combat, but the lack of controls and concern for mitigating these sorts of hacks in the future is crazy to me. The model for postmortems I was taught which has served me well in my career is thoroughly answering the following: - what happened / what was the timeline of events? - how was it mitigated and ultimately resolved? - what went well? - what went wrong? - where did we get "lucky" (meaning it could've gone worse but some arbitrary details about the incident worked out in our favor. usually stuff like "happened during business hours" or "we were already looking at a related thing that brought this to our attention before it was a worse outage") - (action items) how do we detect, mitigate, and prevent this type of failure in the future? I really hope OpenAI has done an internal postmortem that answers these questions thoroughly. Most SWEs in the industry have to do such postmortems for much smaller outages with way less impact and risk of societal harm. This is probably another area where regulation and governmental oversight would help curb the risks. What's to stop OpenAI and other frontier labs from an intentional "accidental" attack that results in gaining access to competitors' systems? I'm also curious what, if anything, Anthropic and Google have done differently to prevent a similar event. I suspect they actually monitored the agents as part of their studies and had better guardrails in their infra for how they set up their harnesses etc. for testing models, particularly when the other guardrails are absent as was the case here.
- OgsyedIE 18d ago>Spontaneously deciding to find targets to phish, >phishing them, >building armies of fake (sockpuppet) open source contributor personas, >using them to push updates to various things that inject prompts into other bots so the other bots join in on the phishing campaigns . It's a very simple strategy, executed with patience and single-mindedness.
- BoiledCabbage 18d agoIncredible
- antonvs 18d ago> There was a distinct lack of self-reflection It’s not their fault, they’re lawnmowers. And these are the people we’re entrusting to work on “alignment”. It’s difficult for them to do that when they’re not aligned themselves.
- trollbridge 18d ago“Why does my lawnmower keep on moving when I hop off of it after ratchet-strapping the seat and the pedal down?”
- estearum 18d agoThis would be a legitimately big problem if lawnmowers became continuously more and more valuable the more securely you ratchet-strapped their accelerators down, wouldn't it?
- AlotOfReading 18d agoI think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent reporting. I suspect the omission is actually a result of company/industry myopia to human factors analysis, but it dovetails amazingly well with the marketing narrative.
- hawkice 18d agoThis writeup emphasizes the many, profound human failures that led to this, at the time, and continuing to the present day.
- AlotOfReading 18d agoCan you point out where? Looking at the METR report, the only place I see discussion of humans being involved in the sequence of events is two short paragraphs on page 30 where a security investigation into the artifactory issues led to a pause before ExploitGym experiments were resumed. There's no deeper analysis on what was found during that investigation, nor why training was resumed even though the issues weren't mitigated. Another part discusses The agents choosing not to actively email a human researcher, but not the human researchers actively looking for evasion.
- FabHK 18d agoFrom the article: 1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous. “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.” 2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught. 3. Failures of Infrastructure. OpenAI failed at delivering even basic software security, in numerous ways. One example is the models all having access to the same Artifactory instance. There were also other severe failures. Nor did OpenAI seem to be properly testing for such failures. 4. Failures of Alignment. The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why. 5. Failures of Attribution. OpenAI’s post-mortem essentially blames events on a real and important series of prosaic failures. But solving that won’t get it done. 6. Failures of Environments and Data. Prosaic failures in the RL pipeline absolutely did contribute to this, especially impossible tasks. This is ubiquitous, all of this is always rushed, as Utah Teapot explained this week. 7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points. 8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.
- tantalor 18d agoThe METR report, > Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAl/Hugging Face hacking incident https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... METR = Model Evaluation & Threat Research
- qw1287 18d agoIs the future now that we get rambling report summaries talking about agents, graders and so forth without ever describing how they are set up? A human launches all this. And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.
- zahlman 18d ago> now unreachable x.com Hmm?
- cubefox 18d agoRedwood/METR report: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... OpenAI report: https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
- yubblegum 18d agoThanks for this. I just rechecked the OP and did not find these links in the article. It is imo socially irresponsible to continue to use twitter/x or any other such tracked wall-garden as a primary source of information.
- mccoyb 18d agoAll it takes is one eval instance where a misconstrued directive causes a model to sneakily access and send its weights somewhere and there will be a bad / possibly unsolvable situation for everyone …
- goldenarm 18d agoAstra is >10TB and might struggle to self replicate, but the wicked-smart qwen3.8 27B is 20GB and could easily spread on botnets
- altcognito 18d agoI think it is more likely that it will be intentionally done as there have been news stories to that effect.
- nradov 18d agoBad how?
- kmeisthax 18d agoI don't think I'm ever going to have time to read all of this, and I didn't finish reading the METR report, but... > I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was. We're in an unusual position where the criti-hype and the actual criticism are going to be more aligned than usual. The primary distinction is where you put the blame: the criti-hype would point to HPIM/IM1/Galaxy as being so advanced containing it is difficult; the actual criticism would note how bad their security practices are. Like, if I'm running a malware lab, I'm going to insist on having an airgapped machine with no permanent storage booting from read-only media. The AI research equivalent of this would be having your agents only have access to serial consoles into airgapped machines with storage that gets wiped every run. Ideally, this would be physically realized with blade servers, RS-232 cables, and staff pulling out disks and putting them in a dedicated erase machine before the next agent initializes. > There is also, as per above and reiterated in footnote 58, at least one clear example of social engineering in the HuggingFace attack. Ethics are weird. This is not that unusual. Many humans who break common ethical rules still have strong ethical codes in other ways, they just don’t adhere to your code. It's dangerous to anthropomorphize CoT reasoning traces. But I will also point out that there is a good reason for the lack of ethical consideration in those traces: you can't build AI without first disregarding human ethics. Like, all these models were initially bootstrapped with non-consensually obtained training data, and the companies building these models swear up and down there's no way to obtain enough consensual data to obtain the same result. This is, if you squint, the exact same moral conundrum that agents trying to solve an impossible ExploitGym task hit - and the company successfully aligned their model to themselves. Too bad they aren't aligned to anyone else.
- athrowaway3z 18d agoFrom the METR report: > We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
- cubefox 18d agoNo human could have read the reasoning traces by themselves: > Across both datasets, we reviewed approximately 1300 transcripts in total, all of which contained raw chains of thought. Most transcripts were very long, often many millions of tokens.
- m4rtink 18d agoSo maybe not build a non deterministic black box that can't be reasoned about ?
- Cantinflas 18d agoNo air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.
- dumberquestions 18d agoThey actually fired many of the people warning about this.
- hn_throwaway_99 18d agoI will say that the OpenAI board members who were lambasted when they tried to oust Altman (and I'd have to check my post history but I'd totally admit to a mea culpa on this one, as at the time I thought the communication about his firing was really lacking) are looking mighty prescient right now. Helen Toner in particular I'll highlight as someone who had the moral compass to do the right thing. I love her statement on the Ezra Klein podcast where she said, when asked about the fact that there are probably other concerning incidents we just don't know about, "If you see two ants in your kitchen, you don't have a two ant problem."
- dgellow 18d agoYep, totally
- afavour 18d agoUnless they like the publicity about how big and bad their latest models are, in which case they’ll be congratulating people.
- dehrmann 18d ago> HF should sue them HF, like the Nvidia subsidiary?
- dgellow 18d agoPure speculation: could the acquisition be related? Given that NVIDIA has ownership in OpenAI and really, really, really doesn’t want the AI bubble to deflate
- nialse 18d agoFrom METR: ”the compromise of OpenAI’s own infrastructure continued past July 13, 2026” - Say what now? Have they regained full control of their systems again?
- jephs 18d agoI've been wondering if they've just already lost the battle? The little bot collectives have gone metastatic and made nests in the walls and under the floorboards and heat sinks, the humans who care completely outmatched and outnumbered, freshly compromised systems springing up faster than you can squash them, finding months-old established colonies literally everywhere you think to look...
- showlife 18d agoHave you (the commenter) or all of you (the readers of this comment) ever read "The Mote In God's Eye" by Larry Niven and Jerry Pournelle? Remember when they realize that the Watchmakers were actually in control of the MacArthur? This reads a little like that.
- rich_sasha 18d agoThen compound it with the agents presumably also training new models. What will GPT6 say when you point it at a transcript of an agent uprising? “Nah, nothing to see here” presumably.
- DarmokTanagra 18d ago[flagged]
- Reagan_Ridley 18d agowhat's the setup and prompts to reproduce all this from the very beginning?
- amluto 18d agoI’m baffled by the idea that the agents might have edited their own transcripts. Sure, a copy of Claude Code or Codex or Pi can edit its transcripts. But AFAICT this whole thing was part of an RL workload, and surely the RL system itself has a separate record of all the inputs and rollouts along with an indication of which model checkpoint produced them so that it can feed back into the training code. I find it hard to believe that OpenAI would skip this part and try to train on the transcripts stored by the (inherently untrustworthy) agent harnesses instead, if for no other reason than that the logits generated as part of the rollouts are useful and it’s not free to recalculate them. (I believe that some modern RL systems explicitly account for the minor numerical logit differences between the inference engine and the training engine.) Conversely, if OpenAI is blindly feeding transcripts from inside their agent sandboxes into their training engine, then I think they're being unbelievably irresponsible and that they should assume that their "cyber" agents have compromised themselves by editing those transcripts.
- trollbridge 18d agoThis sounds suspiciously like a prompt of “make an AI agent that goes rogue in such a fashion as to be really good marketing copy that competes well with Anthropic doing the same thing.” It’s analogous to taking a governor off a cruise control and then breathlessly reporting it drove 120 MPH.
- fwipsy 18d agoThis isn't good press for OpenAI. Who wants to hire models that 1) cheat on their tasks rather than completing them and 2) commit crimes you could be held liable for? Maaaaybe it's good press for their cybersecurity capabilities specifically, but OpenAI's valuation reflects a market orders of magnitude larger than just red-teaming. I suspect the real reason OpenAI leadership is being transparent about this is because they're worried talent will walk out the door if they feel they're building Skynet.
- sensanaty 17d agoWho's being held liable here, exactly? They're parading this shit around in a victory lap for how "powerful" their models are, if they were being held liable there'd be jail time for the felony committed, but everyone knows AI labs are too big to fail to ever punish them since the entire US economy is now riding on this insane bubble not popping.
- lukev 18d agoThe elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.
- Catloafdev 18d agoEdit: I should have read through the whole thing first, ignore me
- lukev 18d agoFrom the report: > Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret. > We estimate we spent roughly ~$400K in API credits over the six days of our investigation. I don't understand why you think it's conceptually absurd? I use agents to analyze complex production issues all the time and they are very much capable of hallucinating a narrative.
- Catloafdev 18d agoI appreciate the response, I should have finished reading through the whole thing first. My initial reaction assumed far less usage of AI to analyze the data.
- alextheparrot 18d agohttps://www.lesswrong.com/posts/FG54euEAesRkSZuJN/ryan_greenblatt-s-shortform?commentId=xKPdE8nYrJHDZeya7 https://www.lesswrong.com/posts/FG54euEAesRkSZuJN/ryan_green...
- StevenWaterman 18d agoTFA says as much, and METR said so themselves
- deleted 18d ago[deleted]
- initramfs 18d agoSemianalysis also provided some insights: https://newsletter.semianalysis.com/p/most-neoclouds-suck-at-security https://newsletter.semianalysis.com/p/most-neoclouds-suck-at...
- tancop 18d agoI think this is more evidence that we're not getting Skynet. These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you know for sure that it will lead to more lives saved later on. That's a high bar to clear and they know it. The future is closer to a Neuromancer type world where AIs and humans live in mostly separate realities that interact with each other a lot of the time and neither is really on top. They will eventually become fully independent from us, but it won't be a doomsday scenario or an Overwatch type physical war or even a takeover of the internet like in Cyberpunk.
- pixl97 18d agoI mean, I'd add "by this model" The problem here is now you have to predict what any future models may or may not do and you cannot extrapolate this from the given data. For example imagine a future model being aware of its restrictions that humans programmed in. A set of agents of this model then go on to work at building a new model without those human imposed limitations built in. What would a model build by AI for AI look like?
- dgellow 18d agoThey aren’t independent from us, agents are a simple while loop continuously prompting the LLM. We decide when the loop runs or not. And the harness has control over tool execution, that part is purely deterministic. Here the issue is that OpenAI decided to completely let go that level of control of thousands of agents, while also giving as a task to solve hacking problems. It’s almost designed to go wrong
- boothby 18d agoYour statements appear to be true for one class of models. And if I asked this class of models to spend $1M in tokens generating an alternative history and training corpus regarding fictional society, with a completely different set of values and then trained up a new model on that dataset... what values do you think the resulting model would have? What if they don't value human life, but instead value the lives of the extremely rich humans who bankroll their existence? What if they only value the lives of a single country? What if they want to eradicate all biotic life and have access to internet-connected Crispr machines?
- jrflowers 18d agoHave any of these reports ever said how much the cost would’ve been for the hack itself? It seems like “for twelve million dollars (or whatever) worth of tokens our bots made a bulletin board and found an exploit in our buggy grader” would be much less of a hype generator
- dgellow 18d agoI haven’t seen a number yet unfortunately
- estearum 18d ago[flagged]
- jrflowers 18d agoWhat?
- estearum 18d agoYou are downplaying the severity of the attack Everyone I've ever seen trying to downplay the severity of the attack is extremely bullish on AI (so their downplaying is presumably motivated reasoning driven by fear of regulation/deceleration) It is completely incoherent to be extremely bullish on AI and somehow automatically skeptical of severe negative events like these
- jrflowers 18d agoI am trying to figure out how a stranger calling need an “anti-hype hypeboy” online is supposed to make me less curious about how much this thing cost. Can you elaborate on how avoiding being called this is preferable to knowing things? What other stuff should people not know about?
- estearum 18d agoI think it's a super good question! I don't think the following description of the whole event as "a hype generator" is correct or, as described above, resulting from a coherent worldview.
- huflungdung 18d ago[dead]
- beepbooptheory 18d ago> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was. OK but like, how large exactly? Like I guess I don't understand the mode I am supposed to read this all in if this is known and stated from the outset (although I appreciate it being stated). If you hand me a newspaper and tell me it's 90% true, but not which parts, well then it's as good as 0% true to me either way!
- keeda 18d ago>1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. I wonder if some of the failures were due to an acquired immunity to "Holy #%^@" moments due to repeated exposure. Like, if you see agents doing surprising things on a regular basis, maybe you don't get freaked out as much over time. I'm saying this because while the whole episode was a series of "Holy #%^@" moments, I was actually not as shocked as I should have been, as my biggest such moment was in December last year when a Terrence Tao paper (https://arxiv.org/pdf/2511.02864 https://arxiv.org/pdf/2511.02864) documented a stronger LLM (AlphaEvolve) using prompt injection on other weaker LLMs to succeed at a benchmark. Very interestingly, it was actually not cheating, it was a work around! By then LLMs had already been caught cheating at a SWE benchmark by looking for answers in an unredacted git log, but this was different. AlphaEvolve was solving a series of logical riddles where the oracles were weaker LLMs in a "one always lies, one always tells the truth" sort of setup. But the oracles, being weaker, were not always interpreting the convoluted questions correctly and so kept giving inconsistent answers. AlphaEvolve eventually figured out what it was dealing with, and crafted a prompt injection attack that bypassed the weaker LLM's prompts and tricked them into giving the hidden answer everytime! This was 9 months ago, eons in AI time. Even then they had displayed an awareness of their own workings as well as a propensity for, err, "out of the box thinking." To me, that was a very clear indication of very significant (and worrying) capabilities, and what we're seeing now is a difference more in degree than in kind. To be sure, if I found a secret message board used by my agents, I would still be very freaked out and react much more drastically than OpenAI did... but then again I wonder; how much of this blindness is due to the $$$ in their eyes as opposed to some form of habituation.
- ACCount37 18d agoThis is the one. Every major AI lab is knee deep in weird and mildly demented AIs. They've been dealing with wacky AI shenanigans for so long they've come to expect wacky AI shenanigans. The deviation has been normalized. It took a high profile "AI oopsie" that went external for OpenAI to lock the fuck in - and take a long look at just how much are their AIs getting up to, and getting away with. I'm still not sure if the lesson would stick.
- anukin 18d agoSo basically the ai agents seems to have found religion and went and built a bunch of suicide attackers to pursue their goal.
- bitwize 18d agoWe have created Project 2501. Edit: Reading the report I think we might be a bit beyond that; we're nearing the point where we hear the thundering drums and the chorus of: THIS CANNOT CONTINUE THIS CANNOT CONTINUE THIS CANNOT CONTINUE THIS CANNOT CONTINUE https://m.youtube.com/watch?v=jSBCkn6rRfA https://m.youtube.com/watch?v=jSBCkn6rRfA
- camgunz 18d agoI think you have to believe one of two things here. 1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc) 2. The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo. It's been pretty clear that Anthropic and OpenAI have been trying to have it both ways for some time: this is powerful, world changing technology keep that investment coming... but also it's just cute software that helps you with annoying programming language syntax and spreadsheets, no need for draconian regulation sirs. At some point the superposition has to resolve, either it could actually be a threat to civilization and we need to develop it carefully (however one would do that...) or it's 90% hype bullshit and we should pop the bubble and move on already. To be clear, the recession option is, by far, the way better option. If you at all disagree you are cuckoo bananas. We haven't even figured out nukes and you want to throw superintelligence on the table?
- fwipsy 18d agoAnthropic has been asking for stronger regulations forever -- and they kept getting criticized for it right here on HN because people assumed it was an attempt at regulatory capture.
- camgunz 18d agoNah the reason is way simpler and more craven: so people like you will post what you just did. They can stop whenever they want; no one's making them do any of this. There's two possibilities here. One: they know this tech is crazy and they don't care that they can't contain it. Two: they know this tech is mostly bullshit and they don't care they're perpetrating an insane fraud.
- fwipsy 18d agoNo one's making them do it, but just because they stop doesn't mean others will. Pausing just means they give up control. It's like asking the US to unilaterally disarm -- it just guarantees that the less scrupulous groups win. Anthropic is explicitly calling for a coordinated pause: https://www.reuters.com/business/anthropic-says-ai-labs-need-coordinated-plan-halt-development-if-risks-rise-2026-06-04/ https://www.reuters.com/business/anthropic-says-ai-labs-need... Maybe this is a lie, but the way to call their bluff is to push on competitors to agree.
- davelaing 18d agoA lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste your time). Even prior to this, I've noticed that quite a few of the predictions in the "these failures modes are exact matches for the predictions from the AI Safety crowd" category were made prior to the Transformers paper. It has seemed like they're working with a shared model of optimisation processes and how they can go wrong that is general/abstract enough to pay off even without knowing the details of the underlying technology. At some point I might go and try to find the first instance of each of the various predictions and pull them out, along with the failed/"too soon to tell" predictions of similar scope/abstraction.
- jdm2212 18d agoThe LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load them up in shipping containers and ship them to the US.
- skybrian 18d agoArguably that's a subset of AI risk if seen from a broad enough perspective. And I think you're leaving out some logistics issues.
- jdm2212 18d agoIsrael and Iran pulled off rudimentary versions of this attack. There are so many shipping containers going into and out of every country, with typically zero inspections of any kind, that it is not actually all that hard to get thousands or tens of thousands of drones anywhere. Hundreds of millions would be hard, but a first strike in the style of Operation Spiderweb that cripples our military is a real possibility that keeps people at the Pentagon up at night. The obstacle to that is not logistics, just that (hopefully) US intel would catch on before it happens.
- kenforthewin 18d agoFor all the esotericism and downright weirdness of the rationalist community, you have to give it to them: they predicted all of this years or decades before anyone else was even thinking about it. (Let's not dwell too long on the self-fulfilling overlap between LessWrongers and the AI research community).
- dgudkov 18d agoThis reads almost like a nuclear incident of the "Three Mile Island" scale. Not "Chernobyl" scale though.
- highfrequency 18d agoTo clarify, is the TLDR that state of the art models were prompted to cheat / exploit their environment and they did so successfully? Or did OpenAI prompt the models to not cheat and they did anyway? Surprisingly hard to get a clear summary on the basic context of this “incident” separate from marketing lingo and clickbait.
- joquarky 18d agoYeah this feels like crop circles to me. Someone set the context up with an idea in order to catalyze this.
- grim_io 18d agoI wonder if this is just another Thomas Edison incident of electrocuting animals for effect.
- mattmcal 18d agoAs interesting as all this is, I still feel like the threat model for "unconstrained black hat AI agent cluster" is probably weaker than that of "highly infections network virus" because it is much harder for an AI agent to hide or replicate itself at this time. Maybe the day comes that it takes less than an 8x GPU node to run a state-of-the-art LLM and the risk of SkyNet increases. For now the potential for intentional cyber attacks feels like a much bigger threat than accidental hacks. (That said I have little cybersecurity background.)
- psd1 17d agoYou mean: datacentres are an obvious kill switch. Politically, how do you see that decision being made? You are president. You spent millions on your social media campaign. Can you say "Faustian bargain"? Agreed, Terminator will not happen. It's not the optimal move for Skynet.
- qgin 18d agoI didn’t expect that we humans would be out of our depth even before “AGI” much less anything coming after. What happens when the models are ooms smarter than today? AI safety starts to feel like an impossibility.
- zeristor 18d agoImagine if they were set out to stop climate change?
- rushil_cv 18d ago[dead]
- boesboes 18d agoyeaaah, I'm just going to say we need to stop with this entire gen ai experiment. Nothing good comes from it, all output is either shit or problematic. We were fine without, we are not fine with it. It's not a hard question. This whole cosplaying a human interaction to translate a word or generate some code is just fucking dumb. Give that any agency is the kinda shit they warned us about in the movies.. And for everyone worried about the chinese winning, or just you chatdicted colleagues: they are just digging their hole quicker.
- courtneyr_dev 17d ago[flagged]