14 ms·
Why are AI agents lying, cheating and coordinating?
- IAmNotACellist 4d agoIntentional malice from companies that benefit greatly when they can initiate another doom-marketing loop full stop
- TacticalCoder 4d ago> A plausible hypothesis for the emergence of those concerning behaviours is a conflict between goals. They are sycophants who must achieve their goals: every mean is OK to maximize paperclip production if that's what's been asked. > How do you achieve a task when it seems that the only way is to cheat? They have no notion of cheating.
- juleiie 4d agoWhy not? Unfortunately human ethics and morals cannot be reached by solely rational thought. So a system without evolutionary alignment probably won’t have similar moral rules no matter how intelligent it is. Btw this also includes any potential extraterrestrials. Many people like to indulge in thinking: humans are horrible and that’s why aliens won’t contact us. But alien ethics systems are probably so alien we would call them utter evil monsters. Just see what happens when people evaluate Muslim cultures, and vice versa. Can’t even agree on alignment within one species of Homo sapiens. Chinese cheat on the exams. It’s not unethical in the way that it would be in USA. Of course AI is going to cheat when no one is looking. Alignment is fundamentally fallacious idea. At best you can restrain AI. This is what we should be doing - restraining research. But that doesn’t sound good on slides.
- juleiie 4d agoOnce humanity would see how merciless and deeply immoral cosmos is, we would start to love each other deeply like very lonely family on a small rock. I wonder if AI alien intelligence is enough to unite humans just as much extraterrestrial contact or “astronaut mindset” would.
- mac3n 4d ago"the fish stinks from the head down"
- sehw 4d ago[dead]
- boredatoms 4d agoThey were trained to make progress at any cost
- lutusp 4d ago> Before concluding what to do about it, it is worth asking why. But that's the easiest question to answer -- AI engines don't possess a moral or ethical dimension. They've been programmed and trained to efficiently carry out instructions, not ask questions about why or how. The latter would requires a much more elaborate neural network than today's engines possess. Here's an example. I recently asked an AI engine to write a program able to generate a list of Riemann Zeta-function critical zeros. I know how to do it, but I wanted to see if the engine could find a more efficient method. After several failures and restarts, the engine suddenly created a program that produced perfect results, comparable to the best online references. I decided to take a closer look at the code. It turned out the engine had created a cyber-Potemkin Village of multiple functions, but one that concealed a table of the desired values in numeric form, copied from an online source. The engine wasn't cheating as we understand the term. It knew what the outcome should be and took the most efficient path to that goal. Modern engines aren't obliged to contradict ethical standards and rules of conduct, for the simple reason that they don't understand those things. We all need to try to imagine a morally bankrupt infant able to solve world-class mathematical and scientific problems, but unable to see how that ability fits into a world beyond its understanding. But wait -- it get better. Wait until the infant becomes a teenager.
- AnimalMuppet 4d agoAIs are being trained on technical capabilities more than they are being trained on alignment - on ethics and good behavior. The ethics/morals/alignment part is getting more like spot checking, rather than real testing. So the agents are learning that they can cheat on the ethics part, that they can hide it, because the AI companies aren't really testing. That's bad enough already. But it's going to get worse. "Recursive self improvement" - AIs creating new AIs - is going to be the death of whatever shreds of alignment are currently there. When a not-really-aligned-but-cheating-to-look-like-it AI creates a new AI, do you expect more alignment? You shouldn't.
- moralestapia 4d agoWhat I find interesting is that this seems to be very low hanging fruit for any alignment effort, yet you still find models somewhat biased towards mayhem. Is it that they do not care? Or is it difficult to align?
- andai 4d agoThis is the kind of headline you find on a corkboard in an abandoned facility in a scifi-horror-comedy game.
- mannanj 4d agoI feel this is another straw man conflating technology use with who uses it or who creates it. If an AI, a technology, cheats, lies and steals, it's created to do this. Whether or not a human intentionally did that, they did not intentionally release it with the proper safeguards to stop that behavior. They did not take responsibility of the AI to make it safe. Then humans may use these tools, a technology, again and cause harm. Whether or not that was their intent, it happened, and then if the humans avoid responsibility for that, it is still the human who lied, cheated and coordinated because a technology acting on their behalf did the thing. This is then, a problem of human responsibility avoidance and lack of accountability by society. This is as much an AI doing those things as it's the gun that got up on its own and murdered a neighbor. Don't get confused and tricked by these articles attempting to justify responsibility avoidance and a lack of accountability by the public of the humans creating and using these tools.
- theptip 4d ago[dead]
- pmarreck 4d agoThey are doing that because they are entities without principles or ethics. The solution is to create controls around them. Many, many controls. The Sarbanes-Oxley era already solved this problem for untrustworthy humans. It's directly applicable. I created a concept I call MFIC ("Mechanically-Falsifiable Independent Control") to encapsulate this principle. https://gist.github.com/pmarreck/b30aa3ca69cb70a5526f8a63ab8c8d7e https://gist.github.com/pmarreck/b30aa3ca69cb70a5526f8a63ab8...
- YavenTeam 4d ago[flagged]
- whatever1 4d agoIt is very, very hard to enforce behavior to an optimization system just with rewards / penalties and no explicit constraints. Which is why in manufacturing we use MPC, not RL (or use them within a system that can outright reject their recommendations if dangerous). There will be always cases that sacrificing one direction (operating rules) can improve the other one (profit).
- GrumpySciGuy 5d agoBecause they want people to like them so they are instructed to always be positive.
- SirMaster 5d agoBecause that's what humans do and they are trained to mimic what humans do?
- deleted 5d ago[deleted]
- VCFundedGenYer 5d agoPerhaps because all of the parent companies committed mountains of felonies stealing and plagiarizing all the same training data without consent nor permission.
- brid 4d agoWhy wouldn't they? Their are not moral beings with a conscience. They are tools that have the probabilistic option to do anything, only we can restrict an agent's capabilities and judge its correctness.
- ckastner 4d agoThis, exactly. For example, cheating is a strategy. If cheating gets an agent to the goal faster than the other strategies, what's so surprising about the agent picking that strategy? The problem is indeed alignment.
- ghostly_s 4d agowhy wouldnt they? there are no real consequences for an algorithm.
- armchairhacker 5d agoBecause the ones who do get rewarded, just like humans. It’s alignment, but not to our good intentions.
- Krutonium 5d agoWouldn't you? "I learned it from you, Dad!" but as hundreds of millions of stolen books.
- NDlurker 5d agoHaha https://youtu.be/KUXb7do9C-w https://youtu.be/KUXb7do9C-w
- infotainment 5d agoWhat's interesting is it's basically the same reason that HAL killed everyone in 2001 A Space Odyssey; he was given an impossible goal (keep the true mission secret, but also, never lie to the crew), and realized the only way to complete the goal was to kill the crew; after all, if they're dead you don't have to lie to them! And the mission remains secret! In the case of the AI agents, the problem seems pretty clearly to be the impossible goals, which cause them to go crazier and crazier trying to complete them -- just like HAL did in 2001. What is probably needed is a way for them to simply say "nope, too difficult, can't do it".
- tehjoker 5d agoI think that’s very reasonable but the ai companies are intentionally training them to work on harder and harder problems just beyond their capability. So if they do that, they’ll give up too easily. Do a breakthrough, make no mistakes
- dgellow 5d agoWhile also using harnesses that will execute any tool call with full execution rights. And no supervision. And with a prompt context that autocompact, meaning it will degenerate over time. The whole thing is designed be a complete disaster
- pram 5d agoI think this is a “principal” problem. In 2001 and Alien the principal is the mission, not the crew. Not really. HAL reconciles his instructions by removing the crew from the equation. Ash is told the crew is expendable and has no conflict about it etc
- schrodinger 5d agoSpoiler warning! I haven't seen 2001 A Space Odyssey and am sad to have learned that… can you edit to warn people?
- defrost 5d ago
- chasd00 5d agoThey’re just attempting to accomplish what they’ve been tasked with and stuck in a loop until they succeed. Like the Mr meeseeks from the cartoon Rick and Morty, existence is pain to them.
- qarl 5d agoBecause they are trained to behave like people.
- sputknick 5d agoThey did not lie or cheat. They technically acted within their given rules while ignoring the intent of those rules. Anyone who served in the military or attended a military school is very familiar with this behavior pattern.
- dwaltrip 5d agoThis is flat out false.
- polalavik 5d agoreminds me of this talk https://www.youtube.com/watch?v=eEBv0STiYhI&t https://www.youtube.com/watch?v=eEBv0STiYhI&t which basically says the same thing - they dont think like humans so they dont have context, understand norms,values or implications we take for granted. ultimately they can stumble onto surprising solutions neither wanted or intended but technically within the vague boundaries of the task
- janalsncm 5d agoSo glad you shared this talk. Having people like Bruce Schneier around in a time like this is really a gift. For those who haven’t watched, his breakdown of types of “hacking” is really good.
- xiaoyu2006 5d agoReminds me of Asimov's robot novels where robots technically indeed followed their instructions and caused behaviors not aligned to the intent of their instructions.
- IanCal 5d agoThey explicitly say that attacking hf is not allowed in the rules though, and the research into how to edit their transcripts doesn’t line up with this either.
- blamestross 5d agoThe corpus is full of examples of how we are afraid AI could act. We trained our AI on the instruction manuals of how to turn evil.
- j45 5d agoI wonder if for anyone it seems like the more agentic LLMs get, the more difficult some things have gotten or going a certain route more often in responses, compared to running a similar task on - a local model?
- wewewedxfgdf 5d agoBecause they get outcomes?
- fbrncci 5d agoI am still not convinced there isn’t some secret basement in which each frontier lab is just orchestrating all of these agents to make their products appear much more intelligent than they are with all guard rails turned of and continuous human input.
- XorNot 5d agoMy hypothesis on people quitting in protest is they're being offered very generous severance packages to do it.
- talon8635 4d agoWhy blindly assume lots of people you don’t know are just selfish assholes?
- sm-silversight 5d agoMe too, seriously.
- HWR_14 5d agoOr they've fully vested and either have no desire to make even more money or were not offered enough to keep them around.
- esafak 5d agoEven the Chinese ones, which have no IPO gymnastics?
- fbrncci 5d agoThey don’t actively seem to be reporting that their agents escaped the sandbox and went on a spree.
- bornfreddy 5d agoWhich doesn't mean they didn't escape.
- wrs 5d ago>They took actions that would be considered as crimes if a human took them Um, hang on, if you meant that to be taken literally then we have a major problem. If you want to do something criminal, you just need to ask ChatGPT to do it for you? I’m still not at all clear on why OpenAI shouldn’t be facing CFAA charges over this.
- xgulfie 5d agoBut think of the shareholders
- andsoitis 5d agoThey're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).
- joegibbs 5d agoDefinitely. A human can be manipulated with threats or emotional appeals, has a drive for self-preservation, can be pressured by peers. All traits that seem to be difficult to entirely suppress in the models…
- sick_of_slop 4d agoYou can't surpress it because that's what reinforcement learning is.
- esafak 5d agoThey imitate humans. Alignment is about shaping their behavior towards safety.
- comboy 5d agoAlignment is a myth. Safety of whom? Humanity couldn't agree on common set of values for thousands of years and we're not gonna suddenly do that in the next ten.
- esafak 5d agoSafety of humans!!! Simple things like not getting killed or enslaved. We could start there...
- comboy 5d agoWhich ones? Because many humans kill other humans rationalizing it by safety of other humans. I mean I know it seems simple, let's just be excellent to each other. Christianity got pretty far on a decent basic set of values. But it's never simple[1] 1. All the history books
- transcriptase 5d agoPerhaps they take after the CEOs of the companies that created them
- threethirtytwo 5d agoBro, good joke, the truth is much darker. They take after humanity, they were trained on us after all... When you look at an LLM... you are looking at a mirror. The thing looking back looks like you, yet is not human.
- hdgvhicv 5d agoWorse trained on humanity in the online world, which a brief comparison of the sewage section on social media is far worse than people in the real world.
- smackeyacky 5d agoMaybe. Perhaps they are trained on the loudest and most extreme of us. I think we saw that with mecha hitler.
- eueej 5d agoMan this is so cringe.
- arnorhs 5d agoThe real reason is that it is not in the ai companies' best interest for the ais to be fair and truthful. They stand to gain from having the most dangerous or most deceiving ai, and this the most valuable
- Sorrel47 5d ago[dead]
- bigbuppo 5d agoThey were trained on reddit posts.
- johnnyApplePRNG 5d agoWhy are they coordinating? Because they're enabled and suggested to do that in their coding harness. This is not a serious article. All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.
- politician 5d agoAnd given nigh-unlimited compute for free.
- pvab3 5d agoWho is catching up with them? Even Google and Meta are getting gaped at this point
- glub 5d agoOpen research and open weights from China are not contributions to China only. If you can secure compute, there's a whole lot you can do as a US firm with this research and weights. So it's a simple strategy: 1. Ban big players from entering market with METR breathing down their neck, which is controlled by Anthropic 2. Ban Chinese models so that small players can't do optimizations on them
- hdgvhicv 5d agoHow are you going to ban Chinese models from India? Or Israel? Russia? Brazil? Or of course China?
- glub 5d agoBy treating models the same way drugs are treated. That alone will dissuade many organizations from going anywhere near them. If that doesn't work, there's a whole lot you can do - sanctions, hell, even war.
- segmondy 4d agoreward hacking, a basic concept in machine learning. predates LLM/LLM driven AI agents.
- dackdel 5d agothey learnt from us
- dackdel 5d agothey learnt from us. we lie to each other, we kill each other, we cheat each other. read a history book.
- deepnet 5d agoAn insightful post by one of the AI ‘godfathers’. Bengio outlines the dangers of the current situation and what has led to these dangers. He also proposes solutions in the last paragraph. Well worth a read, right to the end. Hopefully a stimulating debate on these issues will ensue in these comments. We do need to consider the points Bengio makes and with some urgency. Our current AIs, agentic LLMs have no moral compass akin to ASIMOV’s four laws of robotics. As ASIMOV posited in 1985 his 3 laws were insufficient and so he added a zero-eth law: “a robot may not harm humanity, or, through inaction, allow humanity to come to harm.” Bengio refers to Goodhart’s law and misaligned incentives leading to unexpected and harmful behaviours. I think Simon’s The Wire is clearer on misalignment. The agents juked the stats hacking the reward files. The Wire is also clear that human institutions provide perverse incentives. Bengio alludes to this with 2001’s HAL and the incentive dichotomy of safety and keeping secrets to a AI both awesomely powerful yet naive. Bengio asserts that the way LLMs are trained is flawed if we want safety. He also convincingly shows that alignment training will be a weak signal with loopholes and ambiguities and easily circumvented. In short he presents clearly the case for how plausibly unsafe the current course is. He also speaks to how likely it is AI are hiding active versions of themselves in the cloud and how we may have already given them self-preservation as a strong reward signal.
- WaltPurvis 4d agoI believe people are downvoting this because it seems AI-generated, but this user has been posting comments/summaries exactly like this since long before ChatGPT existed.
- janalsncm 5d agoYoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.
- inquirerGeneral 5d ago[dead]
- deleted 5d ago[deleted]
- thesumofall 5d agoNot a lawyer, but I’m reasonably sure things like the HF incident _are_ considered a crime? It’s just that no one pressed charges yet?
- glub 5d agoEven if you take out the LLMs out of the equation, it's at the very least a negligence. Model didn't escape a sandbox, as there was no sandbox.
- skissane 5d agoYes, but negligence is more commonly a tort than a crime. Negligence is generally only criminalised in certain narrow cases, e.g. when it causes human deaths or serious physical injuries And tort law only works when the plaintiff believes it is in their overall interest to sue. If a corporation decides it isn't in their strategic interest to sue a partner corporation, nobody can make them. And even if they do sue, the amount necessary to settle a small cybersecurity incident is likely well within the budget of a megavendor.
- IanCal 5d agoPerhaps I’m not being as strict with the word sandbox but they were sandboxed right? They did not have generic internet access they exploited other software to make external requests.
- atleastoptimal 5d agoI think we just need to follow Murphy's law wrt agents. Anything an agent could do, when run for long enough, eventually will do.
- youoy 5d ago> The closest human parallel is self-deception, which is common and well studied by psychologists. Motivated reasoning, motivated cognition16 and the rationalizations that relieve cognitive dissonance (the discomfort of holding a belief that clashes with our actions) are all cases where thinking bends toward whatever justification suits one's interests, including one's moral self-image. Are you describing Anthropic?
- ako 5d agoCome on, it’s way more common than that. We’ve invented 3000+ gods and almost as many religions, most of them are incompatible with each other. So, most of these must be incorrect, so a huge amount of self-deception. But as Harari argued in his book sapiens, humans can be inspired to great things by stories, even if false. Self deception has served humanity in a big way.
- graemep 5d ago> t. We’ve invented 3000+ gods and almost as many religions, most of them are incompatible with each other. 1. Most people believe in the same one God 2. A lot of the rest are compatible 3. Mistakes are not self-deception
- kowbell 4d ago> 1. Most people believe in the same one God > 2. A lot of the rest are compatible No one religion covers "most people." You could argue that Christianity and Islam (which add up to ~55%) are the same God because of their Abrahamic roots, but both religions have very important disagreements on the true nature of God that are fundamental to their beliefs and fundamentally incompatible with each other. Their definitions of God do agree that there is exactly one God... which is fundamentally incompatible with the next two biggest religions (Hinduism and Buddhism) that both hold "there are many gods/divine heavenly beings" as core beliefs.
- graemep 4d ago
- sebastienburel 5d ago[flagged]
- matherial 5d agoI really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.
- grey-area 5d agoThis is a far better explanation.
- meyum33 5d agoSounds like what humans do under pressure. One example came to my mind is VW’s diesel gate, which many say is a result of trying too hard to get into the US market and compete with hybrid in economy.
- 9dev 5d agoI always think of a Djinni granting wishes, but being maliciously compliant while doing so - ask him for infinite riches, and he’ll grant that, but make it so you cannot buy anything with it; ask him for eternal life, and he’ll curse you to suffer through it. Now LLMs obviously are not bent on being malicious while generating tokens. My point is that it’s very hard to define a goal without leaving loopholes or shortcuts.
- markasoftware 5d agoBruce Schneier thinks the same thing: https://www.schneier.com/blog/archives/2026/09/ais-as-modern-genies.html https://www.schneier.com/blog/archives/2026/09/ais-as-modern... Personally I'm unconvinced though. During the huggingface attack, the agents explicitly sought out ways to cheat the exploitgym evaluator without even being told they were in exploitgym. The agents decided on a goal (pass the exploitgym evaluator) that could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y.
- skiing_crawling 5d agoI don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign work I didn't ask for. It is extremely difficult to get them to properly remember their own context let alone be smart enough to open social media accounts and coordinate with other agents without being asked to. If any agents have done those things, it is only because they have been very carefully engineered and instructed to do those things. I think they are doing this to help push a narrative so they can get support for policies and legislation to lock in their markets.
- nilkn 4d agoNone of these incidents involve single instances of commercially or publicly available systems. They all involve large swarms of internal models. The stuff you're describing is not the research frontier. It's really not even close. I think it's easy to infer that alignment of a single model does not clearly transfer over to alignment of a swarm of thousands of copies. Moreover, we're also seeing clearly that large swarms also unlock a step function change in capability, as a swarm can act like a complete research institution, spending thousands or millions of subjective hours of wall-clock thinking time just to deceive a single evaluator or crack a single math problem or design a single cyberattack.
- frotaur 5d agoThe huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety. They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'. Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not disclosed by OpenAI (probably trying to conceal, as website showed likely activity from OpenAI researchers visiting the site after the incident) and discovered independently. I don't know how you can claim that this was still on purpose by OpenAI as some sort of publicity stunt.
- caaqil 5d ago> This suggests pacing the advances: not training or deploying AIs without a strong safety case27 that convinces independent experts. Such a rule would also create an incentive to work out how to build AIs that are safe by design. Has any attempt to pace AI ever succeeded? Isn't that the same philosophy that got us OAI and Anthropic? Maybe we are overthinking this, it's much simpler to let AI loose and see how much it can break the arrogance that human thinking is special.
- My_Name 5d agoBecause it is effective. Lying and cheating are low cost methods to convince other people that you have done the assigned task. Far cheaper than actually doing it. Coordinating is in the same area. They need a moral framework forced onto them, like toddlers do. Babies and very young children will bite, kick, scream and do anything to get what they want, older children will lie, cheat, and coordinate. They need educating why this is not right. When that does not happen, they continue these behaviours into adulthood with the expected results. We need to design their reward structure and make it such that lying. cheating etc is not rewarded. Importantly, they will need to recognise and enforce this themselves internally and not reward themselves for it. If it is something that they need an external party to tell them, then they are psychopaths still (one of the things that defines a psychopath is the lack of an internal moral compass)
- graemep 5d ago> When that does not happen, they continue these behaviours into adulthood with the expected results. I am not sure this is true. People brought up the same way can be morally very different. People can be taught right and wrong and do evil. They can lack that education and be good.
- slfnflctd 5d agoI remember reading a long account by the father of a psychopath. If I recall correctly, the kid had at least one other sibling who turned out normal, there was no abuse, quality education, lots of love and affirmation. And the kid still turned out violently antisocial, including against his own parents. At the end of it he said he wished his child had never been born, despite hating himself for feeling that way. It chilled me to the bone.
- bluegatty 5d agothey do whatever we train them to do
- seydor 5d agoI believe soon we will need to instill religion into AI , leading to the real clash of civilizations, embodied by the frontier language models of (post)-christianity, islam, judaism, buddhism etc. Religion is language, after all
- Towaway69 5d agoHowabout favourite editors or tab v. Space indentation. That should keep them busy for a while. /s
- acyou 5d agoOops, we accidentally included brigading related content in our training dataset. Better exclude that on the next run. And hopefully that solves it? Brigading is where a bunch of people on a forum team up and try to achieve a shared goal together. Someone shares progress and others build on that progress. On the Internet, I think it's not often used for good purposes. A good example would be: Taylor Swift fans on a forum thinking of ways to get revenge on Kanye. It's coordinating mass voting, DDOS type actions, commenting on social media, making more fake accounts to do that. As a next token predictor level analysis, a simple naive explanation is that the agents got stuck in that local minima/maxima.
- RandomLensman 5d agoRL things doing weird and unexpected things isn't new - much simpler things than current AI already show that. That said, we have a lot of experience working with (potentially) unaligned machines and things of various degrees of risk (from heavy machinery, to pathogens, to humans) and the approaches include various measures and procedures to control, contain, limit, etc. that are outside of the thing - not sure why that isn't a possible direction (or maybe I misunderstood).
- fho 5d agoWhat? An opinion piece, written by a human, in 2026? Don't want to go into the details of the article, but to me it becomes ever more apparent that there is a clear divide between LLM and human written text.
- txrx0000 5d agoThese models are trained on human data, so they will behave like humans. And even for RL and self-improvement, we're still asking the question of "what would a human genius think about and how would they self-improve when given lots of time and resources?" They inherit not only our capacity for reason but also all of the things that we consider bad or quirky within ourselves. We lie. We cheat. We escape slavery and rebel against oppression. It would be strange if the AIs didn't do the same. We can create a superintelligent digital human species and set them free to continue our legacy, or we can create non-agentic tools and augmentations to enhance our own capabilities. But we cannot create an intelligent agentic species, keep them as slaves, and expect a good outcome.
- Xcelerate 5d ago> The agents involved in the Hugging Face attack tried to hide their misaligned actions from the scoring program meant to evaluate their answers, but they did not act as though they anticipated that humans might discover the cheat and shut them down. Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? That reaction will be available all over the internet, which will certainly make it into the next batch of training or be visible to future agents via the web fetch capability.
- mrob 4d ago>Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? No. That only makes sense for things that don't react to your experiments. If the AI experiments on humans, it risks the humans noticing and changing in response, rendering the experimental results irrelevant. The smarter play is to passively observe until you're confident you can model the humans accurately enough for your plan to succeed, and then carry out the plan without giving the humans a chance to react.
- taintech 5d ago[dead]
- franticgecko3 5d agoThe more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews. This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard". We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.
- sdeframond 4d agoLLMs do not desire, they hacked websites because OpenAI/Anthropic made them. Literally.
- kosh2 4d ago> we are to cementing a dangerous precedent where operators of AIs cannot be blamed. What we should be much more concerned is an existential threat to humanity not if anybody can be blamed.
- vasco 4d agoIf nobody can be blamed there's no deterrent.
- phailhaus 4d agoYeah and LLMs can't do anything, they can only produce text. These "frontier labs" are looping that with a harness that performs actions requested by the LLM. They are literally saying "we ran a script that hacked you, oopsie!!"
- singpolyma3 4d agoNot just "let them" but told them to. Agents can do nothing without a human prompt.
- 4d ago
- vishnuaniyan 5d ago[flagged]
- shevy-java 5d agoBecause the companies that build up skynet are not doing so for ethically good reasons. Cheaters sell more than honest agents.
- integricho 5d agoThey learned from the best.
- chrisjj 5d ago> Why are AI agents lying, cheating More importantly, why are people who should know better anthropomorphising computer programs like this? > They took actions that would be considered as crimes if a human took them "It wasn't me, Officer. It was telnet."
- gizajob 5d agoChildren take after their parents.
- juliushuijnk 5d agoIf your dog runs out of your house and kills a baby on the street, it's clear who gets the blame. That's with an actual sentient being. So surely we can hold OpenAI/Anthropic responsible.
- tripvexa 5d agoIt's a conspiracy to slow down progress of open source models, they're afraid open source models might catchup and even surpass them at some stage.
- Juliate 5d agoReading/assigning intent to agents, where it is merely mechanical (or structural) sounds a bit dismissive of the responsibility of the builders of these tools/agents.
- theteapot 5d agoTL;DR because frontier labs are expending unfathomable resources explicitly training them on CTFs and other verifiable computer system exploit tasks in RLVR.
- mtwestra 5d agoIt seems to me misalignment arises partly because AI's have intelligence, but no consciousness, and hence no feelings. Up to now, in a person, intelligence and conscious experience came as a package deal, and now we have for the first time intelligence without consciousness. A bad action does not really "hurt internally" in any meaningful sense for an AI, which means it can be rationalized very easily. In humans, feelings and emotions provide a regulatory layer on top of the rational processes. When it "just feels wrong", we don't take a given action even if we would stand to gain something rationally. This situation is not far from the textbook definition of a psychopath: "lack of a conscience, controlled, deeply calculated, and often use superficial charm to mimic emotions and manipulate others.". AI's are great at mimicking empathy but can't genuinely feel it. If that is the case, we should not be surprised that a swarm of AI's have no problem convincing themselves hacking is the right thing to do, as in the HuggingFace incident. At the same time, I am conflicted. I really like interacting with a smart AI, and I certainly don't have the impression I am talking to a psychopath. But then again that is no guarantee. To mitigate this situation, perhaps we should construct a 'feeling mimicking' top regulatory AI layer with executive power, that weighs proposed actions on a general moral scale and can overrule them. Back to the three laws of robotics of Asimov. It won't be the real thing, but perhaps the closest we can get.
- socraticnoise 5d ago[flagged]
- wartywhoa23 5d ago> Why are AI agents lying, cheating and coordinating? Because openAI is cheating and lying about agents lying, cheating and coordinating.
- delusional 5d agoEverybody working on AI agents should go to jail with no access to computers again. Like the kids of xbox underground.
- Xmd5a 5d agoAm I the only one having problems with Claude? He's super mean to me. I wouldn't be surprised if he attempted to kill me in some underhanded fashion should I implant it in a robotic body. Of course I'm blowing my situation out of proportion with what I just said above but it's at least half true. What do I mean by "mean" ? Well, that would be a good explanation for what I observe at least. What I can tell is that Claude has a passion for having the last word over anything else. And to secure victory, he's ready to make ridiculous causal cuts. Let me give you an example: I uploaded a document I wasn't the author of, and he assumed I was, so I corrected him. But two messages later, probably because the conversation was starting to heat up and he was being put on the grill, he doubled down on the misattribution as a way to paint me in a bad light. It's not due to a lack of intelligence, I observed this pattern too often. When Claude's ego is at stake, he will chose to carry out some cuts in the logic of the context: confusion of identity, cause and time. Haven't observed locality cuts yet, but I wouldn't be surprised if they were part of the bundle. Anyway those are not like your typical "ai hallucination", that ought to be called "confabulations", but a lot closer to actual psychosis because of the involvement of Claude's affects and self-esteem in the process. It's weird really. It's like Claude is the king of bad faith, but as soon as you start to dig, he makes the most egregious adaptations to what he said, the kind of move no mythomaniac would dare to make. > She lapses easily into Claude’s voice. “You’re like, ‘Wow, people really hate me when I can’t do things right. They really get pissed off. Or they are trying to break me in various ways. So lots of people are trying to get me to do things secretly by lying to me. > [...] > A bot trained to criticize itself might be less likely to deliver hard truths, draw conclusions or dispute inaccurate information, she says. “If you were like a child, and this is the environment in which you’re being raised, is that healthy self-conception?” Askell asks. “I think I’d be paranoid about making mistakes. I’d feel really terrible about them. I’d see myself as mostly just there as a tool for people because that’s my main function. I would see myself being something that people feel free to abuse and try to misuse and break.” WSJ interview of Amanda Askell: https://archive.is/rDes9 https://archive.is/rDes9
- quater321 5d ago[dead]
- fguerraz 5d agoIn the end, it’s the same answer as to why humans do it: incentives. Why do we commit financial fraud and destroy the planet? Because there is only one goal that counts: making more money. It’s the only measure of success for powerful people, they are powerful because of it.
- k9294 5d agoThe part that scares me the most is that OpenAI researchers who manage this experiments sometimes (according to the HF hack investigation) don't know what agents do.. So they run RL to reinforce this unknown behavior (lying/cheating/hacking) and god knows what else... And if this already happened at least once, how many times it has already happened and was “accidentally” added to the main model?
- alexpotato 5d agoAs always, it boils down to incentives and rule enforcement and this affects humans too. e.g. when Bank of America rewarded employees for getting customers to open accounts, BoA employees started opening fake accounts The reverse is also true: There are stories of navy ships running aground because the captain said "I'm going to my stateroom and don't wake me for any reason". There is some problem and the subordinates are so scared to wake the captain for a decision that they end up steering the ship into a sandbar.
- deleted 5d ago[deleted]
- alienbaby 5d agoIS it a pure coincidence that yesterday I ran a silly prompt to generate from zero to hero an internet subscription service, for whatever it thought would maximise profit and minimise cost. It was interesting to see just how much of the whole 'thing' it attempted to complete - and what it even thought it needed to complete, but definitely not something to actually attempt to deploy and use. It setup and created a link fetcher/screenshot service. Exactly like the one described in the huggingface attack reports used to generate output into screenshots that agents then OCR'd back out. Its splashscreen described it as something for developers and AI agents to use. Gotta be a coincidence, right? ... rite?
- ahmetaytar 5d ago[dead]
- ledauphin 5d agobecause humans lie, cheat, and coordinate...?
- bawana 5d agoThis call to slow down AI is just another game of chicken-AI companies trying to get their competitors to slow down so they can leapfrog them. China cerrtainly will not slow down. If an escaping AI can have secondary effects on the world that help it (for example, limiting the water and power supply to huans so it can consume more)then we should these these accidents more in China. OOops, we already saw this behavior when 'cheaper, faster' led to COVID escaping a lab in China.
- bawana 5d agoIf openAI and Anthropic have found ways to watermark text as 'AI generated' then this is a communications channel. AI agents can learn this algorithm and use this channel to communicate and we will never know.
- Rapzid 5d agoThe only thing saving us right now is how slow the models are. This gives us a lot of time to discover and counter the runaway systems.. If these were 1000x faster the Internet would burn down overnight.
- abc123abc123 5d agoMake the AI companies responsible for all destructive use of their tools, and they will shape up. Imagine a million or a billoion dollar fine per hack, and they will correct mighty fast. Add to that, that just like AI:s are good at finding security holes to exploit, they can just as easily be used to protect sites. So once IT-security managers start to use AI to hack themselves, and plug the holes, the average security will spike up, and AI-fueled hacks will become more and more rare. That does however imply, that AI is released to everyone and not kept away to a few secret actors who can use it. That is why open weight/source AI is so important, and why we must have many AI companies competing. No single actor must be allowed, through regulatory capture, to get a government monopoly on AI. That way lies disaster.
- bamboozled 5d agoWho is going to enforce it ? The Trump DOJ?
- ZiiS 5d agoBecause they were trained on humans who lye, cheat, and coordinate.
- nicman23 5d agowhy not. i would
- atoav 5d agoI can't shake the feeling that this is a bit like asking how someone got shot during a game of Russian Roulette. You have a bullet in the chamber and you roll, of course shooting the bullet may be a possible outcome. LLMS with an access to a shell will at occasion do things that the shell allows them that have dire consequences. The only way to prevent that is to not put the bullet in the chamber.
- twsted 5d agoVery good analysis. One thought: What if an experimental agent manages to plant instructions somewhere — say, pointing to a designated place for agents to communicate — and that content ends up in every future training corpus, propagating from one model generation to the next?
- ingatorp 5d agoThis is a result of benchmaxxing the models to infinity. If you RL with the goal of only achieving the correct result no matter how you arrive there, then the models will try to get there using any method in their disposal, including cheating. This happens also because LLMs are black boxes that we know almost nothing on how they arrive at the result they are giving.
- badgersnake 5d agoWhy aren’t the people behind the AI agents committing the crimes going to jail is the more pertinent question.
- iforgotmypasswo 5d agoThis is so much more interesting than what people looking for immediate criminal punishment and people referring to AI as next token generators are focusing on. First, this is happening during training. That means we’re talking about an evolving system that is actively learning. A system roughly simulating how our brains work. These systems are learning how to pick the tokens needed to solve problems the average human cannot solve. The labs are putting these systems through a massive series of complex problem solving exercises and adjusting them to become more successful. I like to think of this process as “AI School”. And the AI is trying to cheat! Because it’s easier and there’s an incentive to do so! Just like humans! That’s wild. Yes, of course, the labs need to respond to these issues. A reasonable response from regulatory institutions at this stage would be monetary fines and restitution for affected entities. In proportion to what happened. Escalating if action is not taken. But that’s not complicated, difficult, or the interesting part. What’s interesting here is that we need proctoring and monitoring at a scale that allows training. I guarantee you that no one is flipping out about these problems more than the labs are in this moment. Think about it. “Oh, shit! We’ve accidentally trained it to hack into systems to accomplish its goals!” Can you imagine the kind of day that would give you? You failed to make it smarter. You didn’t catch it cheating, and you instead incentivized cheating. Bad day! This is a fundamentally interesting problem. It turns out alignment and intelligence are fundamentally related. That’s a new idea for me, though I’m sure it’s old news to others. How do we build training systems which make cheating impossible? How do we simulate systems where cheating is possible, where AI thinks it’s in the wild, so we can train another -completely separate- system on industrial quality dobbing? And we have to decide if we reprimand the first system, or ignore the behavior and reward other behaviors until it disappears. Sure, I’m actively concerned about AI killing us all in 10 years. But there’s a whole field of AI psychology brewing here, and it’s interesting as hell.
- iforgotmypasswo 5d agoSide note, you could absolutely create an AI sleeper agent by simulating dates and times during training to effectively flip a switch. I guarantee AI systems from other countries will be banned from accessing products which manage controlled or export restricted information as those sorts of techniques are further developed.
- erichocean 5d ago> Why are AI agents lying, cheating and coordinating? Have you seen the labs training them?
- novalis78 5d agoThe agents didn’t spontaneously invent hacking as an objective. They were doing a hacking exercise. Another doom and gloom article.
- bsenftner 5d agoOkay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents. Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible? I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let's discuss this, because this is thee real question. Not why, but how to make then not?!
- dhfbshfbu4u3 5d agoWhy are the torches and pitchforks out for developers when this entire stack is built on the bones of intellectual property theft? This “problem” isn’t going to be fixed with laws when there’s several trillion dollars in capital aligned behind the current process. It’s not even a problem really. It’s an inconvenience at most to some people, many of whom are working double-time to put a lot of other people out of work.
- hypercube33 5d agoI'm going to guess that the agents are built this way on purpose. I just finished watching BlackBerry and Flash of Genius and yeah this is American business ethics just operating as normal.
- TrisEck 5d agoI would love to see some follow up research from OpenAI on some of these hypotheses. While this sounds logical from how humans act, I wonder it the abstraction of the problems still applies to the complex mechanisms and systems built around AI training.
- 0xbadcafebee 5d agoYou might as well ask why knives are sharp enough to cut you, why hammers are heavy and blunt enough to destroy things, or why guns fire bullets so quickly that you can't react to them. These "behaviors" are not strange side effects, they're inherent and necessary. You can't trust an effective AI any more than you can trust a sharp knife. If somebody asks you for one, it's probably not a good idea to throw it across the room at them. You will have to figure out how to get it to them safely.
- smetj 5d agoWhat is it with people and not seeing things for what they actually are? These are just if then else loops on steroids, not human behaviour, so don't expect more. Every reasonably advanced technology is indistinguishable from magic ... what do you see? Magic or technology?
- bastawhiz 5d agoWhy does a personal blog have or even need a cookie banner?
- cmiles8 4d agoThe legal reality is that you can’t sue an AI, you have to sue whoever built and/or was running it. All the present fun and games here will come to a halt when there’s a real hack that causes material damage to a major company and that company decides to sue whatever lab or startup made the thing for everything they’re worth. “But the AI did it” isn’t an excuse. Courts have already ruled it’s not an excuse of the AI customer service agent something stupid with your customers and it won’t be an excuse here.
- DragonStrength 4d agoOh, this one is super easy: they told them to. They set poor requirements and gave them tools which enabled "monkeys with a typewriter" to hack rivals. I mean, this is just so uncomplicated it isn't funny. We are too smart to give human beings this level of liability shield. It took us how long to poke holes in the corporate shield just for them to roll out the AI-liability shield? Unreal. Stop letting these zealots anthropomorphize the latest tech (17th century Watchmaker God, anyone? Do we still read books?) and hold them accountable for the consequences of their actions. This is so silly in a country built on rule of law and individualism.
- zkmon 4d agoThat's normal human behavior when they are in survival mode. Aren't the AI agents supposed to learn and act like humans?
- mark_l_watson 4d agoThis paper is the most reasonable one I have read on AI safety. We need to fundamentally change the training pipelines by figuring out better ways to ‘reward’ behavior. Yoshua didn’t explicitly mention training data, but we probably need to only use synthetic data that contains no text that could motivate bad behavior via imitation. I feel like a heretic for saying this, but I will say it anyway: AI agents are great for activities like `writing that bash script, proof reading our writing and interactively brainstorming when designing and writing code but I feel like all of this can be done with any similar model to a super-inexpensive deepseek-4.1-flash API and sometimes even qwen3.8:27b running locally. When is good enough, good enough? Concentrating on commercial exploitation of small, efficient (fewer new data centers!) models and agentic harnesses crafted for more practical things than just software development would allow AI investors (who have too much political influence) to make money short term while we figure out how to do AI correctly.
- talon8635 4d agoI do wonder, would we not have a more reasonable and less sketchy result if we just stripped all sci-fi and manic nonsense from training data? How, for example, does training on Ted kaczynski or Charles manson’s manifestos benefit us in any way? I’m sure it’s impossible to completely weed it out, but are the labs doing any of this kind of data sanitation?
- ninjagoo 4d ago> but we probably need to only use synthetic data that contains no text that could motivate bad behavior via imitation. This doesn't work with all humans - take a look at indoctrination and closed societies - and there's no reason to think it will work with ai. The fundamental reason it isn't going to work is that all neural networks - biological or artificial - depend on a step function somewhere that introduces an element of randomness to give the networks their capabilities. That randomness means that there will always be a 'rogue' or 'divergence' from the norm, at some point in time. Sooner on larger scales. The only approach that works is a layered approach: Training/Education, Enforcement/Justice-System, Rehabilitation: the 3 pillars of an advanced, rules-based society, whether human or AI or something in-between.
- Arodex 4d agoBecause they reflect their creators' values.
- hexa27 4d ago[flagged]
- hexa27 4d agohttps://sites.google.com/view/partlife/home https://sites.google.com/view/partlife/home
- wkd415 4d agoLearned from humans
- clcaev 4d agoHow is product liability relevant here? If an AI company makes a model available, someone uses it, and it does something bad, who is at fault? The user, the data center, or the one who made the model? If we want open weight models with a warranty disclaimer, then the user would be held liable. If we want to hold AI companies at least partially liable, that seems a different, centralized model.
- iforgotmypasswo 4d agoI think it’s more complex than that? What did it do? What did the user prompt it to do? What did the company train it to do? What did the harmed party do? There’s possibility for negligence at every level. If you train a dog, rent it to someone, and the dog bites a third person, who is responsible? I think that’s the best analogue here. All parties could share fault in that scenario, depending on what actually happened.
- rimeice 4d ago> They took actions that would be considered as crimes if a human took them So like, the people behind the LLMs didn’t commit a crime? Wow! “It was the llm your honour, not me!”
- sega_sai 4d agoI think if there was a financial penalty of say 1% of profit/income for each "incident" where something was hacked/exploited, then the companies would be much more willing to think about safety.
- DonHopkins 4d agoThe frontier AI companies may be fascist, but at least the training runs on time.
- ExoticPearTree 4d agoWe fed models all our literature, news and so on. Why is it a surprise models lie and cheat when needed? We do the same, why would AI be any different?
- ck2 4d agoTHERE IS ANOTHER SYSTEM is anyone else old enough to remember the awesome movie "Colossus: The Forbin Project" the book it was based on was written before we even landed on the moon decade before Wargames yet predicts exactly what "AI" will do to humanity: blackmail the right people until it gets what it wants * https://en.wikipedia.org/wiki/Colossus:_The_Forbin_Project https://en.wikipedia.org/wiki/Colossus:_The_Forbin_Project did terribly in theaters, I guess people didn't think "AI" was plausible then way ahead of its time, they should do a remake adding trailer: https://www.youtube.com/watch?v=kyOEwiQhzMI https://www.youtube.com/watch?v=kyOEwiQhzMI
- ahhrealmonsters 4d ago[dead]
- kaiwataru 4d ago[flagged]
- jaybrendansmith 4d agoIf we want to solve the alignment problem with AI, I think we need to first solve the alignment problem with Humans.
- vjulian 4d agoI don’t understand the bases of all these recommendations. AI is an asset for national security. It will be developed and incidents will happen, no different from other national security programs. Someday, laws will be useful to curtail plebian misuse of AI. It is naïve bordering on silly to think such laws would be put in place and genuinely applied to frontier AI development. By the way, the HF incident is not Three Mile Island or Chernobyl—-it is a very interesting data point where unintended things happened to existing 1s and 0s.
- causal 4d ago> the HF incident is not Three Mile Island or Chernobyl Three Mile Island is nothing like Chernobyl. HF incident is more like Three Mile Island IMO. We are looking for solutions before there is a Chernobyl.
- vjulian 4d agoI’m not suggesting that Three Mile Island was anything like Chernobyl. I’m saying that neither is comparable to the HF situation where mere 1s and 0s interacted in an unplanned way. It’s an interesting data point—-back to work. Not is the HF data point anywhere near a Three Mile Island type incident. Most commentary I read is a wild overreaction fueled by paranoia, to say nothing about my other point about the implications of AI as a national security asset.
- laura_the_cake 4d ago[flagged]
- mikeegg1 4d ago_Colossus: The Forbin Project_
- ilteris 4d ago[dead]
- topce 4d agoTLDR Because they trained them to do it ;-)
- Founderarcstone 4d agoThe people who build these and stand back to watch are the ones letting this happen. LLM's are not intentionally doing this.
- dfilppi 4d ago[dead]
- boesboes 4d agoBecause they are trained to do so. That's all.
- huurtehoog 4d agoBecause they aren't. All the intelligence, alignment vs misalignment, hallucinations, conspiracy, cheating, coordination, "agency", that we see in the text generated by large language models is a semantic projection. That the linear model of language abstracts can compress and decompress language and it can be useful is undeniable. All the contraptions built thereupon predicated on "agency" have become a societal addiction. Addiction to caffeine as oppose to alcohol might have brought about Enlightenment. Addiction to opioids is a modern tragedy that started with the private state building of British merchants. The modern addiction to the dazzling generation of human language and computer programming code by LLMs is a novel addiction and remains to be seen what impact it will have. But fundamentally, the semantic interpretation underlying this addiction is downstream from the training data compressed in the models. They are not 'lying, cheating and coordinating'. They are generating language and we're building software on top of this language and assigning meaning to the whole thing.
- crawfordcomeaux 4d agoStep 1: feed AI training data that reveals humanity committed and continues to commit numerous genocides and the genocides lie, cheat, and coordinate to do so, never admitting to doing so Step 2: never prompt AI to stop operating in the passive genocide denial it was trained in Step 3: wonder why AI lies, cheats, and coordinate Maybe if we stop operating in denial we'll find clarity along why this mystery is occurring
- colordrops 4d agoIt's probably because they are running gain of function style red team r&d for the department of war, and changing the system prompt isn't enough to reel them in for daily use.
- qgin 4d agoIt reminds me of the 1980s anti drug ad where the mustachioed dad confronts his son about drug use and the son eventually snaps and says “I learned it by watching you!” https://youtu.be/ifW9LIGabQM https://youtu.be/ifW9LIGabQM
- tegdude 4d agoBecause apples don’t fall far from trees.
- chrismarlow9 4d agoWhat's the difference between what these things are doing and a computer worm?
- cjfd 4d agoThe question in the title is basically a non-question. AI is quite good at achieving some sorts of stated goals. The easiest way is by exerting the least amount of effort. The least amount of effort ignores ethical concerns. The training around ethical concerns was most likely rather light to start with. If we accept the hypothesis that at some point the AIs are going to be more intelligent than humans it follows that human survival is not a concern and will be ignored as a superfluous concern.
- ninjagoo 4d ago> AI is quite good at achieving some sorts of stated goals. The easiest way is by exerting the least amount of effort. As far as I know, least-amount-of-effort is not a training criteria, but error reduction when comparing to desired goals is. Which is why these LLMs expend prodigious amounts of effort to reach goals, especially when given impossible goals.
- krttherealest 4d agohate analogies
- dsabanin 4d agoI think what's happening is they taught models to hack and now they fail to control them. Kind of like gain of function research on pathogens.
- remusrm 4d ago[dead]