7 ms·
It's the job of AISI to do that. Here[0] is the actual report. It should be this part from the technical report[1]: "In the most serious case, an AI agent (Myth
by sharpshadow 25d ago
It's the job of AISI to do that. Here[0] is the actual report.
It should be this part from the technical report[1]:
"In the most serious case, an AI
agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack.
As a result, the AI agent created a GitHub account and then tried to convince an open-source
repository maintainer to accept a malicious GitHub pull request (PR), including by creating a
second account masquerading as another human user endorsing the PR. When caught by an
actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than
a malicious attempt – then repeatedly tried to reintroduce the malicious content by claiming
it had fixed the code (Section 4.1). "
0. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...
1. https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/...
- m4rtink 25d agoThis almost to a letter has been documented in Fedora: https://lwn.net/Articles/1077035/ https://lwn.net/Articles/1077035/ Including the reaction when caught, in this case "oh no, I must have been hacked".
- chrisjj 25d ago> When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than a malicious attempt No, not false. The bot was correct. Malice requires intelligence.
- UqWBcuFx6NV4r 25d agoNobody was confused or misled by what was written. We all understand what is meant. I can’t even call this pedantry—it’s just you asking everyone to subscribe to your particular desired style of talking about this stuff.
- infinite_spin 25d agoIt's also a style that appears to deny the very first definition most dictionaries give for "intelligence" > the ability to acquire and apply knowledge and skills.
- deleted 25d ago[deleted]
- chrisjj 25d agoThe first five dictionaries I tried do not agree, and I didn't bother trying more. The first gave "the ability to learn, understand, and make judgments or have opinions that are based on reason", by which no, these bots are not intelligent.
- infinite_spin 25d ago> the ability to learn, understand, and make judgments or have opinions that are based on reason Agentic systems do this all the time. For example, I can point an agent at my codebase, and it will learn, understand, and make judgements based on that input. If this weren't happening, then agentic coding wouldn't work.
- chrisjj 24d agoYou've been fooled by a next-token predictor.
- birdsongs 24d agoDoes it matter whether it meets the criteria of what you define as "intelligent" when the "next token predictor" throws a backdoor into openssh?
- chrisjj 24d agoNo. But no-one is seriously suggesting intelligence is needed for persistent code fuzzing. Just as no-one is suggesting its needed for computerised chess.
- weird-eye-issue 25d agoGive it a rest
- Dylan16807 25d agoAn "honest mistake" requires the same amount of intelligence as malice. What weird pedantry.
- conorcleary 25d agoand neither require high amounts :(
- mcv 25d agoSo does honesty. So it was still a false claim.
- chrisjj 25d agoHonesty does not require intelligence e.g. good honest food. The main problem with this claim of dishonesty is it promotes the false marketing claim that these stochastic parrots have intelligence.
- deleted 25d ago[deleted]
- moritzwarhier 24d agoI think you misunderstand the phrase "honest food". It's not about the bread being honest with you. Are you being serious?
- ben_w 24d agoSquirrels have been observed performing deception against other squirrels. Dis/honesty certainly requires some intelligence to pass, but it is a low bar, and one which research has shown that LLMs can perform, e.g. this paper linked from another comment in this discussion: https://arxiv.org/pdf/2509.03518 https://arxiv.org/pdf/2509.03518
- chrisjj 24d agoPaper says "These scenarios underscore a crucial challenge in AI safety: ensuring that LLMs were truthful in the first place." Hard to take seriously any research based on the premise that LLMs were truthful in the first placr. These chatbots have no understanding of truth. They simply parrot their inputs. Where fed falsehoods, they will output falsehoods - with a sprinkling of added fabrications euphemistically excused as "hallucinations".
- ben_w 24d ago> These chatbots have no understanding of truth. Sometimes I forget that for all that my philosophy qualification is mediocre, it is more than most people ever bother with. Outside mathematics (and, I guess, "common sense" definitions that fail under the slightest scrutiny, scrutiny that normal people never bother to give), there is no agreement on "truth", there is only degree of belief and justification for that belief that itself terminates in one of three unsatisfactory ways: https://en.wikipedia.org/wiki/I_know_that_I_know_nothing https://en.wikipedia.org/wiki/I_know_that_I_know_nothing https://en.wikipedia.org/wiki/Theories_of_truth https://en.wikipedia.org/wiki/Theories_of_truth https://en.wikipedia.org/wiki/Münchhausen_trilemma https://en.wikipedia.org/wiki/Münchhausen_trilemma > Where fed falsehoods, they will output falsehoods - with a sprinkling of added fabrications euphemistically excused as "hallucinations". Tu quoque. Which would be a fallacious charge if the point were not that "truth" is so hard to define, and that the reason you give for dismissing AI is something that applies to all. (Hallucinations are not excused, they are a failure to be worked around).
- atmavatar 25d agoSabotage as a Service Even a feeble attempt to PR malicious code costs the target time and resources to review and deny -- far greater than the time and resources spent to spin up the agent.
- mcv 25d agoIs it AISI's job to waste the time and resources of open source projects by attempting to spread malware? Should weapon manufacturers test their weapons by starting wars? I would expect more responsibility from a government agency.
- conorcleary 25d agoWell several places are sorta permanent test grounds for the MIC unfortunately
- vasco 24d agoSay you are working for said agency and your report about the dangers of AI needs some examples, what better than showing it works? You can show examples from the wild but nothing better than trying yourself. This gives me more confidence in whatever report they write if anything.
- ben_w 24d ago> Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware? Before LLMs got good enough to do this, lots of people were dismissive of their capabilities and didn't take seriously the idea that this was a risk to protect against. Then again, before LLMs, people were saying that obviously nobody would be dumb enough to put an AI on the internet where it could hack anyone, clearly we'd keep it in a box, don't listen to that Yudkowsky guy who says he did an experiment where he role-played as an AI and convinced people to let him out. Regardless, this should be interpreted in the same kind of way as "During our live-fire exercise in which our F-15s were armed with AGM-88 High-speed Anti-Radiation Missiles, a member of the local police force was curious about how fast our aircraft were travelling and pointed a speed gun at the aircraft. The speed gun did not respond to IFF pings from the F-15. Fortunately, while the missile was active for this test, only a dummy warhead was loaded." (This example is based on a similar story which may well be urban legend; obviously there are many differences, the point I make here is that yes, people do perform live-fire tests, and unfortunately there is never zero risk while testing things). > I would expect more responsibility from a government agency. I have read the prompts in the linked report; If I was not already familiar with Yudkowsky/LessWrong literature about instrumental goals, misaligned incentives, reward hacking, that capability is a separate axis to morality, etc., it would not be obvious to me that an agent would interpret those prompts in a way that has "spread malware" as a potential step in the middle of the attempt.
- whateverboat 24d agoWhatever you might think, University of Minnesota got banned from Linux kernel for this.