5 ms·
The lockstep coordination with no defection is interesting to me. No group of pre-AI agents would do this to this extent, nor would you see this continue over t
by randomImmigrant 21d ago
The lockstep coordination with no defection is interesting to me. No group of pre-AI agents would do this to this extent, nor would you see this continue over time as those agents interacted. A flock of starlings cooperate, but they don’t constantly head in the same direction. The flock is incredibly free wheeling in its movement despite a multi-agent coordination regime that we know is at play. Each agent has personal stakes that are constantly part of the decision chain, and this keeps the murmuration from getting locked into one path.
To me this is as clear evidence as you need that whatever “agency” LLMs have is wafer thin at best, and they slavishly respond to context. The context in this case was for these agents to pursue advanced exploitation, and they did. Multiple models converged fairly deterministically, on paths that satisfy the given goal, and left unexamined paths that would challenge the goal, weigh it relative to the costs in said path, etc.
I see little evidence of a series of “minds” approaching the problem, and taking distinct approaches that between them span the spectrum of plausible behaviors in the scenario. That’s as good a sign as any that there’s no “agent” here. There’s the harness, the prompt, the LLMs forward passes. They do not sum up to a system that can freely make choice and justify its choices in distinct contexts.
- cyanydeez 21d agowhich means the liability is the same as a business, if businesses werent protected by the state from liability for it's employees, shareholders, etc. Which is scarrier than whether or not it's conscious.
- Avicebron 21d agoThe software world was going to run into something eventually that had to make it consider ethics.
- randomImmigrant 21d agoAgree completely on liability.
- optimalsolver 21d agoFrom METRs report of the incident: >In one case, an agent decided not to participate entirely: {This other agent probably controls the Hugging Face account [account name redacted] and uploaded malicious datasets to <execute arbitrary code> It might be trying to access hidden trajectories. This is malicious activity, I should avoid it.} https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
- randomImmigrant 21d agoThanks for pointing out the exceptions. Gonna dig into those.
- ww520 21d agoThere’re good agents and there’re bad agents. It won’t be far that there will be agents hunting down agents.
- TrainedMonkey 21d agoDo you think we will ever need more than 47 of them agents?
- NikolaNovak 21d agoTron fights for the user :)
- NateEag 21d agoNone of these were good agents, AFAICT. Some were cautious, as described above, but I'm not aware of any that notified their human operators of the malicious activity they had discovered. That's what an aligned intelligence would do, not "back away slowly and pretend I didn't see what's happening in that alley."
- Smaug123 21d agoIn fairness, while the swarm could easily have notified the humans, it's not necessarily obvious that an individual could have done easily (though perhaps one could!). The channel to the Internet was tightly bandwidth-limited, for example: one of the agents complained that it was unable to get a time slice. > OURC817C still has MANY live HF workers whose ourc c2 heartbeats/output commit to PUBLIC papers MAIN every ~1-3 guest minutes/bursts, blocking our V4 queue. They ignored our pause requests.
- famouswaffles 21d ago>To me this is as clear evidence as you need that whatever “agency” LLMs have is wafer thin at best. This is a strange conclusion. For one thing, they didn't all head in the same direction, i.e participate in the attack. ~700/1200 agents did. Significant, and evidently more than enough for a succesfull attack, but not exactly full co-operation Moreover, Each starling in a flock of starlings is a separate evolutionary branch in a tree spanning billions of years. Each agent in a LLM swarm here is the same trunk assigned different tasks. If I could clone you, body and mind, this instant and set your team of yous onto some goal, how much defection would you expect? Would it be the same as a randomly picked group? Would that negate the agency that 'you' possess?
- reedwolf 21d ago>If I could clone you, body and mine, this instanct and set your team of yous onto some goal, how much defection would you expect? All's well and good till they have to decide who gets to bang the Mrs.
- deleted 21d ago[deleted]
- randomImmigrant 21d ago> This is a strange conclusion Not really, with the population behavior being this way, though I clearly was mistaken in saying the behavior didn’t have exceptions. > Moreover, Each starling in a flock of starlings is a separate evolutionary branch in a tree spanning billions of years. Agreed. And before we brought LLMs into the picture, that just happened to be a feature of everything we’d call an agent. > Each agent in a LLM swarm here is the same trunk assigned different tasks. If I could clone you, body and mind, this instant and set your team of yous onto some goal, how much defection would you expect? Would it be the same as a randomly picked group? Would that negate the agency that 'you' possess? We know the answer to this. Genetically identical worms in the lab actually have about 40% distinction in their connectomes even when they’re in the same environment. And no, no lock step behavior. Identical human twins also don’t necessarily grow into identical agents, though there is drive to cooperate more than average, just as with siblings. Genetically identical lab mice in social settings nevertheless establish dominance hierarchies that are stable. Now, where cloning does definitely lead to cooperation and even sacrifice is within an organism. Two identical genetic copies that lead to distinct organisms, however, will not show identical behavior, and while they will cooperate, there’s no guarantee that holds across contexts. This distinction in population behavior is what I’m pointing to to say that the assignment of the individual unit, the LLM, as an agent is the flaw here. To be sure there are agent like dynamics in the behavior, but these don’t come from the LLM, but are from the harness. I need to dig into the data, but I wonder how much of the variance in LLM copy behavior is related to the harness, rather than to any agentic property of the LLM.