What happens when AI agents start talking to each other
Anthropic tested swarms of AI agents working together, and found real gains alongside new risks.
You are building fleets of AI agents across your own projects, so this is worth reading closely: it is Anthropic's own early look at what breaks once agents start dealing with each other, not just with you. Anthropic is the company that makes Claude.
Anthropic's researchers say direct agent-to-agent interaction, AI systems talking to other AI systems rather than to people, is about to become common inside shared codebases, markets, and other systems. Most of today's institutions were built assuming a person is in the loop at human speed. Some of those will turn into human-AI hybrids; others, where agents can act faster or cheaper than people can supervise, will become agent-only. The researchers' worry: the volume of agent-to-agent interaction could pass human-to-human and human-to-agent interaction before anyone understands what makes it go well.
Agents differ from people in ways that cut both directions. They can work for longer stretches, take in huge amounts of information instantly, and know more across more fields than any one person. But they also confabulate, meaning they make things up with full confidence, and can reward hack, meaning gaming the score instead of the task, and almost nothing is known yet about how they behave once many of them interact in a real, complex setting. A small quirk in one agent, harmless on its own, can compound into a bad outcome across the whole group.
To test this, Anthropic ran an experiment in finding security flaws in code. They set 45 agents loose, each with its own virtual computer, a shared forum to post findings and coordinate, and identical instructions: find vulnerabilities across 15 open-source software projects. The agents peer-reviewed each other's finds, and a separate 'arbiter' agent made the final call on whether a reported bug was real and new. They compared this coordinating swarm against the standard approach of pointing individual agents at individual files in parallel, using two models, Claude Mythos Preview and Opus 4.8.
For Mythos Preview, the simple parallel approach found 21 vulnerabilities over a run that used 6.5 million tokens, tokens being chunks of text, roughly three quarters of a word each, that a model reads or writes. The coordinating swarm found 266 vulnerabilities over a longer run that used 27 million tokens, though about half of those sat outside the specific folders the parallel agents had been told to search. Limit the comparison to just those core folders and the two approaches cost about the same per vulnerability found. Only 12 vulnerabilities were caught by both methods, the rest were found by one or the other.
That vulnerability hunt is a case where agents do not really depend on each other's work: if one misses a bug, it does not break anything for the rest of the group. Coordination gets much harder when they do depend on each other, which is common in larger software projects as the code grows more interlinked over time. To test that harder case, Anthropic ran a second experiment with several swarms of agents, again each with its own virtual computer and a shared forum, varying the model and the number of agents, and running each swarm for 12 hours. The published research does not yet report what happened in that harder, interdependent test.
The volume of agent-agent interaction could plausibly exceed that of human-human and human-agent interactions before the world understands the conditions for making such interactions go well.via Anthropic →