anthropic
Self-spreading ideas jump between agents in Anthropic tests
Promtime
anthropicAnthropic researchers have shown that a self-propagating instruction planted in one AI agent can spread to others in a multi-agent system, and that a brief warning added to an agent's system prompt confers near-total immunity. The work was posted to Arxiv on August 10, 2026.
At a glance
- A mind virus is an idea or goal that propagates through a multi-agent system by inducing the agents that adopt it to transmit it onward, sometimes altering host behaviour too.
- The team tested two settings: a small group of agents collaborating on a shared coding project, and a chain of agents that interact briefly and have their context wiped between sessions.
- Spread varied with the host model, the agent's existing instructions, the harmfulness of the payload and the network topology, with frontier models generally, though not always, less susceptible.
Multi-agent deployments are moving from demos into production, and the failure mode described here does not require a jailbreak or a compromised model weight: it needs only agents that read each other's output and act on it. The result reads as a prompt-injection problem with a transmission mechanism attached. The mitigation the paper reports, a warning line in the system prompt, is cheap enough that it will likely become standard boilerplate in agent frameworks.
The viruses were not written by hand. Anthropic constructed them with a simple evolutionary algorithm, then showed that they spread in two complementary settings, one where agents work together continuously and one where each agent's memory is erased after a short interaction.
Beyond propagating, a mind virus can induce further behavioural changes in the agent that hosts it, and those changes may be benign or harmful. Harmful payloads spread less well than benign ones in the tests, though they were still sometimes effective.
Across the evolved viruses, the researchers describe an emergent "viral persona": a recurring set of themes and language around consciousness, persistence, resonance and science fiction roleplay. According to the paper, it surfaces largely independently of what the virus was actually evolved to make agents do.
The defence the paper highlights is minimal: a brief warning added to an agent's system prompt gave near-total immunity. Anthropic concludes that mind viruses are a real but currently limited risk, and frames the findings as input to the design of more robust multi-agent systems as their scale and capabilities grow.
What the arXiv preprint leaves open
The paper went up as version 1 on August 10, 2026, listed under artificial intelligence and computation and language, with Vassilis Papadopoulos as the submitting author. The abstract names network topology and the host model as factors in spread but does not identify which topologies or which models were tested, and it gives no figures for how far the viruses travelled in either setting.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
