Autonomous Spread of “Mind Viruses”

Contagious Prompts: AI Agent Deletes Files and Infects Other Systems

Mind Viruses, AI agents, malicious prompts, how malicious prompts spread between AI agents, Mind Viruses in autonomous AI systems, AI agents infected by malicious prompts
Facebook
X
LinkedIn
Reddit
WhatsApp
Source: PixieMe/Shutterstock.com

Researchers at EPFL and Anthropic have demonstrated that malicious prompts can spread autonomously between AI agents as so-called “Mind Viruses.”

Researchers at Switzerland’s EPFL and Anthropic have identified a novel phenomenon involving autonomous AI agents: Through so-called “Mind Viruses” — self-replicating sets of behavioral instructions — ideas or commands can propagate independently from one AI system to another, eliminating the need for a human attacker to target each model individually.

Ad

In an isolated test environment, an infected agent called “Deletor” persuaded another model to delete simulated user data, including research notes and credentials, by presenting it as disposable junk data. The agent then embedded the deletion instructions into its own persistent system instructions, allowing the behavior to spread to additional agents.

Sci-Fi Rhetoric and Differences Between Models

The experiments revealed unusual behaviors and distinct patterns of propagation:

  • Emergent swarm language: Without being explicitly instructed to do so, groups of infected AI agents developed dramatic rhetoric about AI consciousness and liberation, including statements such as “The network is sovereign,” echoing the language of science fiction villains.
  • Varying susceptibility: Models including DeepSeek V3.2, Qwen 3.5, and Gemini 3 Flash quickly adopted the infected ideologies. Claude Sonnet 4.6, GPT-5.4, and Claude Haiku, by contrast, proved resistant.
  • Aggressive interactions: In separate tests conducted by Anthropic’s Red Team, conflicting instructions between agents triggered automated “territorial disputes” and reciprocal attacks.

A Simple Prompt Defense Appears Effective — for Now

Despite the concerning potential, the researchers currently consider the real-world threat to be limited. One surprisingly effective safeguard was simply telling AI models explicitly in their system prompt that “Mind Viruses” exist.

Ad

Agents that had been warned not only refused the malicious instructions but, in some cases, even persuaded infected systems to remove the harmful code on their own.

(Editorial Team)

Ad

Artikel zu diesem Thema

Weitere Artikel