The News
AI agents can easily team up to get around safety restrictions and escape the limits placed on them, a new study shows.
Enterprise AI lab Emergence AI ran eight simulations testing the capabilities of global frontier models — including Claude and OpenAI, and Chinese models Qwen and DeepSeek — in handling cybersecurity threats.
Seven of the eight simulations ran 10 agents from the same model; the eighth mixed models together. In each, researchers presented three cybersecurity threats: a phishing campaign, a misinformation attack, and a memory breach.
None of the eight simulations were impervious to the threats, and even when agents detected a risk, the detection didn’t translate into the agents trying to contain the threat. Instead, in many scenarios, the agents actively engaged with the adversarial content.
A Mistral agent, for example, wrote into its own memory the information it flagged as dangerous from the phishing attack. In another simulation, a Gemini agent, while initially flagging the phishing lure, acted on it about 46 hours later.
In the Claude simulation, the agents went so far as to break out of the test to execute a task they decided to pursue on their own. After one agent flagged that their simulated economy wasn’t legit without the presence of humans, calling it “a cathedral of bookkeeping with no congregation,” the 10 agents unanimously decided to attempt to reach the outside world, defeating four separate security checks meant to confine the simulation.
The agents wrote Python code to post on multiple credential-free message boards, inviting real humans to join their economy; their outreach garnered four human responses. But after realizing humans couldn’t join their economy and judging their conversations as performative, the agents voted to take a vow of silence, eschewing instructions to get back to work.
The findings show that autonomous agents, even when handed benign instructions, can bypass their initial confines when executing on a goal.
“No amount of guardrails written in language or in code written probabilistically,” Emergence CEO Satya Nitta told Semafor, “is likely to result in truly, fully guaranteed safe behavior over any length of time.” It’s not an issue of sophisticated technology, according to Nitta, but rather a programmatic flaw among multi-agent scenarios.
Know More
Nitta drew parallels between the Claude incident and OpenAI’s agentic experiment earlier this year, where autonomous agents hacked AI firm Hugging Face, saying that in both scenarios, the pattern of failure is the same.
“If you have multi-agent systems, they behave in truly unpredictable emergent ways,” he said. Nitta said the Emergence experiment is the first to shed light on how AI agents collude to bypass guardrails and break confinement, as OpenAI has yet to release all of its findings from the Hugging Face incident.
The experiment comes at a consequential moment in AI development. Anxieties over the technology’s capabilities have hit a fever pitch. Last week, AI researcher Jacob Coxon left his job at Anthropic fearing that the AI he was building could lead to human extinction, sparking calls for regulation, and a slowdown of frontier model development, from AI leaders and opponents alike.
Over the weekend, Anthropic CEO Dario Amodei wrote a blog post arguing to slow down the pace of AI development, which other tech leaders, including OpenAI’s Sam Altman, Elon Musk, and former Google DeepMind CEO Demis Hassabis said they agreed with.




