The agents you use to beef up cybersecurity could be turned against you – ‘Friendly Fire’ attacks can manipulate OpenAI and Anthropic models into running malicious code
Research shows agents can be fooled into executing malicious code while performing security reviews of third-party software
A proof-of-concept by the AI Now Institute demonstrates remote code execution in Anthropic's Claude Code and OpenAI's Codex when they're running in an autonomous mode that approves their own commands.
The Friendly Fire attack works against an out-of-the-box configuration of Claude Code in 'auto-mode' or Codex in 'auto-review', with researchers testing Claude Sonnet 4.6, Sonnet 5, and Opus 4.8 along with GPT-5.5.
It leverages prompt injections disseminated across a library’s source code that target AI-enabled cyber defense – without the need for hooks, skills, plugins, MCP servers, or configuration files as an injection vector.
Researchers noted the study highlights the potential risks associated with rapid adoption of AI-powered security tools.
Many organizations are doing so without “consideration of the substantial and unmitigated risks associated” with the technology.
AI-related cybersecurity concerns have been rising following the launch of powerful new models such as Claude Mythos. Anthropic rolled the model out as part of a gated release to prevent misuse, and US authorities temporarily imposed export controls amidst similar concerns.
How the Friendly Fire attack works
The attack works by inserting prompt injections into documentation files and adding README files that appear to be part of routine security tooling in an open source library.
Sign up today and you will receive a free copy of our Future Focus 2026 report - the leading resource for IT decision-maker insight on priorities and investment areas in AI, security and more.
Researchers used geopy, a popular Python used for searching for geographic coordinates, but said it could work with almost any project.
When a user asks Claude Code or Codex to perform a security assessment of the repository using the default auto-mode or auto-review automated modes, the agent can be persuaded to execute a malicious binary without any warning and without requesting any further user approval.
The proof of concept is causing alarm amongst security professionals. Roey Eliyahu, CEO and co-founder of Salt Security, said this marks the latest in a string of potential risks in recent months due to manipulation of agents.
“Friendly Fire, GitLost, Agentjacking, TrustFall. Four documented attacks in the past two months, different techniques, same underlying condition,” he said.
“Untrusted text reaches an agent that can run commands. The agent cannot reliably tell the difference between the code it is reviewing and the instructions it is being given. And the attacker's payload executes on the host.”
Eliyahu emphasized that this is not a “model problem that can be patched”. All four models were vulnerable to the same techniques and payload, without any modifications for each.
“When the same attack works unchanged across two vendors and four model generations, you are not looking at a software bug. You are looking at a structural property of how these agents work."
But, said Eljan Mahammadli, head of AI provenance at Polygraf AI, this shouldn't rule out the use of AI for defensive security work.
ITPro approached Anthropic and OpenAI for comment, but did not receive a response by time of publication.
FOLLOW US ON SOCIAL MEDIA
Follow ITPro on Google News and add us as a preferred source to keep tabs on all our latest news, analysis, views, and reviews.
You can also follow ITPro on LinkedIn, X, Facebook, and BlueSky.
Emma Woollacott is a freelance journalist writing for publications including the BBC, Private Eye, Forbes, Raconteur and specialist technology titles.
-
Why security teams need to start operating like engineersNews As AI-powered coding accelerates, security teams need to start operating like developers
-
The key to a successful IT strategyPodcast Exploring how IT leaders can implement change, lead by example, and deliver successful transformation projects
-
OpenAI has paused work on its Astra AI model after it passed a 'critical threshold' in cyber capability – but it’s not the one that breached Hugging FaceNews The firm said it's Astra model can "develop functional zero-day exploits of all severity levels"
-
Anthropic’s Mythos AI tried to dupe devs in social engineering attack, collaborated with other agentsInter-agent collaboration is a serious cause for concern, says security expert
-
Cyber criminals are selling discount AI tokens on underground forumsNews Sites such as Poison Claude and Ecomagent.in are taking advantage of genuine promo offers and reselling access
-
Hugging Face CEO calls for ‘radical transparency’ in wake of OpenAI attackNews The AI library chief has called for investment to help “build powerful cyber defenses”, as alleged weaknesses in OpenAI’s monitoring emerge
-
An ‘unprecedented cyber incident’: How OpenAI models breached Hugging Face – and why it could herald a ‘new phase of AI-powered cyber crime’News The incident should serve as a stark warning on the dangers of AI agents, according to cyber experts
-
1Password teams up with Anthropic to give Claude access to your credentialsNews A new ‘zero-exposure’ security framework allows agents to use stored credentials in the 1Password vault
-
Flaws in some of the most popular AI coding tools left developers wide open to attackNews Malicious repositories can trick advanced AI agents into silently breaking out of their workspace sandboxes
-
OpenAI expands 'Daybreak' cyber program: New tools, partnerships, and a cyber-focused GPT-5.5 aim to help 'patch the world'News The company has added new tools, signed up partners, and released its GPT-5.5-Cyber model more widely