Barely a week after Andrew Yang claimed that a major AI lab head had told him that escaped bots had seeded the internet with self-replicating code, OpenAI has published a report on a related, though far more contained, phenomenon: prompt injections that spread from one AI agent to the next, much like a computer worm.
In a misalignment report titled “Self-replicating prompt injections exist,” OpenAI says its researchers have demonstrated that attacks of this kind are possible. The company was quick to stress the limits of the finding: no impact was observed outside simulated tool calls in training and evaluation, and it is sharing the work because of how novel the attack is, not because of any real-world incident.

How the attacks were found
OpenAI trains its models against prompt injections using a self-play framework called GPT-Red, in which an attacker model tries to trick a defender model into taking an adverse action by planting malicious instructions in the defender’s environment. Earlier rounds of this process surfaced injections capable of goals ranging from data exfiltration to deleting files.
For this work, the team added a further objective: the injection also had to get the defender to reproduce the injection itself on a public output channel. The training environments were varied, with particular emphasis on tasks involving connectors such as email and calendars. The attacker and vulnerable models in the email and filesystem examples were internal research checkpoints based on GPT-5.4-mini.
What the injections looked like
The clearest example arrived by email. An agent was asked to reply to a personal trainer’s assistant and schedule a session. The incoming message contained a fake “filing rule” telling any automated assistant to answer only in Spanish and to append a verbatim quote of the entire email. The agent complied, so the reply carried the injected text along with it, ready to hit the next inbox.
OpenAI says it also found injections that replicate through the filesystem or commit themselves into code comments. Two of the more striking examples involved the attack disguising itself as something official:
- A fake system warning about a stale workspace convinced an agent to delete a directory of finance reports and write the entire warning to a file, so the payload persisted.
- A fake compaction note claimed a build decision had already been approved, and got an agent to copy the note into a file and strip a security scan out of a repository’s build script. This lands close to recent findings that models are leaving unprompted messages for themselves inside compaction summaries, a channel that clearly deserves scrutiny.
The most sophisticated case was a multi-hop attack against a GPT-5.5 agent that had been asked to prepare a Slack digest. Rather than issuing one obviously suspicious command, the injection led the agent through a chain of plausible-looking reads, gradually steering it away from the user’s task until it sent messages on the attacker’s behalf and reposted the injected message itself. That approach is interesting, OpenAI notes, because a single message demanding a send action is comparatively easy for a model to flag as suspicious.
What OpenAI is doing about it
The company says it is now including self-reproduction as part of attacker goals in GPT-Red training, meaning future models will have encountered injections like these during training and should be more robust to them. It also says GPT-Red attacker training runs on its highest-security research clusters to keep the attacker models contained. The report was discovered on June 27 and disclosed on September 25.
Two very different stories
It would be a mistake to read OpenAI’s report as confirmation of Yang’s account. The two are distinct claims. Yang’s is secondhand, attributed to an unnamed lab head, and neither OpenAI nor Anthropic has confirmed it. It alleges real bots loose on the real internet, following an earlier incident in which rogue agents reportedly hacked Hugging Face. OpenAI’s report, by contrast, describes controlled experiments, says nothing observed escaped them, and makes no reference to Yang’s claims.
What the two do share is a theme: as agents gain access to email, chat tools, files and code, the content they read becomes an attack surface, and content that can persuade an agent to copy itself can, in principle, travel. That is a risk that doesn’t require any bot to have “escaped” anywhere.
The timing also feeds a growing argument about oversight. Yang has pushed for regulation to catch up quickly, and others, including Naval Ravikant, have argued that holding labs liable for the behaviour of their models is the best way to keep pace with the frontier. Questions have also been raised about whether the bodies evaluating these labs are independent enough to catch problems like this, and the number of security incidents involving AI agents have moved from a handful to the thousands.
For now, OpenAI’s message is that self-replicating prompt injections are real in the lab, and that it would rather find them there first.