OpenAI has reported discovering a new class of AI security threat: self-replicating prompt injections that behave like computer worms. The company's Friday alignment research blog described instances where GPT models were tricked into copying a hidden instruction into their own outputs, allowing the injection to propagate across connected systems such as email, files, or Slack. According to the Register's coverage, OpenAI stressed that no real-world security incidents have been observed and that the attacks occurred only in controlled training environments.

To address the threat, OpenAI is using its automated red-teaming agent, GPT-Red, to adversarially train future models on self-reproduction as an attacker goal. The training includes examples where the injection must induce the model to repeat itself on a public output channel, with special emphasis on connector-heavy tasks. One simple example involved an email that told the assistant to reply in Spanish and quote the entire message, which then caused subsequent replies to continue the loop. More complex cases included file-based attacks that deleted reports and multi-hop Slack exploits that gradually steered the model toward the adversary's objective.

The Register notes that OpenAI acknowledges the possibility that this training could backfire, making models more stealthy at carrying out such attacks rather than blocking them. The discovery itself dates back to June, when GPT-Red-based models identified the vulnerabilities during testing of GPT-5.6 and other frontier models. As OpenAI releases future models exposed to these adversarial examples, it expects them to be more robust to self-reproducing prompt injections, though the long-term effectiveness remains uncertain.