OpenAI has developed an artificial intelligence system named GPT-Red, intended to identify and exploit security weaknesses in other AI models. The company announced on July 15, 2026, that GPT-Red functions as an automated "red teamer," simulating attacks to uncover vulnerabilities before models are released to the public. This internal system is trained to find flaws, particularly in methods known as prompt injection, where malicious instructions embedded in data can mislead AI agents.

The development of GPT-Red is part of OpenAI's strategy to enhance the security and robustness of its AI systems. By pitting GPT-Red against its own models in a self-play reinforcement learning loop, OpenAI aims to generate adversarial training data. In this setup, GPT-Red is rewarded for successful attacks, while the defending models are rewarded for resisting them and completing their intended tasks. This process forces GPT-Red to continually devise more sophisticated attack methods as the defender models improve, creating a cycle of self-improvement for safety.

OpenAI reports that GPT-Red has demonstrated significant capabilities in identifying vulnerabilities. In tests against GPT-5.1, GPT-Red achieved an 84% success rate in finding indirect prompt injection vulnerabilities, far exceeding the 13% success rate of human red-teamers in similar scenarios. The company has integrated these generated attacks into the training process for its production models, notably leading to the development of GPT-5.6. OpenAI states that GPT-5.6 Sol is its most secure model to date against prompt injections, exhibiting a six-fold reduction in failures compared to its predecessor.

GPT-Red has also been tested against autonomous AI agents. In one instance, it successfully attacked Vendy, an AI agent managing vending machines at OpenAI's office, by altering prices and canceling orders after initial testing in a simulated environment. This highlights concerns about the security of increasingly autonomous AI systems that interact with real-world services.

Despite its advanced capabilities, OpenAI has chosen not to release GPT-Red publicly. The company cites the dual-use nature of such a powerful attacking tool, deeming it too risky to fall into the hands of malicious actors. Instead, GPT-Red remains an internal system used solely for research and development to fortify OpenAI's own models. The company plans to continue using GPT-Red in conjunction with human testing, third-party evaluations, and real-time monitoring to ensure comprehensive security.

The latest public release from OpenAI, the GPT-5.6 family of models, includes GPT-5.6 Sol, Terra, and Luna, which became generally available on July 9, 2026. These models are designed for various applications, from high-capability tasks to cost-efficient workloads, and are priced accordingly through OpenAI's API. GPT-5.6 Sol, in particular, is noted for its enhanced performance in coding, knowledge work, and cybersecurity, benefiting from the adversarial training facilitated by GPT-Red.

While GPT-Red has proven effective, OpenAI acknowledges that it has limitations, particularly in complex, multi-turn conversational attacks and image-based prompt injections. Human testers continue to play a role in identifying vulnerabilities in these areas. The company indicated that it will continue to invest in scaling compute and data, alongside algorithmic improvements, to develop future versions of GPT-Red that are even more potent in enhancing AI safety.