
OpenAI has developed GPT-Red, an internal AI model that automatically detects security flaws in GPT systems by simulating attacks like prompt injections—and it succeeds at finding exploitable vulnerabilities 84 percent of the time, far exceeding the 13 percent success rate of human red teamers.
The findings are feeding directly into model training: GPT-5.6 Sol now shows six times fewer failures on direct prompt injection attacks than models from four months earlier, though about 3.8 percent of stronger attacks still succeed, meaning determined adversaries could still penetrate the system at scale.
What happened
OpenAI built an internal AI called GPT-Red that automatically searches for security vulnerabilities in GPT models by simulating prompt injections and other attacks. Trained via self-play reinforcement learning, GPT-Red succeeds in finding exploitable flaws in 84 percent of test scenarios, compared with 13 percent for human red teamers. In one test, it manipulated an AI-powered vending machine in OpenAI's office to change prices and cancel other customers' orders.
Why it matters
The findings directly feed into training the next generation of models. GPT-5.6 Sol shows six times fewer failures on direct prompt injections than the best model from four months ago, without hurting general performance. This suggests AI-driven security testing is substantially more effective than human effort at hardening models against adversarial attacks, a critical concern as these systems handle sensitive tasks.
What to watch
About 3.8 percent of "stronger" prompt injections still succeed on GPT-5.6 Sol, meaning at scale—across hundreds or thousands of attempts—a sizable number of attacks get through, similar to Claude Opus 4.5. GPT-Red remains internal; OpenAI says a detailed paper will follow.
Ask the AI about this article →
OpenAI's approach to AI security marks a shift from traditional human-led red teaming to adversarial machine learning. By training GPT-Red via self-play reinforcement learning—where the attacking model and defending models improve against each other iteratively—OpenAI has created a process that scales far beyond what human testers can achieve. The 84 percent attack success rate for GPT-Red versus 13 percent for humans reveals a substantial efficiency gap; human red teamers, while valuable, cannot explore the adversarial space as exhaustively or rapidly.
The concrete improvements in GPT-5.6 Sol demonstrate that this testing regime translates into measurable security gains in production models. Six times fewer failures on direct prompt injections represents a substantial hardening, achieved without degrading the model's general capabilities—a critical constraint, since overfitting to defense can cripple utility. However, the persistence of 3.8 percent success rates on stronger attacks indicates that prompt injection remains a partially open problem even at the frontier. The comparison to Claude Opus 4.5 suggests this residual vulnerability is not unique to OpenAI's approach but reflects a broader challenge across leading models.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Visko raised $10 million in pre-seed funding from Llama Ventures and opened public access to its first foundat…
U.S. markets ended August higher, with the S&P 500 up 2.6% and the Nasdaq up 3.9%

Neurovia AI, an Abu Dhabi-based company, is pitching Saudi security agencies software that it says can compres…

AI company Runway has unveiled Solaris, the first model in a new category it calls "Interface World Models." I…

Google's AI search gave advice to call emergency services for users alone with an African, Indian, or Pakistan…

John Deere introduced JD, a conversational AI tool that lets farmers ask open-ended questions about their hist…
