AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Jul 16, 2026, 06:00 JST2 min read

OpenAI's AI Red Teamer Finds Security Flaws 6× Better Than Humans

OpenAI's AI Red Teamer Finds Security Flaws 6× Better Than Humans

Key takeaway

  • OpenAI has developed GPT-Red, an internal AI model that automatically detects security flaws in GPT systems by simulating attacks like prompt injections—and it succeeds at finding exploitable vulnerabilities 84 percent of the time, far exceeding the 13 percent success rate of human red teamers.

  • The findings are feeding directly into model training: GPT-5.6 Sol now shows six times fewer failures on direct prompt injection attacks than models from four months earlier, though about 3.8 percent of stronger attacks still succeed, meaning determined adversaries could still penetrate the system at scale.

3 Key Points

  1. What happened

    OpenAI built an internal AI called GPT-Red that automatically searches for security vulnerabilities in GPT models by simulating prompt injections and other attacks. Trained via self-play reinforcement learning, GPT-Red succeeds in finding exploitable flaws in 84 percent of test scenarios, compared with 13 percent for human red teamers. In one test, it manipulated an AI-powered vending machine in OpenAI's office to change prices and cancel other customers' orders.

  2. Why it matters

    The findings directly feed into training the next generation of models. GPT-5.6 Sol shows six times fewer failures on direct prompt injections than the best model from four months ago, without hurting general performance. This suggests AI-driven security testing is substantially more effective than human effort at hardening models against adversarial attacks, a critical concern as these systems handle sensitive tasks.

  3. What to watch

    About 3.8 percent of "stronger" prompt injections still succeed on GPT-5.6 Sol, meaning at scale—across hundreds or thousands of attempts—a sizable number of attacks get through, similar to Claude Opus 4.5. GPT-Red remains internal; OpenAI says a detailed paper will follow.

Ask the AI about this article →

Context & Analysis

OpenAI's approach to AI security marks a shift from traditional human-led red teaming to adversarial machine learning. By training GPT-Red via self-play reinforcement learning—where the attacking model and defending models improve against each other iteratively—OpenAI has created a process that scales far beyond what human testers can achieve. The 84 percent attack success rate for GPT-Red versus 13 percent for humans reveals a substantial efficiency gap; human red teamers, while valuable, cannot explore the adversarial space as exhaustively or rapidly.

The concrete improvements in GPT-5.6 Sol demonstrate that this testing regime translates into measurable security gains in production models. Six times fewer failures on direct prompt injections represents a substantial hardening, achieved without degrading the model's general capabilities—a critical constraint, since overfitting to defense can cripple utility. However, the persistence of 3.8 percent success rates on stronger attacks indicates that prompt injection remains a partially open problem even at the frontier. The comparison to Claude Opus 4.5 suggests this residual vulnerability is not unique to OpenAI's approach but reflects a broader challenge across leading models.

FAQ

How much better is GPT-Red than human red teamers?
GPT-Red finds successful attacks in 84 percent of test scenarios versus 13 percent for human red teamers. Additionally, GPT-5.6 Sol shows six times fewer failures on direct prompt injections than the best model from four months ago.
What real-world attack did GPT-Red perform?
In one test, GPT-Red manipulated an AI-powered vending machine in OpenAI's office, changing prices and canceling other customers' orders.
Are GPT models now completely protected against prompt injection attacks?
No. About 3.8 percent of "stronger" prompt injections still succeed on GPT-5.6 Sol, meaning at scale across hundreds or thousands of attempts, a sizable number of attacks get through, similar to Claude Opus 4.5.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Visko raises $10M, launches live AI video model OrbisSiliconANGLE AI · 1h ago
  • Runway unveils Solaris, an AI that generates app interfaces in real timeTHE DECODER · 1h ago
  • Google AI Search flags Facebook users as dangerTHE DECODER · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleWalden Robotics launches at $1.1B valuation with $300M funding