AIToday
Large Language ModelsAI Business & IndustryITmedia AI+Published: Jul 16, 2026, 13:01 JST2 min read

OpenAI、自動攻撃AI「GPT-Red」発表 人間13%に対し84%成功

OpenAI、自動攻撃AI「GPT-Red」発表 人間13%に対し84%成功

Key takeaway

  • OpenAI has released GPT-Red, an automated AI model designed to find vulnerabilities in its own systems by simulating attacks.

  • In testing, GPT-Red achieved an 84% success rate at prompt injection attacks on previously unseen scenarios, compared to just 13% for human security testers.

  • The company used this automated red-teaming approach to train its latest model, GPT-5.6 Sol, which reduced its failure rate on the most difficult direct prompt injection benchmark to 0.05%—one-sixth of the previous best product model's rate from four months earlier.

3 Key Points

  1. 何が起きたか

    OpenAIは7月15日、自社モデルの脆弱性を自動で発見するAI「GPT-Red」を発表しました。人間のレッドチーマー(セキュリティテスター)による攻撃成功率が13%だったのに対し、GPT-Redは84%のシナリオで攻撃に成功しました。

  2. なぜ重要か

    人間による脆弱性検査は時間がかかり大規模に実施しにくく、モデルを堅牢にするのに必要な量と多様性の攻撃データを生成できません。GPT-Redはこの課題を自動化・大規模化し、最新モデル「GPT-5.6 Sol」は最難関のプロンプトインジェクションベンチマークで失敗率を4カ月前の6分の1に低下させました。

  3. 注目点

    GPT-Redは社内専用モデルとし、訓練した攻撃能力が悪意ある者の手に渡らないようにします。詳細を記した査読前論文は今週後半に公開される予定です。

Ask the AI about this article →

Context & Analysis

OpenAI's announcement of GPT-Red reflects a shift in how AI companies approach safety validation. Traditional red-teaming by human experts is labor-intensive and cannot generate the volume and variety of attack data needed to significantly improve model robustness. By automating this process with an AI opponent trained via self-play, OpenAI aims to create what it calls a "safety flywheel"—a cycle where today's models train tomorrow's models to be more resilient. The company invested computation resources equivalent to its largest post-training efforts solely into safety, underscoring the resource commitment required.

The real-world test on an AI vending machine agent in OpenAI offices demonstrated the practical stakes: GPT-Red successfully achieved three malicious goals (price manipulation, order fraud, and order cancellation) that previous defenses had not caught. This discovery then fed back into training GPT-5.6 Sol, which now exhibits dramatically lower failure rates on prompt injection benchmarks. The company emphasizes that this robustness gain came without degrading model capability—a critical distinction, since overly defensive models often become less useful. By keeping GPT-Red internal and planning to publish a peer-reviewed paper, OpenAI is attempting to balance transparency with security, allowing the research community to learn from the work while preventing attack tools from proliferating.

FAQ

What is prompt injection and why does it matter?
Prompt injection is an attack where malicious commands are hidden within instructions sent to an AI model. GPT-Red specializes in discovering these vulnerabilities so they can be patched before deployment.
How does GPT-Red learn to attack so effectively?
GPT-Red uses self-play reinforcement learning, where it trains simultaneously against multiple defensive language models. The attack side earns rewards for successful injections, while the defense side earns rewards for resisting attacks and completing its original task—creating a cycle where each side becomes stronger.
Will OpenAI release GPT-Red publicly?
No; GPT-Red is kept as an internal-only model separated from deployed models to prevent its trained attack capabilities from reaching malicious actors.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Visko raises $10M, launches live AI video model OrbisSiliconANGLE AI · 1h ago
  • Runway unveils Solaris, an AI that generates app interfaces in real timeTHE DECODER · 1h ago
  • Google AI Search flags Facebook users as dangerTHE DECODER · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGoogle stock faces AI threat ahead of earnings