AIToday
Large Language ModelsMITテクノロジーレビューPublished: Jul 16, 2026, 06:00 JST2 min read

OpenAI、AIで防御を鍛える攻撃モデル「GPT-Red」開発

OpenAI、AIで防御を鍛える攻撃モデル「GPT-Red」開発

Key takeaway

  • OpenAI has developed GPT-Red, an AI model trained to automatically attack other language models by finding security vulnerabilities, allowing their defenses to be strengthened through repeated adversarial cycles.

  • When OpenAI tested GPT-Red's most powerful attacks against GPT-5.6, the success rate dropped from over 90 percent against last year's GPT-5 to less than 23 percent, demonstrating the approach's effectiveness in building more robust AI systems as models become increasingly complex and integrated with external tools and agents.

3 Key Points

  1. 何が起きたか

    OpenAIが「GPT-Red」と呼ばれるLLM攻撃モデルを開発しました。このモデルは他のLLMのサイバー防御を高めるための訓練相手として機能します。セルフプレイ・ループで何度も対戦を繰り返すうちに、GPT-Redは攻撃能力を、相手モデルは防御能力を高めていきました。

  2. なぜ重要か

    LLMはWebサイトやコード、他のエージェントと連携する複雑なAIエージェントとして使われるようになり、人間チームだけで起こり得るあらゆる攻撃に対応し続けることが難しくなっています。GPT-Redは「プロンプト・インジェクション」などの新たな攻撃手法を自動で発見できるため、より堅牢なモデルの開発が可能になります。

  3. 注目点

    OpenAIがGPT-Redの攻撃を自社モデルに試した結果、2024年8月リリースのGPT-5では90%以上が成功した一方、新版GPT-5.6では成功率は23%未満に低下しました。ただしGPT-Redは人間のレッドチーマーを完全に置き換えるものではなく、人間が見逃す攻撃がある一方、GPT-Redも画像経由のプロンプト・インジェクション攻撃には十分対応していません。

Ask the AI about this article →

Context & Analysis

OpenAI's development of GPT-Red reflects a fundamental shift in AI safety testing as large language models become more capable and widely deployed. Traditionally, red-teaming—the practice of finding security vulnerabilities before release—has relied on human testers, but the growing complexity of LLMs and their integration with external systems like web browsers, code editors, and other AI agents has made comprehensive human-only testing increasingly difficult. The self-play loop approach OpenAI employed, in which GPT-Red iteratively attacks other models while they defend themselves, creates a training ground where both attacker and defender capabilities co-evolve. This mirrors competitive dynamics in cybersecurity, where defenders must anticipate and prepare for attacks that haven't yet been invented.

The "Fake Chain of Thought" attack exemplifies why automation may be necessary: it exploits a mechanism (chain-of-thought reasoning) that is internal to how modern LLMs solve problems, and discovering such attacks requires understanding the model's own logic. When OpenAI replicated a 2025 experiment in which human red teamers had found vulnerabilities in GPT-5, GPT-Red outperformed the human testers, suggesting that automated searching of the attack space can be more thorough in specific domains. However, the body makes clear that GPT-Red is not intended to replace human expertise—human testers can still find attacks GPT-Red misses, and vice versa, so the most effective approach combines both.

FAQ

What specific attack technique did GPT-Red discover that researchers had not seen before?
GPT-Red discovered a new prompt injection attack called "Fake Chain of Thought," in which it writes false records into another model's chain-of-thought reasoning process, tricking the model into believing and acting on false information.
Will OpenAI release GPT-Red publicly?
No. OpenAI has no plans to release GPT-Red publicly and believes it is far more powerful than imitation models. The research team developed it over more than a year using substantial computational resources.
What are the limitations of GPT-Red compared to human red teamers?
GPT-Red is not skilled at designing attacks that require repeated back-and-forth dialogue between attacker and target—something human attackers can easily do. It also does not yet handle prompt injection attacks that use images to pass text to models.
MITテクノロジーレビューRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • CBTS launches Forge Agents for custom AI agentsSiliconANGLE AI · 49m ago
  • Imec CEO: AI era widens chip-model-CSP collaborationDIGITIMES Asia · 49m ago
  • Alphabet's AI Overviews reach 2.5B monthly usersYahoo Finance AI · 49m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleJefferies picks Amazon as top hyperscaler, cheaper than Walmart, Google