
OpenAI has developed GPT-Red, an AI model trained to automatically attack other language models by finding security vulnerabilities, allowing their defenses to be strengthened through repeated adversarial cycles.
When OpenAI tested GPT-Red's most powerful attacks against GPT-5.6, the success rate dropped from over 90 percent against last year's GPT-5 to less than 23 percent, demonstrating the approach's effectiveness in building more robust AI systems as models become increasingly complex and integrated with external tools and agents.
何が起きたか
OpenAIが「GPT-Red」と呼ばれるLLM攻撃モデルを開発しました。このモデルは他のLLMのサイバー防御を高めるための訓練相手として機能します。セルフプレイ・ループで何度も対戦を繰り返すうちに、GPT-Redは攻撃能力を、相手モデルは防御能力を高めていきました。
なぜ重要か
LLMはWebサイトやコード、他のエージェントと連携する複雑なAIエージェントとして使われるようになり、人間チームだけで起こり得るあらゆる攻撃に対応し続けることが難しくなっています。GPT-Redは「プロンプト・インジェクション」などの新たな攻撃手法を自動で発見できるため、より堅牢なモデルの開発が可能になります。
注目点
OpenAIがGPT-Redの攻撃を自社モデルに試した結果、2024年8月リリースのGPT-5では90%以上が成功した一方、新版GPT-5.6では成功率は23%未満に低下しました。ただしGPT-Redは人間のレッドチーマーを完全に置き換えるものではなく、人間が見逃す攻撃がある一方、GPT-Redも画像経由のプロンプト・インジェクション攻撃には十分対応していません。
Ask the AI about this article →
OpenAI's development of GPT-Red reflects a fundamental shift in AI safety testing as large language models become more capable and widely deployed. Traditionally, red-teaming—the practice of finding security vulnerabilities before release—has relied on human testers, but the growing complexity of LLMs and their integration with external systems like web browsers, code editors, and other AI agents has made comprehensive human-only testing increasingly difficult. The self-play loop approach OpenAI employed, in which GPT-Red iteratively attacks other models while they defend themselves, creates a training ground where both attacker and defender capabilities co-evolve. This mirrors competitive dynamics in cybersecurity, where defenders must anticipate and prepare for attacks that haven't yet been invented.
The "Fake Chain of Thought" attack exemplifies why automation may be necessary: it exploits a mechanism (chain-of-thought reasoning) that is internal to how modern LLMs solve problems, and discovering such attacks requires understanding the model's own logic. When OpenAI replicated a 2025 experiment in which human red teamers had found vulnerabilities in GPT-5, GPT-Red outperformed the human testers, suggesting that automated searching of the attack space can be more thorough in specific domains. However, the body makes clear that GPT-Red is not intended to replace human expertise—human testers can still find attacks GPT-Red misses, and vice versa, so the most effective approach combines both.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Amazon Web Services (AWS) has integrated its fully managed data warehouse service, Amazon Redshift, with Agent…

Visual Studio Code 1.135 now includes an experimental 'Rubber Duck' feature that lets developers request a sec…

Sonos announced a new app update with generative AI features, a new soundbar called the Beam Ultra, and its se…
