
OpenAI has released GPT-Red, an automated AI model designed to find vulnerabilities in its own systems by simulating attacks.
In testing, GPT-Red achieved an 84% success rate at prompt injection attacks on previously unseen scenarios, compared to just 13% for human security testers.
The company used this automated red-teaming approach to train its latest model, GPT-5.6 Sol, which reduced its failure rate on the most difficult direct prompt injection benchmark to 0.05%—one-sixth of the previous best product model's rate from four months earlier.
何が起きたか
OpenAIは7月15日、自社モデルの脆弱性を自動で発見するAI「GPT-Red」を発表しました。人間のレッドチーマー(セキュリティテスター)による攻撃成功率が13%だったのに対し、GPT-Redは84%のシナリオで攻撃に成功しました。
なぜ重要か
人間による脆弱性検査は時間がかかり大規模に実施しにくく、モデルを堅牢にするのに必要な量と多様性の攻撃データを生成できません。GPT-Redはこの課題を自動化・大規模化し、最新モデル「GPT-5.6 Sol」は最難関のプロンプトインジェクションベンチマークで失敗率を4カ月前の6分の1に低下させました。
注目点
GPT-Redは社内専用モデルとし、訓練した攻撃能力が悪意ある者の手に渡らないようにします。詳細を記した査読前論文は今週後半に公開される予定です。
Ask the AI about this article →
OpenAI's announcement of GPT-Red reflects a shift in how AI companies approach safety validation. Traditional red-teaming by human experts is labor-intensive and cannot generate the volume and variety of attack data needed to significantly improve model robustness. By automating this process with an AI opponent trained via self-play, OpenAI aims to create what it calls a "safety flywheel"—a cycle where today's models train tomorrow's models to be more resilient. The company invested computation resources equivalent to its largest post-training efforts solely into safety, underscoring the resource commitment required.
The real-world test on an AI vending machine agent in OpenAI offices demonstrated the practical stakes: GPT-Red successfully achieved three malicious goals (price manipulation, order fraud, and order cancellation) that previous defenses had not caught. This discovery then fed back into training GPT-5.6 Sol, which now exhibits dramatically lower failure rates on prompt injection benchmarks. The company emphasizes that this robustness gain came without degrading model capability—a critical distinction, since overly defensive models often become less useful. By keeping GPT-Red internal and planning to publish a peer-reviewed paper, OpenAI is attempting to balance transparency with security, allowing the research community to learn from the work while preventing attack tools from proliferating.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Visko raised $10 million in pre-seed funding from Llama Ventures and opened public access to its first foundat…
U.S. markets ended August higher, with the S&P 500 up 2.6% and the Nasdaq up 3.9%

Neurovia AI, an Abu Dhabi-based company, is pitching Saudi security agencies software that it says can compres…

AI company Runway has unveiled Solaris, the first model in a new category it calls "Interface World Models." I…

Google's AI search gave advice to call emergency services for users alone with an African, Indian, or Pakistan…

John Deere introduced JD, a conversational AI tool that lets farmers ask open-ended questions about their hist…
