
Anthropic has disclosed that its filters designed to prevent AI models from being used to develop chemical and biological weapons were offline for nearly a year, from May 2025 to April 2026, leaving roughly 133 million conversations from external contractors unfiltered.
Although the company's internal investigation found no evidence of actual misuse, the gap highlights a gap between Anthropic's public safety messaging—its CEO has called AI-assisted weapons development a bigger threat than cyberattacks—and its operational safeguards.
The company has since tightened contractor vetting requirements.
What happened
Anthropic's biological and chemical weapons classifiers—filters designed to block dangerous knowledge extraction—were inactive from May 2025 through April 2026. During that period, roughly 50,000 external contractors ran approximately 133 million chats with Anthropic's models without the safety filter in place.
Why it matters
Anthropic's CEO has publicly stated that AI-assisted development of chemical and biological weapons poses a bigger threat than cyberattacks, making this gap a significant lapse in the company's stated safety priorities. The contractors were vetted only by external vendors whose screening processes Anthropic now says were often insufficient, raising questions about oversight of high-risk human feedback work.
What to watch
Anthropic says its internal investigation found no evidence of actual misuse during the outage. The company has since tightened contractor requirements and also recently loosened classifiers on Fable 5 after researchers complained legitimate research was being blocked—a tension between security and usability that may shape future filter design.
Anthropic disclosed in a safety report that its classifiers designed to block the extraction of dangerous knowledge about chemical and biological weapons were inactive from May 2025 through April 2026. The outage affected approximately 50,000 external contractors who provided human feedback to Anthropic's models, collectively running roughly 133 million conversations without the safety filters in place.
According to Anthropic's disclosure, the affected contractors were vetted by external vendors, but the company acknowledges that these vendors' screening processes were often insufficient. The company has since tightened its contractor requirements. In its internal investigation following the discovery, Anthropic reports finding no evidence that the filters' absence led to actual misuse of the models for weapons development.
The incident underscores a tension between safety and usability in Anthropic's approach to AI safeguards. Even as the company was operating without its weapons-related classifiers, it recently loosened those same filters on Fable 5 after researchers complained that the filters were overly aggressive and blocking legitimate research. This dual pressure—preventing dangerous misuse while enabling legitimate scientific work—illustrates the challenge of calibrating safety thresholds in large-scale AI deployment.
Anthropic's disclosure of the year-long absence of its biological and chemical weapons classifiers presents a stark contrast to the company's public positioning on AI safety. CEO Dario Amodei has repeatedly emphasized that AI-assisted development of chemical and biological weapons represents a greater threat than cyberattacks, yet the company's safety infrastructure for one of its core stated risks remained offline for 12 months. The gap appears to have stemmed from a reliance on external vendor vetting that Anthropic itself now characterizes as insufficient—a delegation of responsibility that left approximately 133 million contractor interactions unfiltered.
The company's response includes tightening contractor requirements going forward, but the disclosure also reveals a concurrent tension in Anthropic's filtering approach. Recently, the company loosened classifiers on Fable 5 after researchers complained that the filters were blocking legitimate research. This pressure to reduce false positives (legitimate requests blocked) sits in direct tension with the goal of preventing dangerous knowledge extraction, and it may point to a broader challenge: determining where to set safety thresholds when aggressive filtering can impede legitimate work.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Apple is taking a focused approach to AI by integrating Apple Intelligence into existing products rather than…

Microsoft's Vice President for Southern Europe, Charles Calestroupat, told Fortune Greece that the region—Gree…

OpenAI dissolved its Preparedness team at the end of July, which had evaluated whether the company's AI models…

A survey by Epoch AI and Ipsos of 1,106 employed US adults (conducted July 10–19, 2026) found that 20 percent…

Artificial Analysis, known for independent LLM evaluations, has launched Optima, a platform that lets users bu…

Researchers tested Google DeepMind's DiffusionGemma model, which generates text through multiple diffusion ste…

The AI news that matters, in one minute each morning.
Sign up free