AIToday
Large Language ModelsAI Safety & AlignmentTHE DECODERPublished: Aug 16, 2026, 19:01 JST3 min read

Anthropic's bio-weapons filter offline for a year, exposing 133M requests

Anthropic's bio-weapons filter offline for a year, exposing 133M requests

Key takeaway

  • Anthropic has disclosed that its filters designed to prevent AI models from being used to develop chemical and biological weapons were offline for nearly a year, from May 2025 to April 2026, leaving roughly 133 million conversations from external contractors unfiltered.

  • Although the company's internal investigation found no evidence of actual misuse, the gap highlights a gap between Anthropic's public safety messaging—its CEO has called AI-assisted weapons development a bigger threat than cyberattacks—and its operational safeguards.

  • The company has since tightened contractor vetting requirements.

3 Key Points

  1. What happened

    Anthropic's biological and chemical weapons classifiers—filters designed to block dangerous knowledge extraction—were inactive from May 2025 through April 2026. During that period, roughly 50,000 external contractors ran approximately 133 million chats with Anthropic's models without the safety filter in place.

  2. Why it matters

    Anthropic's CEO has publicly stated that AI-assisted development of chemical and biological weapons poses a bigger threat than cyberattacks, making this gap a significant lapse in the company's stated safety priorities. The contractors were vetted only by external vendors whose screening processes Anthropic now says were often insufficient, raising questions about oversight of high-risk human feedback work.

  3. What to watch

    Anthropic says its internal investigation found no evidence of actual misuse during the outage. The company has since tightened contractor requirements and also recently loosened classifiers on Fable 5 after researchers complained legitimate research was being blocked—a tension between security and usability that may shape future filter design.

In Depth

Read the full story

Anthropic disclosed in a safety report that its classifiers designed to block the extraction of dangerous knowledge about chemical and biological weapons were inactive from May 2025 through April 2026. The outage affected approximately 50,000 external contractors who provided human feedback to Anthropic's models, collectively running roughly 133 million conversations without the safety filters in place.

According to Anthropic's disclosure, the affected contractors were vetted by external vendors, but the company acknowledges that these vendors' screening processes were often insufficient. The company has since tightened its contractor requirements. In its internal investigation following the discovery, Anthropic reports finding no evidence that the filters' absence led to actual misuse of the models for weapons development.

The incident underscores a tension between safety and usability in Anthropic's approach to AI safeguards. Even as the company was operating without its weapons-related classifiers, it recently loosened those same filters on Fable 5 after researchers complained that the filters were overly aggressive and blocking legitimate research. This dual pressure—preventing dangerous misuse while enabling legitimate scientific work—illustrates the challenge of calibrating safety thresholds in large-scale AI deployment.

Context & Analysis

Anthropic's disclosure of the year-long absence of its biological and chemical weapons classifiers presents a stark contrast to the company's public positioning on AI safety. CEO Dario Amodei has repeatedly emphasized that AI-assisted development of chemical and biological weapons represents a greater threat than cyberattacks, yet the company's safety infrastructure for one of its core stated risks remained offline for 12 months. The gap appears to have stemmed from a reliance on external vendor vetting that Anthropic itself now characterizes as insufficient—a delegation of responsibility that left approximately 133 million contractor interactions unfiltered.

The company's response includes tightening contractor requirements going forward, but the disclosure also reveals a concurrent tension in Anthropic's filtering approach. Recently, the company loosened classifiers on Fable 5 after researchers complained that the filters were blocking legitimate research. This pressure to reduce false positives (legitimate requests blocked) sits in direct tension with the goal of preventing dangerous knowledge extraction, and it may point to a broader challenge: determining where to set safety thresholds when aggressive filtering can impede legitimate work.

FAQ

How long were the filters offline and how many people and conversations were affected?
The biological and chemical weapons classifiers were inactive from May 2025 through April 2026. During that period, approximately 50,000 external contractors ran roughly 133 million chats with Anthropic's models without the filters active.
Did Anthropic find evidence that the filters' absence led to actual misuse?
No. According to Anthropic, its internal investigation turned up no evidence of actual misuse during the outage.
How were the external contractors vetted?
The contractors were vetted only by external vendors whose screening processes Anthropic says were often insufficient. The company has since tightened contractor requirements.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSouthern Europe's AI lead hinges on skills, not tech—Microsoft exec

The AI news that matters, in one minute each morning.

Sign up free