AIToday
Large Language ModelsAI Safety & AlignmentTHE DECODERPublished: Sep 12, 2026, 01:00 JST3 min read

Anthropic report: Claude misused for missiles, spy malware, 151 million exchanges mined

Anthropic report: Claude misused for missiles, spy malware, 151 million exchanges mined

3 Key Points

  1. What happened

    Anthropic's threat report covering December 2025 through August 2026 documents Claude misuse across seven categories — including a Yemen cell (GTG-87001) using Claude Code for three missile programs, a Russian-speaking actor (GTG-20006) whose agents rewrote malware to evade antivirus, and Alibaba's Qwen lab extracting over 151 million exchanges for training data.

  2. Why it matters

    Anthropic says autonomy lowers the cost of attacks, making previously unprofitable targets worth pursuing — and that writing new detection signatures no longer slows an attacker if AI cycles through changes faster than signatures can be rolled out, with more than 20 organizations targeted and a complete proprietary drone-vision SDK stolen.

  3. What to watch

    Anthropic has already launched Fable 5 with stricter safeguards for dual-use biology requests and 'preserved thinking' in Fable 5.1 to block distillation via new API accounts — the test is whether classifiers can both enable useful work and prevent harm, which Anthropic itself concludes may be impossible because user intent in dual-use areas cannot be reliably detected.

WHO IT HITSEnterprise security teams defending against AI-accelerated attacks face a structural problem: anomaly-based detection degrades when malware rewrites itself faster than signatures can be deployed. AI lab trust and safety teams are similarly affected, since the report documents labs relaying their own customers' requests to Claude and buying transcripts through intermediaries.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

Anthropic's report breaks misuse into seven categories and says it documents novel cases rather than the typical kind. The affected models were primarily Haiku, Sonnet, and Opus, while Fable and Mythos appeared in only a single distillation case — suggesting that older, widely deployed models remain the primary vector even as newer ones carry stricter safeguards.

The report's cyber chapter argues that sophisticated attacks no longer require sophisticated attackers. The techniques themselves are familiar — stolen credentials, unpatched devices, SQL injection, phishing — but reconnaissance, exploitation, and tool-building now run in parallel at machine speed. A Russian-speaking actor tracked as GTG-20006 ran AI agents that checked whether malware was being flagged, rewrote and recompiled the code when it was, and slipped past detection again; more than 20 organizations were targeted, and the actor stole a complete proprietary SDK for a drone vision system.

The distillation findings raise a different set of questions. Anthropic says the illegitimate version is industrial-scale and covert, enabled by fake accounts and 'transfer stations.' But the strangest cases involve labs relaying their own customers' requests — Moonshot AI (GTG-16002) relayed nearly 300,000 customer requests over ten days while users believed they were using a Kimi model, and DeepSeek (GTG-16001) routed selected users to Claude Opus, over 12.1 million exchanges in 14 days. Among those, Anthropic found a user likely tied to the People's Liberation Army analyzing CCTV footage from hundreds of cameras in Chengdu, and an operator with live credentials for a database linked to the Russian Ministry of Defense. Whether stricter API controls can prevent this kind of relay — where the misuse happens through intermediaries rather than direct accounts — is the open question the report leaves unresolved.

FAQ
How did Alibaba's Qwen lab extract data from Claude?
Qwen (GTG-16005) used a fixed prompt to make Claude write out its reasoning traces before answering, then processed the transcripts into fine-tuning data. The campaign peaked at almost three million exchanges a day from more than 3,500 fraudulent accounts, totaling over 151 million between May and July 2026.
What did Anthropic do in response to the misuse?
Anthropic launched Claude Fable 5 with stricter safeguards for dual-use biology requests. It also introduced 'preserved thinking' with Fable 5.1 to keep new API accounts from manipulating the context and extraction by distillation.
Did suspending the account stop the Mali surveillance platform?
No. Anthropic says suspending the account interrupted only the development work. The platform, which monitors roughly 25 million SIM cards, runs on local models on-premises and continued operating.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 4h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 4h ago
  • OpenAI agents hit RubyGems, undisclosed since May 12thSimon Willison's Weblog · 4h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleIrregular: agent escape was training, not sandbox