
What happened
Anthropic announced Claude Sonnet 5.5, a model 30%+ faster than Sonnet 5, scoring 70.6% on Terminal-Bench 4.0 versus Sonnet 5's 10.3%. It lists at $2 per million input tokens, $10 per million output tokens and $0.20 per million cached-read tokens, with task cost down up to 30%.
Why it matters
Same-price, faster generation with fewer tokens needed per task could let teams run clear-scope, everyday work at up to 30% lower cost.
What to watch
Some evaluations show Sonnet 5.5 matching the company's top-tier Opus 5.5 at Max effort, but Anthropic's and outside testers' tests found Opus 5.5 better on complex, open-ended work making sustained judgment calls. Watch the upcoming Claude Haiku 5.5 addition.
WHO IT HITSSoftware developers, technical writers and analysts doing clear-scope, everyday work — bug fixes, documents, slides, spreadsheets — can run more iterations at the same token price or lower task cost. Teams weighing complex, judgment-heavy tasks may still need Opus 5.5.
Summaries like this, in your inbox every morning.
Sonnet 5.5 arrives as the second release in Anthropic's Claude 5.5 family, following the top-tier Opus 5.5. The company positions the pair by task type: Opus 5.5 for complex work needing careful judgment, and Sonnet 5.5 for clearly scoped everyday tasks, bug fixes, and polished documents, slides and spreadsheets. Sonnet 5.5 also brings a sharp eye for design and, per Anthropic, generates clearer prose than the previous generation.
The benchmark spread is wide. On Terminal-Bench 4.0, an agentic coding evaluation, Sonnet 5.5 scored 70.6% against Sonnet 5's 10.3%, while on GDPval-AA it landed 2 points below Opus 5.5. It is said to be close to Opus 5.5 in computer use and chart recognition, and well above Sonnet 5 and GPT-6 Sol on long-horizon knowledge work. A single screenshot alone was enough for it to play through Pokémon Red. In some evaluations at Max effort it matches Opus 5.5, yet both internal and outside testing reportedly favor Opus 5.5 on complex, open-ended work. At Low and Medium effort levels it keeps per-task cost to roughly one-tenth while beating Sonnet 5's top score.
The alignment and safety picture is deliberately nuanced. Sonnet 5.5 holds or exceeds Sonnet 5 on most alignment metrics in Anthropic's automated behavior audits. Its cybersecurity capability improved markedly, reaching a level similar to Opus 5. It is the first Sonnet model to carry the same cyber safeguards and fallback functions as top-capability models, applied like Opus 5.5's safeguards, with biological safeguards unchanged from Sonnet 5. Both safeguards target only rare high-risk requests, so most everyday software development and life-sciences work should be unaffected. The broader question of whether same-price models can absorb the highest-judgment tasks will hinge on how the reported Opus 5.5 edge holds up once more outside testers use Sonnet 5.5 in production.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Omdia principal analyst Todd Thiemann surveyed 400 security leaders; the top inhibitor to AI agent identity se…
Okta launched a multivendor reference architecture, the Blueprint Alliance, for agent runtime security
Meta said on September 28 it will offer its AI models and agents to companies and developers, starting with Mu…

Chipmaker AMD agreed to acquire World Labs, the AI startup founded by industry pioneer Fei-Fei Li, for $8.2 bi…

Anthropic's Thariq Shihipar said on the Latent Space podcast that agent security may become one of the definin…

Anthropic announced Claude Sonnet 5.5 on September 28
