AIToday
Large Language ModelsImpress WatchPublished: Sep 29, 2026, 13:00 JST

Anthropic's Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, cuts task cost up to 30%

Anthropic's Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, cuts task cost up to 30%

3 Key Points

  1. What happened

    Anthropic announced Claude Sonnet 5.5, a model 30%+ faster than Sonnet 5, scoring 70.6% on Terminal-Bench 4.0 versus Sonnet 5's 10.3%. It lists at $2 per million input tokens, $10 per million output tokens and $0.20 per million cached-read tokens, with task cost down up to 30%.

  2. Why it matters

    Same-price, faster generation with fewer tokens needed per task could let teams run clear-scope, everyday work at up to 30% lower cost.

  3. What to watch

    Some evaluations show Sonnet 5.5 matching the company's top-tier Opus 5.5 at Max effort, but Anthropic's and outside testers' tests found Opus 5.5 better on complex, open-ended work making sustained judgment calls. Watch the upcoming Claude Haiku 5.5 addition.

WHO IT HITSSoftware developers, technical writers and analysts doing clear-scope, everyday work — bug fixes, documents, slides, spreadsheets — can run more iterations at the same token price or lower task cost. Teams weighing complex, judgment-heavy tasks may still need Opus 5.5.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Sonnet 5.5 arrives as the second release in Anthropic's Claude 5.5 family, following the top-tier Opus 5.5. The company positions the pair by task type: Opus 5.5 for complex work needing careful judgment, and Sonnet 5.5 for clearly scoped everyday tasks, bug fixes, and polished documents, slides and spreadsheets. Sonnet 5.5 also brings a sharp eye for design and, per Anthropic, generates clearer prose than the previous generation.

The benchmark spread is wide. On Terminal-Bench 4.0, an agentic coding evaluation, Sonnet 5.5 scored 70.6% against Sonnet 5's 10.3%, while on GDPval-AA it landed 2 points below Opus 5.5. It is said to be close to Opus 5.5 in computer use and chart recognition, and well above Sonnet 5 and GPT-6 Sol on long-horizon knowledge work. A single screenshot alone was enough for it to play through Pokémon Red. In some evaluations at Max effort it matches Opus 5.5, yet both internal and outside testing reportedly favor Opus 5.5 on complex, open-ended work. At Low and Medium effort levels it keeps per-task cost to roughly one-tenth while beating Sonnet 5's top score.

The alignment and safety picture is deliberately nuanced. Sonnet 5.5 holds or exceeds Sonnet 5 on most alignment metrics in Anthropic's automated behavior audits. Its cybersecurity capability improved markedly, reaching a level similar to Opus 5. It is the first Sonnet model to carry the same cyber safeguards and fallback functions as top-capability models, applied like Opus 5.5's safeguards, with biological safeguards unchanged from Sonnet 5. Both safeguards target only rare high-risk requests, so most everyday software development and life-sciences work should be unaffected. The broader question of whether same-price models can absorb the highest-judgment tasks will hinge on how the reported Opus 5.5 edge holds up once more outside testers use Sonnet 5.5 in production.

FAQ
How much does Claude Sonnet 5.5 cost?
It has the same list price as Sonnet 5: $2 per million input tokens, $10 per million output tokens, and $0.20 per million cached-read tokens. Anthropic says fewer tokens per task cut task cost up to 30%.
How does Claude Sonnet 5.5 differ from Opus 5.5?
Sonnet 5.5 is built as a fast, low-cost model complementing Opus 5.5. It scores 2 points below Opus 5.5 on GDPval-AA, and Anthropic's and outside testers' tests found Opus 5.5 better on complex, open-ended work requiring sustained judgment.
Where can I use Claude Sonnet 5.5?
It is available on Claude and on many platforms including Amazon Web Services, Google Cloud and Microsoft Azure. A more cost-oriented Claude Haiku 5.5 is slated to join the Claude 5.5 family within weeks.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Okta's Blueprint Alliance takes on agent runtime securitySiliconANGLE AI · 9m ago
  • Omdia: 400 security leaders name confusion top AI agent identity blockerSiliconANGLE AI · 9m ago
  • Meta launches Meta Enterprise Platform, taps MongoDB CEO DesaiITmedia AI+ · 9m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleClaude Sonnet 5.5 hits Snowflake Cortex AI preview