
What happened
Anthropic released Claude Haiku 5.5, a small model for repetitive work, with list prices of 10 cents per million input tokens and 50 cents per million output for prompts up to 100,000 tokens. The company also halved cache-read prices on Claude Sonnet 5.5, from 20 cents per million tokens to 10 cents.
Why it matters
The Sonnet 5.5 cut is expected to take about 20% off the cost of most agentic work on that model.
What to watch
Whether these price cuts hold as Anthropic's newest small model takes on OpenAI's GPT-6 Luna, which charges the same 10-cent and 50-cent rates. Watch whether the quoted savings survive the new tokenizer, which uses slightly more tokens per task.
WHO IT HITSTeams running high-volume summaries, classification, live customer support, and browser automation — plus coding teams using small models as subagents — are the main beneficiaries of the cheaper Haiku 5.5 rates.
Summaries like this, in your inbox every morning.
Anthropic's newest small model arrives two weeks after Opus 5.5 launched Sept. 22 and days after Sonnet 5.5 launched Sept. 28, making Haiku 5.5 the third model in the 5.5 generation. The company is aiming it at repetitive work, with high-volume summaries and classification as the main target, and says no Anthropic model runs faster at standard speed — which is why it suggests the model for live customer support and browser automation. Haiku 4.5, released last October, sets the pricing baseline the new model undercuts: $1 per million input tokens and $5 per million output, versus 10 cents and 50 cents for prompts up to 100,000 tokens.
The 10-cent and 50-cent rates match what OpenAI Group PBC charges for GPT-6 Luna, the low-cost model it launched last month. Anthropic's published benchmarks put Haiku 5.5 ahead of Luna on all six tests where both have a score, including 72.4% versus 48.9% on OSWorld 2.1 and 39.2% versus 16.4% on Terminal-Bench 4.0. Anthropic still points customers to Sonnet 5.5 and Opus 5.5 for complex agentic coding, suggesting the small model is not meant to replace the larger ones.
Asana Inc., which tested the model before release through its AI Teammates agent evaluation suite, reported task completion latency more than 30% lower than the model it uses today and inference on each agent turn up to 2.5 times faster. On safety, Anthropic said alignment testing turned up far fewer instances of misaligned behavior than Haiku 4.5 showed, with cybersecurity safeguards that allow more defensive work than Sonnet 5.5 permits — though penetration testing remains blocked. Whether the price advantage translates into durable savings may hinge on how the new tokenizer, which uses slightly more tokens per task, affects real-world usage patterns.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Preferred Networks began offering a major update to its Japan-focused AI translation service PLaMo翻訳, claiming…

Using only NVIDIA's official figures from P100 (2016) to Rubin (2026), the same 10 years divides out to roughl…

A Playwright (1.56.1) script swept 320–1440px in 40px steps, then refined to 1px

US nonprofit Common Sense Media said ChatGPT's parent-notification function is not working, and that conversat…

SOMPO Holdings set AI risk governance in March 2025, requiring risk assessment and model output testing for gr…

Canon began offering a new support service for home inkjet printers that uses a monitoring function and genera…
