
What happened
Anthropic began offering Claude Haiku 5.5, its smallest and fastest model, with average operating costs roughly 75% lower than Haiku 4.5, priced at $0.1 per million input tokens and $0.5 for output.
Why it matters
For businesses running high-volume tasks such as summarization, data compression, database queries and classification, the lower cost per task could make large-scale AI operations substantially cheaper — provided the model's quality holds up in production.
What to watch
The pricing edge hinges on whether the newer model's quality claims translate into real-world workloads, since its benchmark scores still trail the more expensive Sonnet 5.5. Also note that once a request exceeds 100,000 tokens, input pricing rises to $0.5 and output to $2.5 per million.
WHO IT HITSCost-sensitive engineering and data teams running high-volume AI workloads — such as summarization, classification and database queries — stand to gain from the reduced token pricing, and the model's availability on AWS, Google Cloud and Microsoft Azure may make it easier for enterprise buyers already using those platforms.
Summaries like this, in your inbox every morning.
Anthropic's Claude line has seen a steady cadence of updates to higher-end models such as Fable, Opus and Sonnet, while Haiku — the smallest and fastest tier — had gone roughly a year without a refresh. The arrival of Haiku 5.5 therefore fills a gap at the cheap end of the lineup, where customers doing repetitive, high-volume work had been running an older generation.
The model is designed for tasks where speed and cost matter more than peak reasoning: summarization, data compression, database queries and classification. In coding, it can act as a sub-agent alongside Opus 5.5 or Sonnet 5.5, handling cheaper steps so the expensive models are reserved for harder work. Its benchmark scores show a large jump over Haiku 4.5 and even a lead over GPT-6 Luna, though they remain below Sonnet 5.5.
Haiku 5.5 also introduces effort-level settings to the Haiku line for the first time, letting users trade cost against capability. The real test is whether the price advantage holds for large production workloads, where token volumes can push requests past the 100,000-token threshold — a point at which the cost per token rises, though it is still half of Haiku 4.5. Separately, Sonnet 5.5 now has its cache-read fee cut in half, which Anthropic says makes it about 20% cheaper to run for most agent work.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Episil Technologies Chairman Rex Hsu said silicon carbide capacity utilization rebounded from approximately 30…

OpenAI rolled out GPT-6 and Intelligent UI to Plus, Pro, Business, and Enterprise plans worldwide starting Oct…

SpaceX reportedly is arranging about $10 billion in bank loans and $30 billion in investment-grade bonds to bu…

Microsoft announced a Windows strategy it calls hybrid intelligence, plus new Copilot abilities to read local…

Microsoft began preorders for two RTX Spark machines on October 7, 2026 — the Surface Laptop Ultra, with up to…

At its first live event in two years, Microsoft unveiled the Surface Laptop Ultra with Nvidia's RTX Spark chip…
