AIToday
Large Language ModelsAI Business & IndustryZenn AI/MLPublished: Oct 2, 2026, 22:00 JST

Anthropic Message Batches API cuts LLM token costs 50% off

Anthropic Message Batches API cuts LLM token costs 50% off

3 Key Points

  1. What happened

    Anthropic's Message Batches API offers a 50% off rate, takes up to 10,000 requests per batch, and returns results normally within 1 hour (24 hours max), working with all Claude models.

  2. Why it matters

    Switching asynchronous workloads to this API halves the token price, so businesses running large batch jobs no longer pay full synchronous rates for processing that does not need real-time responses.

  3. What to watch

    The discount applies only to asynchronous processing, so the savings hinge on whether a job can tolerate results arriving within 1 hour or up to 24 hours. Watch the maximum 24-hour turnaround when planning time-sensitive work.

WHO IT HITSEngineering teams running non-real-time workloads such as bulk document classification, tagging, or overnight report generation can cut token costs by half by switching from synchronous API calls to the Batches API.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The article is framed around a common engineering pain point: running a synchronous for-loop over tens of thousands of Claude Sonnet calls and watching the projected monthly bill grow before the job even finishes. The author argues that continuing to use real-time APIs for non-interactive work is pure waste.

The Batches API addresses this by accepting an array of requests with custom IDs, processing them asynchronously, and returning results as a stream that can be handled one by one without loading everything into memory. The implementation steps are framed as three core changes: passing an array to create, polling until the status becomes "ended", and consuming results as an AsyncIterable.

The stakes hinge on whether teams correctly judge which workloads truly need immediate responses. Chat, live API responses, and other real-time features cannot use this path. For workflows such as bulk classification, summarization, translation, data enrichment, and overnight evaluation reports, the halved token price may make previously expensive jobs affordable. The main practical constraint is the 1-hour typical and 24-hour maximum turnaround, which could matter for scheduling nightly or CI-integrated jobs.

FAQ
What is the maximum batch size and how long does it take?
Each batch can contain up to 10,000 requests. Results are typically returned within 1 hour, with a maximum of 24 hours.
When should I not use the Message Batches API?
It cannot be used for use cases that require chat or real-time API responses. It fits jobs where immediacy is not required.
Which models does it support?
The article states it works with all Claude models.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleMicrosoft expands biomimicry to more than 20 data center sites