
What happened
Anthropic's Message Batches API offers a 50% off rate, takes up to 10,000 requests per batch, and returns results normally within 1 hour (24 hours max), working with all Claude models.
Why it matters
Switching asynchronous workloads to this API halves the token price, so businesses running large batch jobs no longer pay full synchronous rates for processing that does not need real-time responses.
What to watch
The discount applies only to asynchronous processing, so the savings hinge on whether a job can tolerate results arriving within 1 hour or up to 24 hours. Watch the maximum 24-hour turnaround when planning time-sensitive work.
WHO IT HITSEngineering teams running non-real-time workloads such as bulk document classification, tagging, or overnight report generation can cut token costs by half by switching from synchronous API calls to the Batches API.
Summaries like this, in your inbox every morning.
The article is framed around a common engineering pain point: running a synchronous for-loop over tens of thousands of Claude Sonnet calls and watching the projected monthly bill grow before the job even finishes. The author argues that continuing to use real-time APIs for non-interactive work is pure waste.
The Batches API addresses this by accepting an array of requests with custom IDs, processing them asynchronously, and returning results as a stream that can be handled one by one without loading everything into memory. The implementation steps are framed as three core changes: passing an array to create, polling until the status becomes "ended", and consuming results as an AsyncIterable.
The stakes hinge on whether teams correctly judge which workloads truly need immediate responses. Chat, live API responses, and other real-time features cannot use this path. For workflows such as bulk classification, summarization, translation, data enrichment, and overnight evaluation reports, the halved token price may make previously expensive jobs affordable. The main practical constraint is the 1-hour typical and 24-hour maximum turnaround, which could matter for scheduling nightly or CI-integrated jobs.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
DIGITIMES estimates global AI server shipments will exceed 2.5 million units in 2026, including 2.37 million h…

Google Cloud revenue grew 82% to $24.77B on Alphabet's TPU-powered stack, while NVIDIA's revenue rose 105.8% a…

The SEC revealed cases where fund advisors allegedly sold investors pre-IPO shares of SpaceX, OpenAI and other…

Testing Azure API Management's llm-token-limit policy at 800 tokens per hour, actual consumption hit 1,472 tok…

Qwen released Qwen3.8-Flash-Next on August 27, 2026, calling it a preview of the architecture planned for Qwen…

A migration writeup lists seven breakages for code moved to claude-opus-5 or claude-sonnet-5, all in the offic…
