AIToday
Tomasz Tunguz (Theory Ventures)Published: Jun 26, 2026, 06:01 JST1 min read

Sail Research closes Series A funding to offer batch inference for AI agents at 6× lower cost than real-time alternatives.

Sail Research closes Series A funding to offer batch inference for AI agents at 6× lower cost than real-time alternatives.

Key takeaway

  • Sail Research, a startup offering cheaper batch inference for AI agents, has closed Series A funding.

  • The company routes requests to open-source models at dramatically lower cost than real-time services—GLM-5.1 on Sail costs 6× less per token than Anthropic's Haiku—by using spare server capacity and queuing requests instead of reserving capacity per user.

  • As AI agents move from chat assistants to background workers processing data overnight, this batch-focused approach may become the dominant inference model.

3 Key Points

  1. What happened

    Sail Research, founded by Neil Movva and Samir Menon, announced a Series A investment alongside Kleiner Perkins, Redpoint, and Sequoia. The company routes asynchronous inference requests across open models like DeepSeek, Qwen, Kimi, and GLM, selecting the cheapest capable model for each task. GLM-5.1 on Sail costs 6x less per token than Anthropic's Haiku.

  2. Why it matters

    As AI agents shift from chat assistants into background workers running overnight tasks, most inference workloads will likely flow through batch queues rather than real-time systems. Batch inference costs far less because it uses spot capacity and idle server time instead of reserving capacity per request, making it economically viable for long-running tasks like code review, research, and document processing.

  3. What to watch

    Sailboxes—cloud computers that hold state across agent tasks, pause during inference waits, and resume in seconds—let customers pay only for active compute time. Sail has already served trillions of tokens to customers in code review, deep research, and cybersecurity.

Ask the AI about this article →

FAQ

How much cheaper is Sail's inference than alternatives?
GLM-5.1 on Sail costs 6x less per token than Anthropic's Haiku. The same token cost is 6x less when waiting two minutes instead of two seconds for a code review.
What models does Sail support?
Sail distributes requests across open models including DeepSeek, Qwen, Kimi, and GLM, selecting the cheapest capable model for each task.
How does Sail keep costs low?
Sail uses spot capacity when available and fails over to reliable compute when it is not, using fleet-aware orchestration to keep utilization high and cost low. Customers pay only for active time in Sailboxes, not idle time.
Tomasz Tunguz (Theory Ventures)Read Original Article

Get AI news like this every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Next articleRobust.AI picks Aptiv's PULSE sensor for Carter robot