AIToday
AI Business & IndustryLarge Language ModelsSiliconANGLE AIPublished: Sep 5, 2026, 10:02 JST2 min read

Gimlet Labs raises $300M for disaggregated AI inference

Gimlet Labs raises $300M for disaggregated AI inference

Key takeaway

  • Gimlet Labs raised $300M at a $3B valuation.

  • Its platform speeds up AI inference by splitting models across chips.

  • Backers include top investors and Arm, Samsung, and Microsoft funds.

3 Key Points

  1. What happened

    Gimlet Labs Inc., a startup that helps developers speed up inference workloads, raised $300 million in funding at a $3 billion valuation. Andreessen Horowitz led the Series B round, joined by Arm Holdings Inc., Samsung Ventures, Microsoft Corp.'s M12 fund, and others.

  2. Why it matters

    The company's software automatically breaks an LLM (AI that understands and generates text) into modules and places each on the chip best suited to it, which can speed up the process. Customers include one of the world's largest cloud providers and a top three AI lab.

  3. What to watch

    The funds will expand the serverless platform's infrastructure by several hundred megawatts of computing power. Gimlet also plans to enter custom hardware, building an inference-optimized server without a motherboard for use outside data centers.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

This round comes as AI developers face growing pressure to cut the cost and latency of running large models in production. Gimlet’s approach addresses a core inefficiency: an LLM is not a single uniform workload but a set of parts with very different hardware demands, so running the whole thing on one chip type can waste resources. By disaggregating the workflow and matching each piece to a specialized accelerator, the platform aims to make inference faster and cheaper—an outcome that could be significant for businesses that rely on AI at scale.

The company's method includes both the common PD disaggregation (splitting prefill and decode phases) and more granular splits, such as dividing the decode phase into smaller tasks or using a lightweight 'drafter' model alongside a larger one. Its software optimizes each module for its target chip using AI agents and a custom compiler. The funding will let Gimlet expand its serverless capacity by several hundred megawatts and move into custom hardware, indicating the company is betting on continued demand for specialized inference infrastructure.

FAQ

What does Gimlet's platform do?
It automatically breaks an LLM into modules and deploys each on the chip architecture that best fits its hardware needs, potentially speeding up inference. For example, memory-heavy parts go to accelerators with large onboard RAM.
How much funding has Gimlet raised in total?
After this round, the company's total outside funding stands at $392 million.
What will the new funding be used for?
It will support growing the serverless infrastructure by several hundred megawatts and expanding into custom hardware with an inference-optimized, motherboard-less server.
SiliconANGLE AIRead Original Article

Get the latest AI Business & Industry news every morning

For example, today's edition would include:

  • JM-Applied order book tops NT$3B on AI gas gear demandDIGITIMES Asia · 16m ago
  • OpenAI says GPT-6 Astra 'low' beats GPT-5.6 Sol 'high'ITmedia AI+ · 16m ago
  • Nvidia's $99B Portfolio Puts Intel, CoreWeave to the TestYahoo Finance AI · 16m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleBEXCO: South Korea's Safety AI Is Top-Down