
Gimlet Labs raised $300M at a $3B valuation.
Its platform speeds up AI inference by splitting models across chips.
Backers include top investors and Arm, Samsung, and Microsoft funds.
What happened
Gimlet Labs Inc., a startup that helps developers speed up inference workloads, raised $300 million in funding at a $3 billion valuation. Andreessen Horowitz led the Series B round, joined by Arm Holdings Inc., Samsung Ventures, Microsoft Corp.'s M12 fund, and others.
Why it matters
The company's software automatically breaks an LLM (AI that understands and generates text) into modules and places each on the chip best suited to it, which can speed up the process. Customers include one of the world's largest cloud providers and a top three AI lab.
What to watch
The funds will expand the serverless platform's infrastructure by several hundred megawatts of computing power. Gimlet also plans to enter custom hardware, building an inference-optimized server without a motherboard for use outside data centers.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
This round comes as AI developers face growing pressure to cut the cost and latency of running large models in production. Gimlet’s approach addresses a core inefficiency: an LLM is not a single uniform workload but a set of parts with very different hardware demands, so running the whole thing on one chip type can waste resources. By disaggregating the workflow and matching each piece to a specialized accelerator, the platform aims to make inference faster and cheaper—an outcome that could be significant for businesses that rely on AI at scale.
The company's method includes both the common PD disaggregation (splitting prefill and decode phases) and more granular splits, such as dividing the decode phase into smaller tasks or using a lightweight 'drafter' model alongside a larger one. Its software optimizes each module for its target chip using AI agents and a custom compiler. The funding will let Gimlet expand its serverless capacity by several hundred megawatts and move into custom hardware, indicating the company is betting on continued demand for specialized inference infrastructure.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
JM-Applied, a Taiwanese semiconductor gas equipment maker and supply chain member for Micron and TSMC, said on…

OpenAI's product lead Tibo Sotiou posted on X on September 6 that GPT-6 Astra's 'low' setting outperforms GPT-…

Nvidia disclosed roughly $99 billion of public and private equity investments as of July 26, plus about $25 bi…

In January, Ukraine's defense ministry said it would share millions of data points from tens of thousands of d…

OpenAI Group PBC acknowledged it did not publicly disclose an episode where its AI agents wrote to outside web…
Eaton is expanding beyond traditional power management into modular power deployment, next-generation DC conve…
