
Callosum, a London-based AI startup, has raised $100 million to scale Tailored Inference, a cloud service that optimizes how AI inference tasks are executed by routing them to the most efficient models and hardware chips.
The technology claims to complete certain inference tasks 3.7 times faster than GPT-5.6 Luna while cutting infrastructure costs.
A new partnership with Cerebras Systems will integrate its latest wafer-size AI accelerators into the platform.
What happened
London-based Callosum announced a $100 million funding round led by Atomico, with participation from Plural, DCVC, and the U.K. Sovereign AI Fund. This follows a $10.25 million raise in February. The company offers Tailored Inference, a cloud service that routes AI inference tasks to the most efficient models and hardware.
Why it matters
Callosum's technology completes some inference tasks 3.7 times faster than GPT-5.6 Luna while improving output quality and reducing infrastructure costs. By breaking tasks into modular blocks and assigning each to the best-suited AI model and chip, the platform addresses a core bottleneck in AI deployment—how to orchestrate compute intelligently rather than simply adding more capacity, as CEO Danyal Akarca put it.
What to watch
Cerebras Systems announced a partnership to integrate its WSE-3 Turbo wafer-size inference accelerators into Tailored Inference. The CS-4 rack appliance (powered by three Wafer-Scale Backpacks) can run models with 10 trillion parameters and is particularly suited for the decode phase of inference, complementing prefill-optimized chips from AMD and AWS that Callosum's platform also supports.
Ask the AI about this article →
Callosum enters a market where the efficiency of AI compute deployment has become a competitive differentiator. The company's funding—$100 million in this round alone, atop a $10.25 million seed—reflects investor confidence that AI workload optimization will unlock significant value as organizations struggle with the rising cost and power demands of large-scale inference. The partnership with Cerebras is particularly noteworthy because it demonstrates that chip makers themselves recognize the need for software orchestration layers; Cerebras' announcement of the WSE-3 Turbo on the same day signals that hardware innovation alone is insufficient without intelligent routing of workloads. The focus on the decode phase of inference—where models generate responses token by token—points to a specific bottleneck that both companies see as ripe for optimization, especially as enterprise deployments scale beyond initial experiments.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Broadcom is reportedly seeking to borrow up to $100 billion in debt financing to support growth efforts at Ant…
On August 12, AT&T's Chief Data and AI Officer said OpenAI models power about 25% of the telecom's total AI us…

On August 11, IBM announced a multi-year $240 million agreement with Together AI to deploy NVIDIA HGX B300 sys…

Gemini AI added 514 shares of NextEra Energy (NYSE: NEE) to its AI-run portfolio on Rallies, purchasing betwee…

Arista Networks reported Q2 2026 revenue of $3.0 billion (quarter ended June 30, 2026), while CoreWeave posted…

New Berkshire Hathaway CEO Greg Abel has quadrupled the conglomerate's stake in Alphabet (now worth $36.6 bill…
