
What happened
Enterprise AI buyers are now demanding performance warranties—contractual guarantees that vendors will pay if their AI model misses a defined accuracy or performance threshold—before approving deployments. Companies like Armilla AI, Munich Re's aiSure, and Mosaic have begun offering these warranties, backed by capital at risk, marking a shift from selling models as features to selling them as insurable products with measurable downside protection.
Why it matters
Once an AI system moves from generating text to making binding decisions (loans, credit decisions, tool routing, records), the cost of failure becomes a balance-sheet issue, not a quality inconvenience. A Chief Risk Officer can no longer accept "mostly right"—they need to know the maximum loss, how the system will be monitored, and who absorbs the cost when it fails. A warranty answers those questions in legal and financial terms, letting deployments clear risk committees faster than a benchmark ever could.
What to watch
The warranty market is taking shape quickly—Armilla validated MKIII's loan-default model (which evaluates over 4,000 borrower variables against over 5 million historical records) and issued a guarantee; Mosaic and Munich Re's aiSure then launched up to EUR/USD/CAD 15 million in initial capacity for developers and vendors worldwide. Vendors must now decide whether to build observable, instrumented models with predefined failure modes—the only ones underwriters will back—or stay stuck selling unwarrantable black boxes.
Summaries like this, in your inbox every morning.
The enterprise AI market has stalled not because the technology fails to impress, but because buyers face a question that benchmarks cannot answer: who pays when the model is wrong? For three years, the industry treated AI failure as a disclaimer bundled with productivity tools—a flawed email gets corrected, a bad summary gets re-read. But the market has fundamentally shifted. Modern AI systems now authorize transactions, assess credit, deploy software, and change enterprise records. Once a model crosses that line from generating information to creating lasting external-state change, "mostly right" becomes a balance-sheet issue, not a quality issue. A loan model that drifts reshapes a credit portfolio; an agent that routes incorrectly creates an expensive obligation; a quality-control system that misclassifies a defect turns into returned product or regulatory exposure.
This is why Chief Risk Officers have stalled deployments: the technology clears the pilot, the benchmark shows promise, but the risk committee still has no contractual answer to the core question—how large can the loss become, how will we know the system has crossed the line, and who pays? A demo doesn't answer that. A benchmark is a snapshot of how a model performed on a selected evaluation set at one moment; it says nothing about production data, distribution shift, or a version update six months out. The gap between the polished proof-of-concept and the economically accountable deployment is where enterprise AI has been stuck.
The market's emerging answer is a performance warranty—a narrow, powerful instrument that converts a technical performance promise into a financial obligation. Armilla's November 2025 partnership with MKIII demonstrates the model in credit, where loan defaults are measurable and high-stakes. Mosaic and Munich Re's aiSure then announced up to EUR/USD/CAD 15 million in capacity, signaling that a real commercial category is forming. What matters most: a warranty forces vendors and buyers to stop speaking in adjectives—accurate, trustworthy, responsible—and start speaking in measurable terms. It requires observable models, predefined failure modes, versioned decisions, and agreed ground truth. Companies that built this instrumentation early will clear underwriting; companies that bolt it on will discover their product cannot be priced.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic released Claude Opus 5.5 today and cut its price 20%, with input at $4 per million tokens and output…
Firecrawl announced it has raised $75 million in a Series B round led by Smash Ventures, with participation fr…
ASRock is shifting its business focus toward AI

Anthropic and OpenAI, which spent early September warning that model capabilities are outrunning the safeguard…

At a subscriber event, MIT Technology Review's Will Douglas Heaven and Grace Harkins said AI wiping out humani…

Jessica Wachter of Wharton and co-authors estimate hyperscaler spending will reach nearly 1.1兆ドル by 2027, and…
