Amazon's AI director says the bottleneck stopping enterprises from deploying AI agents is not raw capability but reliability—the ability to perform consistently, robustly, predictably, and safely in real-world conditions.
While 85% of enterprises are testing AI agents, only 5% have moved them into production, a gap he attributes to agents acing internal benchmarks but failing when customers actually use them.
What happened
At VB Transform 2026, Bryan Silverthorn, Director of AGI Autonomy at Amazon, told the audience that enterprise adoption of AI agents is stalled not by capability gaps but by reliability problems. He introduced a framework breaking reliability into four dimensions—consistency, robustness, predictability, and safety—credited to Princeton research.
Why it matters
Cisco data shows 85% of enterprises are piloting AI agents, yet only 5% have shipped them to production. Agents frequently pass internal benchmarks but fail when deployed to real customers, suggesting that traditional performance metrics are missing what actually matters for business deployment.
What to watch
Silverthorn, who joined Amazon through its 2024 acquisition of Adept AI, leads multimodal agent training in Amazon's AGI lab. His four-dimensional reliability framework may reshape how enterprises and vendors evaluate readiness for production use.
Ask the AI about this article →
The enterprise AI sector faces a paradox: broad experimentation without production traction. Cisco's finding that 85% of enterprises pilot AI agents but only 5% deploy them to production reflects a critical disconnect between lab performance and real-world reliability. Silverthorn's framing addresses this head-on by arguing that the problem is not benchmarking sophistication—it is that traditional evaluations measure narrow capability (how well an agent solves a test problem) without measuring the four dimensions that actually determine whether an agent will work reliably when a customer depends on it. His four-part framework—consistency (does it behave the same way each time?), robustness (does it handle edge cases?), predictability (can you anticipate its failures?), and safety (does it avoid causing harm?)—reorients the conversation from "how smart is it?" to "can I trust it?" This distinction is significant because it implies that vendors and enterprises may be measuring success on the wrong metrics, explaining why agents pass internal evals and then "collapse in the wild." Silverthorn's background—he came to Amazon through the acquisition of Adept AI, a startup that built autonomous AI agents—suggests Amazon is embedding this thinking directly into its AGI development strategy.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Phonely Ltd. launched Alma, a large language AI model built for voice agents and trained on over 10 million re…
Aranya Inc., a startup founded last year, launched today with $11 million in funding
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Sarah O’Connor's book 'We Are Not Machines' explores how mechanization and AI have transformed the workforce…
