
Enterprise leaders speaking at VB Transform 2026 revealed that evaluating AI agents by scoring individual conversations—even if each looks perfect—fails to catch broken products. Companies are shifting toward comparing cohorts of users against a baseline and adopting cheaper, narrower judge models, though the industry still struggles to balance scalable but ungrounded automation with labor-intensive human review.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Leaders from LangChain, Conviva, and CoreWeave said at VB Transform 2026 that enterprises are moving away from scoring individual AI agent conversations in isolation—which can appear flawless but mask broken products—toward comparing groups of users against a baseline. The industry is also adopting cheaper, narrower judge models instead of larger ones.
Why it matters
A single conversation that looks perfect on its own does not tell you whether your AI product actually works for real users. This shift means enterprises can now catch systemic problems that individual-trace scoring misses, improving how well AI agents perform in production.
What to watch
The tension between automated judging (whether by AI agent or large language model) and human review remains unresolved; Zhang noted that automated methods are scalable but hard to ground in real outcomes, while human review is thorough but does not scale.
At VB Transform 2026, three enterprise technology leaders addressed a critical blind spot in how companies assess AI agent quality. Harrison Chase (CEO of LangChain), Hui Zhang (CTO and co-founder of Conviva), and Emmanuel Turlay (director of engineering at CoreWeave) described a fundamental shift in enterprise evaluation practices. The problem they identified is deceptive: a single AI agent conversation can appear flawless when scored in isolation and still mask a broken product. This gap is reshaping how organizations think about AI quality assurance, moving away from individual-trace scoring and toward comparing cohorts of users against a baseline. Alongside this methodological shift, enterprises are adopting cheaper, narrower judge models—specialized AI models trained to evaluate agent output—rather than relying on large, general-purpose language models. However, the industry faces an unresolved tension between the two main judging approaches. Agent-as-judge (using one AI agent to evaluate another) has not replaced LLM-as-judge, which Chase noted remains the default. The deeper challenge, Zhang explained, is the contradiction between scalability and grounding: automated judging—whether by agent or LLM—can evaluate outcomes at scale but struggles to anchor those evaluations in real-world truth. Human review, by contrast, can ground verdicts in observed reality but cannot scale to the volume of conversations production systems generate. Zhang summarized the dilemma: "You have scalable but ungrounded, whether it's agents as judge or LLMs as judge, you grade the outcome, you grade the work. It still is very difficult to ground it and then you use humans and that's just not scalable."
The core tension revealed at VB Transform 2026 reflects a maturing realization in enterprise AI deployment: a polished single interaction does not validate a product. Scoring individual traces—the traditional approach—creates a false sense of security. Harrison Chase, Hui Zhang, and Emmanuel Turlay highlighted a fundamental trade-off in AI evaluation: automated judging (whether by agent or large language model) offers scalability but struggles to ground its verdicts in real user outcomes. Human review, by contrast, can root findings in ground truth but does not scale to the volume of conversations modern applications generate. The shift toward cohort comparison and cheaper judge models represents a pragmatic middle ground—replacing expensive, monolithic evaluation with lighter-weight but distributed baselines that capture product health across populations rather than polishing individual cases.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion




Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack