AIToday

Expedia's AI Chief: Evals Are the New Product Design

VentureBeat AI6h ago
Expedia's AI Chief: Evals Are the New Product Design

Key takeaway

Expedia's chief AI officer told industry leaders at VB Transform 2026 that evaluation frameworks—not traditional product design documents—are becoming the core way companies specify what AI products should do. As AI-generated code takes over, the strategic thinking shifts into building rigorous evals that encode security, functionality, and red teaming upfront, with production deployment decisions increasingly tied to evaluation results.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Xavi Amatriain, Expedia Group's first chief AI and data officer, told the VB Transform 2026 audience that product requirements are shifting: companies now encode what they want AI products to do through evaluations (evals) rather than traditional design documents, with evals including red teaming and security requirements built in from the start.

  • Why it matters

    As AI-assisted or AI-generated code becomes the future, the thinking and design work moves upstream into evals rather than into code itself. This represents a fundamental change in how AI product teams structure their work—the evaluation framework becomes the actual product specification.

  • What to watch

    VentureBeat's VB Pulse research found that 66% of the 157 enterprises surveyed already permit some production deployment tied to these evaluation frameworks, signaling broad industry movement toward this model.

In Depth

At VB Transform 2026, held in Menlo Park, Xavi Amatriain, Expedia Group's first chief AI and data officer, articulated a fundamental rethinking of product development in the age of AI-generated code. "The new PRD are the evals," he stated, referring to product requirement documents—the traditional cornerstone of product planning. Instead, Amatriain explained that companies now encode product intentions through evaluation frameworks, which might include red teaming evals and other validation mechanisms, with security requirements already embedded before coding begins.

He pushed the reasoning further: "With AI-assisted or AI-generated code, that's gonna be the future. It's like all your thinking is gonna go into the evals." This signals a shift in where strategic product work happens. Traditionally, product managers and engineers write specifications, then engineers code against those specs. In an AI-generated-code world, the thinking must move upstream into the evaluation framework itself, because the code-writing step is no longer the primary locus of human decision-making.

Amatriain brings significant credibility to this observation. Before joining Expedia in December 2025, he served as VP of AI and Compute Enablement at Google, where he oversaw the AI and compute infrastructure powering Gemini and Google Search. He has also mentored engineers who went on to found Perplexity and Scale AI, companies working at the frontier of AI evaluation and data infrastructure respectively.

VentureBeat's VB Pulse research lends empirical weight to Amatriain's assertion. Of 157 enterprises surveyed, 66% already permit some form of production deployment in conjunction with evaluation frameworks, indicating that the shift from traditional PRDs to evaluation-driven specification is not theoretical but already underway across industry.

Context & Analysis

Xavi Amatriain's remarks at VB Transform 2026 capture a significant shift in how enterprises approach AI product development. Rather than writing detailed product requirements first and then building to specification, teams are now encoding expectations into evaluation frameworks—sets of tests that check whether an AI system behaves as intended, including security and red teaming scenarios. This inversion of the traditional product development workflow reflects the reality that AI-generated or AI-assisted code is becoming the delivery mechanism; if the code is generated, the human creativity and specification work must happen earlier, in the design of the evaluations themselves.

Amatriain's background at Google, where he led AI and compute infrastructure for Gemini and Google Search, positions him to recognize this pattern across large-scale AI deployments. His point is not merely theoretical: VentureBeat's research indicates that two-thirds of surveyed enterprises have already begun permitting production deployment decisions based on evaluation results, suggesting the industry is actively adopting this model. The stakes are concrete—security requirements, functional correctness, and product behavior are all being embedded into evals before any code is written, collapsing the traditional gap between specification and implementation.

FAQ

What is Xavi Amatriain's background before Expedia?
Amatriain served as VP of AI and Compute Enablement at Google across the platforms powering Gemini and Google Search before his December 2025 appointment at Expedia. He has also mentored talent who went on to found Perplexity and Scale AI.
What does the VB Pulse research show about enterprise adoption?
Of the 157 enterprises surveyed, 66% already permit some production deployment in conjunction with evaluation frameworks.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →