Expedia’s Xavi Amatriain says evals are the new PRD, not afterthoughts
His playbook: encode goals, security, and red-teaming into evals before coding, then enforce governance with risk-based toll gates.

Xavi Amatriain, Expedia Group’s first chief AI and data officer, told VB Transform 2026 that evals should replace the traditional PRD workflow. For decision-makers, this reframes AI governance as a design-time system problem, where feedback loops and risk-based checkpoints determine whether agents behave in the real world.
At VB Transform 2026 in Menlo Park, Expedia’s first chief AI and data officer, Xavi Amatriain, delivered a line that should reset how most teams write AI requirements: “The new PRD are the evals.” His point is simple and sharp. You do not wait until after you start coding to figure out what “good” looks like. Instead, you “encode what you want the product to do through your evals,” including red teaming evals and other security requirements, and you embed that into the PRD before any real implementation work begins.
Amatriain then pushed it further, essentially arguing that evals are becoming the place where thinking goes to be tested. “With AI-assisted or AI-generated code, that’s gonna be the future. It’s like all your thinking is gonna go into the evals.” The operational consequence is huge: when production deployment is increasingly automated, the only way to keep quality and safety from drifting is to define success, failure, and adversarial conditions up front, and to continuously feed what happens in the wild back into the evaluation suite.
This matters because the stakes around evaluation are not theoretical. VentureBeat’s VB Pulse research on the evaluation gap surveyed 157 enterprises and found that 66% already allow some production deployment without human review or plan to do so within the next 12 months. But only 5% fully trust the automated evaluations that would make that decision. The mismatch is uncomfortable for leadership teams: half of respondents have shipped an agent that passed internal evals but then failed with a real customer. That is the exact failure mode Amatriain is trying to prevent by treating evals like a first-class product design artifact, not a late-stage QA checkbox.
A big part of his argument is about guardrails. Amatriain warned that adding too many brittle constraints can backfire by distorting learning from user feedback. “The more guardrails and artificial business rules and sort of rules that you put into the system, the worse off,” he said. “Not only because they’re brittle, but also because they actually mess up with the feedback loop. You are actually biasing the user and the feedback you get from the user, and then you’re learning that in the wrong way.” He called guardrails “a necessary evil,” with the goal of minimizing their impact over time.
Still, not everyone at the conference agreed on how “necessary” that evil should be. Other speakers argued that the highest-risk actions still require very firm guardrails. Expedia’s response is a structured governance approach built around three layers. First are principles, communicated broadly. In Amatriain’s framing, principles are meant to guide distributed decision-making, because in large organizations, the reality is that “those principles” often do not automatically embed into culture. Principles are followed by processes and tools that enforce them, because “principles look really nice on a picture on some wall, but you need to then give them teeth.” On top sits automation, calibrated to risk.
That’s where Expedia’s “agent release toll gates” come in. These are checkpoints adjusted to the risk level of each agent release, with evaluation rounds, red teaming, and security review tied to that risk. “Governance needs to correlate to the risk,” Amatriain said. If something is low risk, you do not need excessive governance. If the stakes are high, the checks move from recommended to required as risk increases. In practice, that also means governance is no longer a generic blanket process that slows everyone down equally. It becomes something teams can actually design for.
Expedia’s architecture reinforces the same philosophy. Amatriain told the audience he does not buy the idea of AGI as a single unified model or “a singleton.” “I think it’s much better to think of it as composition,” he said, describing specialized agents that are composed into a system. Expedia builds at the component level: tools become skills, skills become sub-agents, and sub-agents get orchestrated into the full agentic system. He argued that principles have to unify how the system behaves, including tone, how the user is addressed, and how context and memory are handled. And he framed security as easier when teams scope narrowly: individual agents can be evaluated and “lock[ed] down” before composing them into the larger system.
He illustrated how design-time decisions change both reliability and safety with Expedia’s travel use cases. Real time pricing, flight availability changes minute to minute, and hotel reviews routinely contradict supplier claims. For this, Amatriain described a system that blends retrieval-augmented generation with direct API tool calls, choosing the approach based on latency. For example, if a user asks how much a four-star hotel usually costs in Chicago in July, the agent should answer immediately because the result can be cached and does not need real-time data. But for a specific hotel justification, like whether reviews contradict supplier claims about a pool, the agent needs more than what a supplier self-reports. He described cross-referencing Expedia’s own review corpus to avoid the “generic chatbot” failure mode.
At the same time, he drew a non-negotiable boundary around agency. “We don’t want the agent to book the hotel or to buy you a plane ticket for you,” Amatriain said. The user has to hit the click. That constraint is not just a product stance, it is also positioned as a security decision: once those design principles are established, he argued, you do not have to bolt guardrails on after the fact.
And that brings us to what he called the next attacker: not just humans, but other AI systems. “Security needs to be a principle that is shifted as left as possible and as part of the design itself,” he said, adding that needing guardrails usually means the team “has not thought about it early on.” In production, Amatriain described a feedback loop where monitoring signals flow back into the eval suite, with the goal of automating as much of that cycle as possible. “But having that whole feedback loop from real signals, from your operating AI system, all the way into being reported and fixed as quickly as possible is going to become essential.”
VentureBeat’s separate June Pulse survey on agent security underscores why this timing matters. From 107 enterprises, 54% reported an agent security incident or near-miss already. 59% plan to adopt, add, or replace agent security tooling within 12 months, and 29% plan to move this quarter. Incident rates climbed with organization size, reaching 63% for enterprises with more than 1,000 employees versus 49% for companies with 101 to 1,000. And sandbox isolation, a post-breach control that limits damage, dropped from 35% adoption at smaller companies to 20% at the largest. In Amatriain’s view, threats will increasingly come from other external agentic systems poking at everything, making not only detection but time to fix essential.
So here is the real takeaway for executives: if evals are the new PRD, then governance is no longer a policy document. It is how you design your product’s definition of “correct,” “safe,” and “resilient,” before you generate code and compose agents. And because trust in automated evaluations is currently thin, the boards and leadership teams that win will be the ones that build evals, red teaming, and feedback loops into the workflow early enough to prevent internal success from turning into customer failure.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Business

Anthropic’s Levant Alpöge cracks the Jacobian conjecture after 87 years
A Harvard valedictorian used Claude to hit a 1939 breakthrough, but the missing “why” is the real problem.

Uber buys Delivery Hero for nearly $15B, vaulting to top food delivery outside China
The deal doubles Uber's dual-services footprint and pushes a ride-and-eats bundling play into 50 more markets.

Epic and Google drop settlement bid, forcing rival Android app stores by July 22
Google told the court it is ready to carry third-party app stores starting Wednesday, July 22.

