
Agent chats can score perfect and still fail: LangChain, Conviva, CoreWeave warn executives
Evaluation needs cohorts, baselines, and monitoring, not just “looks good” trace scoring and cost-heavy judge models.
By Yousef Al-Zahrani·· 4 min

Curating from trusted global sources…
2 briefings · “vb transform 2026”

Evaluation needs cohorts, baselines, and monitoring, not just “looks good” trace scoring and cost-heavy judge models.

At VB Transform 2026, Zillow’s engineering chief laid out how persistent context beats raw data for agentic AI ROI.