Agent benchmarking

From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI

From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI

Rustem Feyzkhanov discusses the critical need for companies to build private, production-aligned benchmarks for AI agents. He explains how to turn agent traces into repeatable simulations, why public benchmarks are insufficient, and how a CI pipeline for agents, integrating observability and experimentation, can ensure reliable evaluation, continuous improvement, and effective release management, moving beyond simple pass rates to measure cost, latency, and policy adherence.