
AI Agents are non-deterministic, tool-using systems, so typical unit tests and quality monitoring frameworks are not up to the task. In this talk, I will share a practical, end-to-end framework for evaluating our AI agents at PagerDuty, both offline and online.
Murilo Venturin is a Machine Learning Engineer specializing in AI agents, LLMs, and generative AI systems. At PagerDuty, he designs and deploys multi-agent AI architectures, integrates reasoning workflows, and builds evaluation frameworks to measure performance of generative AI in production.
Ricardo Moreira is an AI Engineer and Full-Stack Developer specializing in LLMs, deep learning, and scalable AI systems. He is a Senior Applied Scientist at PagerDuty, where he built the company’s unified AI agent testing and evaluation framework.