
It's your first week on-call and you get paged at 10am. You're scrambling through runbooks, searching error messages, trying to understand dependencies in a web of microservices. After talking to a few teammates and gaining context on the system, you resolve the issue, but not before billing services went down for 15 minutes. Now management wants an RCA. The core problem isn't just the incident. It's that you had to manually hunt through logs, metrics, and traces across dozens of services to understand what happened. Modern observability generates data at a scale that makes manual analysis impractical. A single incident might involve correlating thousands of log lines, hundreds of metrics, and traces spanning 20+ services. This is a context engineering problem: How do we automatically extract relevant signals from massive telemetry datasets, understand relationships between events and services, and build actionable incident context? In this talk, we'll examine how agentic AI systems apply context engineering to observability at scale. We'll look at how these systems automatically navigate telemetry data to provide targeted, contextual information that helps SREs resolve incidents faster and write better RCAs.
Grant Griffiths is a Founding Engineer at Neubird.ai, where he leads Infrastructure and Platform Engineering. Prior to Neubird, he spent over five years at Portworx (acquired by Pure Storage), serving as Technical Lead for the Portworx CSI Driver and led Kubernetes open source community engagement. As a core contributor and reviewer in the Kubernetes community, Grant has delivered talks at KubeCon, GopherCon, and HashiTalks on topics ranging from CSI drivers to resilient data pipelines.