
Every team has an engineer who makes on-call look easy. They know the service that quietly degrades every Monday morning, the log pattern that precedes a cascade, and which alert has been firing incorrectly since the 2023 migration. When they leave, MTTR doubles and not because anything broke, but because the knowledge that made fast debugging possible was never captured. This talk is about making incident expertise durable. We cover how to extract the decision patterns that live in senior engineers' heads, encode them into investigation workflows that surface automatically during incidents, and measure whether the transfer is working. We walk through what worked, what created new toil (the "automated runbook" trap), and one incident where the system suggested the wrong root cause and why that was still a net win and more
Pratik has been in the observability space for more than a year now. With a background spanning SRE and platform engineering, he focuses on OpenTelemetry in production, Kubernetes fleet operations. He is an active open source contributor and member at Opentelemetry.