
On-call is stressful enough without worrying whether AI hallucinated the root cause.
In this talk, an SRE who's been handling production incidents since before the role had a name will share what it's like to build an AI system that triages incidents and monitors system health — and why it's not as simple as throwing telemetry at an LLM.
He'll walk through how he worked with AI engineers to transfer decades of troubleshooting instincts — and where the AI deliberately diverges (like parallelizing hypothesis exploration). He'll also discuss how we build trust through evaluation — comparing accuracy, latency, and reasoning against ground truth from SREs.
This is a practitioner's view of what it takes to build an AI investigator that engineers can rely on when things are on fire at 3am.
Jason Schechner is a seasoned Site Reliability Engineer and technical leader with deep expertise across infrastructure, including compute, storage, networking, security, and automation. He currently serves as a Member of Technical Staff at Traversal, where he works on AI-driven systems for diagnosing and resolving real-world SRE incidents. Previously, he held senior SRE and production engineering roles at firms such as Teladoc Health, Citadel, The D. E. Shaw Group, and Goldman Sachs, leading large-scale, high-availability systems and infrastructure initiatives. Known for blending hands-on engineering with team leadership, he has a strong track record of building resilient platforms and guiding complex technical projects.