SREday

Site Reliability, DevOps and Cloud

February 26, 2026 Harness, New York, US

1
Day
10+
Speakers
1
Track
80+
Attendees

What Makes an AI SRE Trustworthy? An SRE’s Take

Jason Schechner
Traversal
Abstract

On-call is stressful enough without worrying whether AI hallucinated the root cause.

In this talk, an SRE who's been handling production incidents since before the role had a name will share what it's like to build an AI system that triages incidents and monitors system health — and why it's not as simple as throwing telemetry at an LLM.

He'll walk through how he worked with AI engineers to transfer decades of troubleshooting instincts — and where the AI deliberately diverges (like parallelizing hypothesis exploration). He'll also discuss how we build trust through evaluation — comparing accuracy, latency, and reasoning against ground truth from SREs.

This is a practitioner's view of what it takes to build an AI investigator that engineers can rely on when things are on fire at 3am.

Bio

Jason Schechner is a seasoned Site Reliability Engineer and technical leader with deep expertise across infrastructure, including compute, storage, networking, security, and automation. He currently serves as a Member of Technical Staff at Traversal, where he works on AI-driven systems for diagnosing and resolving real-world SRE incidents. Previously, he held senior SRE and production engineering roles at firms such as Teladoc Health, Citadel, The D. E. Shaw Group, and Goldman Sachs, leading large-scale, high-availability systems and infrastructure initiatives. Known for blending hands-on engineering with team leadership, he has a strong track record of building resilient platforms and guiding complex technical projects.

Video

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one