
Production incidents don't respect role boundaries. What looks like an infrastructure failure often requires reading code. What looks like an application bug often requires understanding distributed system architecture. The tooling shift toward traces, release tracking, and code-level observability hasn't just made incidents easier to diagnose - it's quietly collapsed the boundary between what an SRE needs to know and what a developer needs to do. This session is a tactical field report on how tooling is collapsing traditional role boundaries, what that looked like in practice across two real incidents, and what it means for the evolving SRE skillset. Both incidents illustrate the same dynamic playing out in different directions — and what we've learned since. But the more interesting story is what happened after. When on-call responsibility spreads beyond a single ops team, something shifts - and it points somewhere we think is inevitable. This session is for SREs who are already crossing the boundary whether they mean to or not - and for anyone thinking about what it means to build reliable systems when the line between "who owns this" and "who can fix this" keeps moving
Khanh Nguyen is an Engineering Leader at Sentry, focused on infrastructure and reliability platforms. Based in San Francisco, he leads teams building scalable systems that improve performance and operational resilience.