Why do enterprise implementations of SLOs fail? The failure is rarely a cancellation or a reversal in direction. It usually masquerades as something more sinister — lip service and minimal effort, where outcomes get sacrificed for the sake of output, and the difference that could have been made for our teams and customers quietly never arrives. What makes it hard to catch is that the early part goes well. Dashboards appear, percentages get published, coverage gets reported, and none of it is fake. Then the program never lifts off the runway, or it does and it's carrying nothing. This session walks through several failure patterns I've seen first hand, and what each one costs you. Some are corrosive. Some just leave a better outcome on the table in exchange for something expedient. All of them are avoidable once you can see them for what they are. These aren't hypotheticals — they come from multiple years driving SRE adoption, culture, and practice across different organizations, and I've been in the middle of every one of them. Before we can fix something we first have to be able to recognize it for what it is and then articulate why it's a pattern to be avoided. If you see yourself, your team, or your organization in any of this — good. You're not alone.
With over a quarter-century spent in the trenches of production systems, Jason has navigated everything from network engineering and cloud architecture to large-scale SRE programs and teams. He has spent years driving reliability and resiliency for mission-critical infrastructure within Fortune 500 landscapes, leading engineering efforts across both the U.S. and globally. His focus spans all three domains of reliability success: People, Processes, and Platforms. One driving truth: the cultural layer is just as critical as the compute layer.