
On-call incidents don’t fail because teams lack dashboards. They fail because observability systems slow down under real investigative load.
As telemetry volumes grow and retention windows expand, SRE teams are being asked to run deeper, broader investigations—often under time pressure—on platforms that were designed for steady-state monitoring, not bursty incident response. Tightly coupled observability stacks bind storage, compute, and query together, forcing teams to overprovision infrastructure, limit retention, or accept degraded performance during incidents.
In this talk, we’ll explore why decoupling observability architectures is becoming essential for SRE teams operating at scale. Using a real incident investigation workflow, we’ll break down how separating data storage, compute, and interaction layers allows teams to keep fast, reliable monitoring while elastically scaling investigations when incidents occur.
We’ll connect these patterns to lessons learned in other data-intensive systems, but stay grounded in the day-to-day realities of on-call life: faster root cause analysis, fewer tradeoffs during incidents, and observability systems that hold up when you need them most.
Eric is Chief Architect at Imply and a driving force behind Imply Lumi, an Observability Warehouse. His work focuses on keeping more data searchable at a lower cost while accelerating investigations with no workflow disruptions. Eric is one of the original authors of the open source Apache Druid® project. Eric's expertise spans his roles as a Fellow at Splunk, Distinguished Engineer at Yahoo Inc, member of the founding team at Tidepool, among other roles.