
On-call incidents don’t fail because teams lack dashboards. They fail because observability systems slow down under real investigative load
As telemetry volumes grow and retention windows expand, SRE teams are being asked to run deeper, broader investigations—often under time pressure—on platforms that were designed for steady-state monitoring, not bursty incident response. Tightly coupled observability stacks bind storage, compute, and query together, forcing teams to overprovision infrastructure, limit retention, or accept degraded performance during incidents.
In this talk, we’ll explore why decoupling observability architectures is becoming essential for SRE teams operating at scale. Using a real incident investigation workflow, we’ll break down how separating data storage, compute, and interaction layers allows teams to keep fast, reliable monitoring while elastically scaling investigations when incidents occur.
Ravin Trivedi is a Senior Customer Architect at Imply with over 15 years of experience in the data space, specialising in real-time analytics, distributed systems and observability. He works closely with organisations across different industries to design scalable, performant and cost-efficient data architectures. With a strong customer focus, Ravin helps break down complex technical challenges and supports data/observability teams in turning them into practical, future-ready solutions that deliver measurable business value.