SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.
Companies presenting:
CANAL+ Group, Contentful, Criteo, Enix, ewake, Gatling, Imply, kladriva, LoopStudio, OCTO Technology, PagerDuty, Pruna AI, Red Hat, Spitzkop, vCluster
Production is a wilderness. The terrain shifts every minute, meaning human and agent failures are not caused by missing intelligence. They come from acting on stale context.
Discovery takes time, and in production it is often the most expensive part of the workflow.
In this talk, I show why AI SRE agents fail when they rely on runtime discovery, and what changes when they operate on live context instead of snapshots.
The short version: an AI SRE that admits uncertainty is the only kind you can safely let run.... Read more
On-call incidents don’t fail because teams lack dashboards. They fail because observability systems slow down under real investigative load
As telemetry volumes grow and retention windows expand, SRE teams are being asked to run deeper, broader investigations—often under time pressure—on platforms that were designed for steady-state monitoring, not bursty incident response. Tightly coupled observability stacks bind storage, compute, and query together, forcing teams to overprovision infrastructure, limit retention, or accept degraded performance during incidents.
In this talk, we’ll explore why decoupling observability architectures is becoming essential for SRE teams operating at scale. Using a real incident investigation workflow, we’ll break down how separating data storage, compute, and interaction layers allows teams to keep fast, reliable monitoring while elastically scaling investigations when incidents occur.... Read more
SRE used to be about keeping systems alive under scarcity: limited machines, limited telemetry, limited deployment safety. Now the risk is abundance: too much cloud, too much telemetry, too many abstractions, too much automation and too many AI systems acting faster than humans can understand.... Read more
AI developers want a GPU the way they want a kind cluster: kubectl apply and it's there, kubectl delete and it's gone, no console in between. The cloud default punishes that: provisioning latency, sticky nodes, per-tenant security setup. We give developers a kubectl context with an auto-provisioning GPU under it: a T4 or L4 in ~3 minutes, gone in ~30 seconds, zero infra to touch.
Live demo: a dev creates a GPU-template tenant cluster, applies a pod, and a node auto-provisions in GCP. The pod runs, is deleted, the node disappears 30s later. Under the hood is a four-piece contract: tenant Kubernetes for isolated RBAC and node pools; Karpenter for auto-provisioning; an idle policy (consolidateAfter 30s) collapses nodes on idle, not on cluster delete; and a tunnel keeps the experience kubectl-native, not console-native.
We close with the failure modes (Karpenter template/provider drift, NodeProvider creds silently disabled) and the GitOps boundary that keeps multi-team setups intact.... Read more
Running ML workloads on Kubernetes can be a burden to manage, especially when the usages tend to diversify. Batch inference will involve different constraints as experimentation, evaluation or serving ML models for user-facing applications.
* How to assign the correct GPUs for each usage ?
* How to scale differently workloads with flexibility ?
* How to efficiently scale down to 0, when no GPUs are needed ?
To answer those questions, we will see in this talk how combine the standard orchestration approach of Argo Workflows with the autoscaling super powers of Karpenter, a tool developed by AWS and given back to the community. We'll dig into the difficulties of scaling GPUs and how Karpenter solves them, with a return of experience of Pruna's computing platform.
... Read more
AI has dramatically reduced the complexity of starting a new code project. Today, you don't read the docs first - you ask your coding agent.
In this talk, we'll show how this applies to performance testing. How developers with zero Gatling experience can write realistic load tests in minutes, using AI as their entry point instead of documentation, and why this changes how teams adopt performance testing altogether.... Read more
Unexpected production errors often come from a familiar place: shared environments that fail to accurately validate real-world conditions.
Ephemeral Environments (EEs) offer a different approach. Instead of relying on fragile, long-lived shared environments, teams can validate each change in an isolated, production-like environment that includes code, infrastructure, and configuration.
In this talk, we’ll explore how EEs reduce unexpected issues in production by improving quality validation, eliminating cross-team interference, enabling safer automated testing, and accelerating deployments.
But prevention is only part of the story.
We’ll also touch on how the same environments can be used after an incident to reproduce failures, reconstruct system state, and support deeper root cause analysis and post-mortems.
This session focuses on how SRE teams can use an incremental and adaptable ephemeral strategy to connect prevention and incident response into a continuous reliability loop.
... Read more
It’s 3:45 AM, and you get paged. Still half asleep, you reach for your laptop. Turns out something is wrong, so you acknowledge the incident and embark on a journey to figure out what is happening, mobilize the right people, and eventually fix the issue. But what if we could make it easier? In this talk, I’ll show how AI agents can help you fast-track triage, coordinate responses, debug problems, and run fixes. To wrap up, I will show how we can go even further and move from merely reacting to incidents to proactively preventing them. In the spirit of shifting left, let’s understand how we can track future incidents our code may cause, right before we commit it!... Read more
Every regulated org has the same scar: a beautiful governance tool - a CMDB, a compliance tracker, a project portal - that describes what production should look like, sitting next to a real platform that actually runs something else. Two sources of truth. Manual sync. Tickets. Drift.
The compliance team trusts the tool; the SREs trust the cluster; nobody trusts the gap between them. This talk is a field report on closing that gap. Drawing on a teardown of a large enterprise platform and the sovereign platform we're building at SPITZKOP, I'll show how to make the control plane connected to the execution plane: GitOps as the single source of truth, policy-ascode (Kyverno) as the admission backstop, schema-first config generation (CUE) shifting compliance left, and a conformity engine that turns regulatory requirements (NIS2, DORA, EUCS) into executable, attested deployment steps — not Excel rows.
You'll leave with a concrete pattern for making compliance a property of your pipeline, not a parallel universe.... Read more
As organizations process ever larger volumes of real time data, building a reliable streaming platform becomes a significant engineering challenge. In this session, Qinghui Xu shares Criteo’s experience operating Apache Flink at scale, covering the architectural decisions, operational lessons, and platform practices that enable large scale stream processing in production. Attendees will gain insight into the challenges of performance, reliability, and observability in distributed streaming systems, along with practical takeaways for teams running critical data workloads on Apache Flink.
... Read more
In this session, I’ll walk through a production rollout system that combines versioned deployments, gradual traffic ramps, error-budget-aware gates, and automatic rollback on the Cloudflare Developer Platform.
We’ll cover the hard parts: defining trustworthy signals, preventing flappy rollbacks, and handling emergency bypass paths safely.
You’ll leave with concrete patterns to increase deploy frequency while reducing risk and MTTR.... Read more
kube-image-keeper [kuik, https://github.com/enix/kube-image-keeper](https://github.com/enix/kube-image-keeper) makes sure that your workloads keep starting even when the upstream registry is down, Docker Hub rate-limits you, or an image tag gets deleted. It’s an open-source Kubernetes image resiliency solution that provides transparent image routing across multiple registries, as well as selective caching and replication of the images used by your workloads into local or external registries. On 100+ production clusters it has saved us countless incidents during registry outages and rate-limit storms, and made node autoscaling actually reliable. The talk covers the concrete usecases it solves, what we got wrong in v1, why we rewrote it as v2, and the edge cases that bit us along the way: mutating webhooks, garbage collection, and keeping the cache honest at scale.... Read more
IT certifications are everywhere in today’s industry, especially around cloud, Kubernetes, Linux, security, and DevOps. Some engineers see them as valuable learning tools and career accelerators, while others consider them disconnected from real production work.
In this talk, I will share a practical and balanced perspective based on both certification experience and real-world production environments. We will discuss what certifications are actually good for, where they fall short, and how engineers can use them effectively without relying on them blindly.... Read more
FinOps Patterns for Kubernetes in Production Running Kubernetes on AWS doesn't automatically mean running it efficiently. Most teams focus on stability and delivery and discover their cloud bill three months later, when the damage is done. In this talk, I share concrete lessons from optimizing EKS infrastructure costs in production: what to measure, where the waste actually hides, and which levers move the needle, without breaking anything. We'll cover: 1. Why standard cloud cost tools (Cost Explorer, Kubecost) only give you half the picture. The underestimated impact of Karpenter, instance diversification, and Spot strategy on real workloads. 2. How right-sizing Kubernetes requests saved more than any Reserved Instance commitment. 3. The FinOps discipline that stuck vs. the optimizations that quietly regressed No vendor pitch. No theoretical framework. Just patterns, numbers, and honest trade-offs from the field including a -20% reduction on a production EKS cluster. You'll leave with a checklist you can apply to your own cluster next week.... Read more
Modern systems are built in a culture of technological excess: infinite scaling, endless tooling, growing abstraction layers, and permanent optimization.
For decades, the industry relied on Moore’s Law and continuous hardware growth. But as software complexity keeps increasing, SREs increasingly face the hidden costs of this gluttony: cognitive overload, operational fragility, rising cloud costs, operational entropy, and systems that nobody fully understands anymore.
Drawing from real-world experience in SRE, GreenOps, and FinOps initiatives, this talk explores operational sobriety as an engineering discipline. Inspired in part by the “erooM” perspective, improving systems through software efficiency rather than endless hardware expansion, we will discuss how reducing unnecessary complexity can help build systems that remain understandable, resilient, and sustainable over time.
The future of reliability may depend less on endless accumulation than on rediscovering the culture of economy, efficiency, and resilience that originally shaped engineering itself.... Read more
AI agents are no longer just chatbots. They now execute code, access infrastructure, call APIs, and trigger real production actions. What works perfectly in a demo can quickly become a reliability and security nightmare at scale.
In this talk, Walid Mansia explores the hidden operational challenges behind autonomous agents, MCP servers, and LLM systems: infinite loops, hallucinated actions, exploding token costs, broken automations, and observability gaps traditional SRE tooling cannot explain.
Walid will also share practical patterns for safely operating AI systems in production, including sandboxing, execution limits, least-privilege access, and AI-native observability.... Read more
19:00
Wrap up
Scan each other's QR codes & head to a nearby pub!
Three Minutes Up, Thirty Seconds Down: Disposable GPU Clusters for AI Devs
Abstract
AI developers want a GPU the way they want a kind cluster: kubectl apply and it's there, kubectl delete and it's gone, no console in between. The cloud default punishes that: provisioning latency, sticky nodes, per-tenant security setup. We give developers a kubectl context with an auto-provisioning GPU under it: a T4 or L4 in ~3 minutes, gone in ~30 seconds, zero infra to touch.
Live demo: a dev creates a GPU-template tenant cluster, applies a pod, and a node auto-provisions in GCP. The pod runs, is deleted, the node disappears 30s later. Under the hood is a four-piece contract: tenant Kubernetes for isolated RBAC and node pools; Karpenter for auto-provisioning; an idle policy (consolidateAfter 30s) collapses nodes on idle, not on cluster delete; and a tunnel keeps the experience kubectl-native, not console-native.
We close with the failure modes (Karpenter template/provider drift, NodeProvider creds silently disabled) and the GitOps boundary that keeps multi-team setups intact.
Bio
An active contributor to OpenSource projects on GitHub, blogger and content creator, focusing on practical, scalable solutions in cloud-native environments. DevOps and Platform Engineering practitioner and advocate. Visit: cloudrumble.net
Amine Saboni
Pruna AI
Scale your GPU workloads with Karpenter and Argo Workflows
Abstract
Running ML workloads on Kubernetes can be a burden to manage, especially when the usages tend to diversify. Batch inference will involve different constraints as experimentation, evaluation or serving ML models for user-facing applications.
How to assign the correct GPUs for each usage ?
How to scale differently workloads with flexibility ?
How to efficiently scale down to 0, when no GPUs are needed ?
To answer those questions, we will see in this talk how combine the standard orchestration approach of Argo Workflows with the autoscaling super powers of Karpenter, a tool developed by AWS and given back to the community. We'll dig into the difficulties of scaling GPUs and how Karpenter solves them, with a return of experience of Pruna's computing platform.
Bio
MLOps Engineer working at Pruna AI, I am interested in efficient, reliable & robustly beneficial AI. My work is centered around building ML development platforms, and provide relevant tooling to build & deploy robust ML models in applications.
Paul-Henri Pillet
Gatling
Don't RTFM, use AI instead for your first Gatling tests!
Abstract
AI has dramatically reduced the complexity of starting a new code project. Today, you don't read the docs first - you ask your coding agent.
In this talk, we'll show how this applies to performance testing. How developers with zero Gatling experience can write realistic load tests in minutes, using AI as their entry point instead of documentation, and why this changes how teams adopt performance testing altogether.
Bio
Paul-Henri Pillet is the CEO of Gatling, one of France’s most successful open source technology companies and a leading provider of high performance load testing solutions. Since joining Gatling in 2014, he has helped grow the platform into a globally trusted tool used by hundreds of thousands of organizations to ensure the scalability and reliability of their applications. Based in Paris, Paul-Henri combines a background in business, public policy, and entrepreneurship, with degrees from HEC Paris and Freie Universität Berlin.
Marcos Novelli Harispe
LoopStudio
The Reliability Loop: Bridging Prevention and Incident Response with Ephemeral Environments
Abstract
Unexpected production errors often come from a familiar place: shared environments that fail to accurately validate real-world conditions.
Ephemeral Environments (EEs) offer a different approach. Instead of relying on fragile, long-lived shared environments, teams can validate each change in an isolated, production-like environment that includes code, infrastructure, and configuration.
In this talk, we’ll explore how EEs reduce unexpected issues in production by improving quality validation, eliminating cross-team interference, enabling safer automated testing, and accelerating deployments.
But prevention is only part of the story.
We’ll also touch on how the same environments can be used after an incident to reproduce failures, reconstruct system state, and support deeper root cause analysis and post-mortems.
This session focuses on how SRE teams can use an incremental and adaptable ephemeral strategy to connect prevention and incident response into a continuous reliability loop.
Bio
Marcos Novelli Harispe is a software engineer specializing in AWS, serverless architectures, and data-driven solutions. He has delivered scalable notification systems, analytics platforms, and automation tools across freelance and full-time roles, with a focus on reliability, efficiency, and measurable business impact.
Daniel Afonso
PagerDuty
Incident Response Reimagined: Accelerating Resolution with AI Agents
Abstract
It’s 3:45 AM, and you get paged. Still half asleep, you reach for your laptop. Turns out something is wrong, so you acknowledge the incident and embark on a journey to figure out what is happening, mobilize the right people, and eventually fix the issue. But what if we could make it easier? In this talk, I’ll show how AI agents can help you fast-track triage, coordinate responses, debug problems, and run fixes. To wrap up, I will show how we can go even further and move from merely reacting to incidents to proactively preventing them. In the spirit of shifting left, let’s understand how we can track future incidents our code may cause, right before we commit it!
Bio
Daniel Afonso is a Senior Developer Advocate at PagerDuty, SolidJS DX team member, Instructor at Egghead.io, and Author of State Management with React Query. Daniel has a full-stack background, having worked with different languages and frameworks on various projects from IoT to Fraud Detection. He is passionate about learning and teaching and has spoken at multiple conferences around the world about topics he loves. In his free time, when he's not learning new technologies or writing about them, he's probably reading comics or watching superhero movies and shows.
Eugene Ngontang
Spitzkop
Your compliance tool describes the platform. It doesn't run it
Abstract
Every regulated org has the same scar: a beautiful governance tool - a CMDB, a compliance tracker, a project portal - that describes what production should look like, sitting next to a real platform that actually runs something else. Two sources of truth. Manual sync. Tickets. Drift.
The compliance team trusts the tool; the SREs trust the cluster; nobody trusts the gap between them. This talk is a field report on closing that gap. Drawing on a teardown of a large enterprise platform and the sovereign platform we're building at SPITZKOP, I'll show how to make the control plane connected to the execution plane: GitOps as the single source of truth, policy-ascode (Kyverno) as the admission backstop, schema-first config generation (CUE) shifting compliance left, and a conformity engine that turns regulatory requirements (NIS2, DORA, EUCS) into executable, attested deployment steps — not Excel rows.
You'll leave with a concrete pattern for making compliance a property of your pipeline, not a parallel universe.
Bio
Eugène Ngontang is Founder & CTO of SPITZKOP, where he's building ESSINGAN — a sovereign platform-engineering product for regulated and critical workloads. He's spent years operating production-critical infrastructure for large enterprises (PwC, NCR Atleos and others), with a focus on GitOps, platform engineering, confidential computing, and measurable sovereignty. He previously presented sovereign network automation (GitOps + YANG) at FRnOG-42. He works at the intersection of reliability, compliance, and cloud-native engineering.
Qinghui Xu
Criteo
Building Streaming Infrastructure at Scale - Criteo's Journey with Apache Flink
Abstract
As organizations process ever larger volumes of real time data, building a reliable streaming platform becomes a significant engineering challenge. In this session, Qinghui Xu shares Criteo’s experience operating Apache Flink at scale, covering the architectural decisions, operational lessons, and platform practices that enable large scale stream processing in production. Attendees will gain insight into the challenges of performance, reliability, and observability in distributed streaming systems, along with practical takeaways for teams running critical data workloads on Apache Flink.
Bio
Qinghui Xu is a Senior DevOps Engineer at Criteo, where he has spent nearly a decade building and operating large scale infrastructure. Prior to Criteo, he worked as a Software Engineer at Spark Archives. Qinghui holds engineering degrees from École Polytechnique and Télécom Paris, with a background spanning distributed systems, machine learning, and data-driven technologies.
Francois Le Pape
Contentful
Ship Fast, Break Nothing: Auto-Rollback with Cloudflare workers
Abstract
In this session, I’ll walk through a production rollout system that combines versioned deployments, gradual traffic ramps, error-budget-aware gates, and automatic rollback on the Cloudflare Developer Platform.
We’ll cover the hard parts: defining trustworthy signals, preventing flappy rollbacks, and handling emergency bypass paths safely.
You’ll leave with concrete patterns to increase deploy frequency while reducing risk and MTTR.
Bio
Senior DevOps/Software Engineer at Contentful. Working with Kubernetes and Cloudflare workers. Focus on Developer Experience and SRE.
Solvik Blum
Enix
kube-image-keeper: keeping workloads alive when registries fail
Abstract
kube-image-keeper kuik, https://github.com/enix/kube-image-keeper makes sure that your workloads keep starting even when the upstream registry is down, Docker Hub rate-limits you, or an image tag gets deleted. It’s an open-source Kubernetes image resiliency solution that provides transparent image routing across multiple registries, as well as selective caching and replication of the images used by your workloads into local or external registries. On 100+ production clusters it has saved us countless incidents during registry outages and rate-limit storms, and made node autoscaling actually reliable. The talk covers the concrete usecases it solves, what we got wrong in v1, why we rewrote it as v2, and the edge cases that bit us along the way: mutating webhooks, garbage collection, and keeping the cache honest at scale.
Bio
With over 15 years of experience in infrastructure and production engineering, Solvik has always worked in highly scalable and business-critical environments.
He has worked at Scaleway and Dailymotion, where he designed, operated and secured large-scale cloud-native platforms, with a strong focus on Kubernetes, bare-metal infrastructure and production reliability.
He has now joined Enix, where he leads Production and Security topics with a pragmatic approach focused on long-term operations, open source, bare metal and cost control.
Luckas Bosch
Red Hat
Are IT Certifications Actually Worth It?
Abstract
IT certifications are everywhere in today’s industry, especially around cloud, Kubernetes, Linux, security, and DevOps. Some engineers see them as valuable learning tools and career accelerators, while others consider them disconnected from real production work.
In this talk, I will share a practical and balanced perspective based on both certification experience and real-world production environments. We will discuss what certifications are actually good for, where they fall short, and how engineers can use them effectively without relying on them blindly.
Bio
Luckas Bosch is an OpenShift Consultant at Red Hat and a former Site Reliability Engineer with experience managing Kubernetes and cloud-native platforms in production environments, particularly for streaming and broadcasting workloads. He holds multiple certifications including CKA, CKS, RHCSA, RHCE, and Kubestronaut, and also hosts the French tech podcast La Tangente.
Rodrigue NDE
kladriva
FinOps Patterns for Kubernetes in Production - From Cost Blindness to -20%
Abstract
FinOps Patterns for Kubernetes in Production Running Kubernetes on AWS doesn't automatically mean running it efficiently. Most teams focus on stability and delivery and discover their cloud bill three months later, when the damage is done. In this talk, I share concrete lessons from optimizing EKS infrastructure costs in production: what to measure, where the waste actually hides, and which levers move the needle, without breaking anything. We'll cover: 1. Why standard cloud cost tools (Cost Explorer, Kubecost) only give you half the picture. The underestimated impact of Karpenter, instance diversification, and Spot strategy on real workloads. 2. How right-sizing Kubernetes requests saved more than any Reserved Instance commitment. 3. The FinOps discipline that stuck vs. the optimizations that quietly regressed No vendor pitch. No theoretical framework. Just patterns, numbers, and honest trade-offs from the field including a -20% reduction on a production EKS cluster. You'll leave with a checklist you can apply to your own cluster next week.
Bio
Rodrigue NDE is a Cloud Platform Engineer with 7 years of experience, specializing in Kubernetes, AWS infrastructure, and FinOps. He has led the design and industrialization of production-grade EKS platforms for multiple organizations, with a focus on reliability, security, and cost efficiency. His background as a Java Tech Lead before moving into Cloud Engineering gives him an unusual perspective: he builds platforms with the developer experience in mind, not just the infrastructure metrics. Recent highlights include a 20% reduction in AWS costs on an EKS production cluster, a platform SLA improvement to 99.9%. Rodrigue is currently working toward his next role as a Senior Platform Engineer, with a particular interest in FinOps, Kubernetes security, and platform self-service.
Brice Le Roux
OCTO Technology
The Art of Being Sober in an Age of Gluttony
Abstract
Modern systems are built in a culture of technological excess: infinite scaling, endless tooling, growing abstraction layers, and permanent optimization.
For decades, the industry relied on Moore’s Law and continuous hardware growth. But as software complexity keeps increasing, SREs increasingly face the hidden costs of this gluttony: cognitive overload, operational fragility, rising cloud costs, operational entropy, and systems that nobody fully understands anymore.
Drawing from real-world experience in SRE, GreenOps, and FinOps initiatives, this talk explores operational sobriety as an engineering discipline. Inspired in part by the “erooM” perspective, improving systems through software efficiency rather than endless hardware expansion, we will discuss how reducing unnecessary complexity can help build systems that remain understandable, resilient, and sustainable over time.
The future of reliability may depend less on endless accumulation than on rediscovering the culture of economy, efficiency, and resilience that originally shaped engineering itself.
Bio
Brice Le Roux is a cloud architect, GreenOps expert, and SRE/DevOps consultant focused on sustainable and resilient information systems.
He works on reliability engineering, sustainable infrastructure, cloud governance, and operational sobriety, and trains architects and engineers on Kubernetes and eco-designed architectures.
Walid Mansia
CANAL+ Group
The SRE Nightmare Nobody Talks About
Abstract
AI agents are no longer just chatbots. They now execute code, access infrastructure, call APIs, and trigger real production actions. What works perfectly in a demo can quickly become a reliability and security nightmare at scale.
In this talk, Walid Mansia explores the hidden operational challenges behind autonomous agents, MCP servers, and LLM systems: infinite loops, hallucinated actions, exploding token costs, broken automations, and observability gaps traditional SRE tooling cannot explain.
Walid will also share practical patterns for safely operating AI systems in production, including sandboxing, execution limits, least-privilege access, and AI-native observability.
Bio
Walid Mansia is an AI Solutions Architect and Platform Engineering specialist focused on autonomous agents, MCP servers, and AI systems in production. He currently leads next-generation AI and data platform initiatives at CANAL+ Group, with previous experience across AWS, DevOps, cloud infrastructure, and large-scale platform engineering. Walid has spent over a decade designing reliable, scalable systems at the intersection of AI, cloud, and SRE.
Poone Mokari
ewake
KeynoteProduction is a wilderness. Treat it like one.
Abstract
Production is a wilderness. The terrain shifts every minute, meaning human and agent failures are not caused by missing intelligence. They come from acting on stale context.
Discovery takes time, and in production it is often the most expensive part of the workflow.
In this talk, I show why AI SRE agents fail when they rely on runtime discovery, and what changes when they operate on live context instead of snapshots.
The short version: an AI SRE that admits uncertainty is the only kind you can safely let run.
Bio
Co-founder and CEO of ewake, Poone Mokari builds AI SRE agents grounded in a live system map. Prior to founding ewake, she worked as a Site Reliability Engineer at companies including Criteo, where she gained hands-on experience managing large-scale production systems and incident response.
Peter Marshall
Imply
KeynoteDecoupling Observability for Incident Response at Scale
Abstract
On-call incidents don’t fail because teams lack dashboards. They fail because observability systems slow down under real investigative load
As telemetry volumes grow and retention windows expand, SRE teams are being asked to run deeper, broader investigations—often under time pressure—on platforms that were designed for steady-state monitoring, not bursty incident response. Tightly coupled observability stacks bind storage, compute, and query together, forcing teams to overprovision infrastructure, limit retention, or accept degraded performance during incidents.
In this talk, we’ll explore why decoupling observability architectures is becoming essential for SRE teams operating at scale. Using a real incident investigation workflow, we’ll break down how separating data storage, compute, and interaction layers allows teams to keep fast, reliable monitoring while elastically scaling investigations when incidents occur.
Bio
Peter Marshall is an award-winning speaker, technology leader, and community builder with 25 years' experience in data architecture and digital transformation. As Director of Developer Relations at Imply, he leads programs that grow and engage global communities through education, support, and events. With experience across startups, enterprises, and the public sector, Peter brings technical expertise and strategic vision to help organizations leverage real-time data and observability technologies. He holds a BA in Theology and Computer Studies from the University of Birmingham
Maxime Brugidou
Criteo
KeynoteReliability in the Age of Abundance
Abstract
SRE used to be about keeping systems alive under scarcity: limited machines, limited telemetry, limited deployment safety. Now the risk is abundance: too much cloud, too much telemetry, too many abstractions, too much automation and too many AI systems acting faster than humans can understand.
Bio
Maxime Brugidou is VP of Engineering, Platform at Criteo, where he leads large-scale platform and infrastructure engineering teams focused on reliability, scalability, and developer productivity. With a background spanning SRE, distributed systems, and data platforms, he has spent more than 15 years helping build and scale high-performance engineering organizations. Maxime is also active in the AI and cloud-native community, regularly supporting technical meetups and knowledge sharing initiatives in Paris.