SREday

Site Reliability, DevOps and Cloud

June 16, 2026 Criteo, Paris, France

1
Day
10+
Speakers
1
Track
100+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

CANAL+ Group, Contentful, Criteo, Enix, ewake, Gatling, Imply, kladriva, LoopStudio, OCTO Technology, PagerDuty, Pruna AI, Red Hat, Spitzkop, vCluster

Topics so far:

This is a past event, what's next?

Schedule

June 16, 2026 • single track • 9AM - 7PM • Paris, in-person
view as table
main room • Track 1

09:00

Poone Mokari

KeynoteProduction is a wilderness. Treat it like one.

ewake
Production is a wilderness. The terrain shifts every minute, meaning human and agent failures are not caused by missing intelligence. They come from acting on stale context. Discovery takes time, and in production it is often the most expensive part of the workflow. In this talk, I show why AI SRE agents fail when they rely on runtime discovery, and what changes when they operate on live context instead of snapshots. The short version: an AI SRE that admits uncertainty is the only kind you can safely let run.... Read more

09:30

Peter Marshall

KeynoteDecoupling Observability for Incident Response at Scale

Imply
On-call incidents don’t fail because teams lack dashboards. They fail because observability systems slow down under real investigative load As telemetry volumes grow and retention windows expand, SRE teams are being asked to run deeper, broader investigations—often under time pressure—on platforms that were designed for steady-state monitoring, not bursty incident response. Tightly coupled observability stacks bind storage, compute, and query together, forcing teams to overprovision infrastructure, limit retention, or accept degraded performance during incidents. In this talk, we’ll explore why decoupling observability architectures is becoming essential for SRE teams operating at scale. Using a real incident investigation workflow, we’ll break down how separating data storage, compute, and interaction layers allows teams to keep fast, reliable monitoring while elastically scaling investigations when incidents occur.... Read more

10:00

Maxime Brugidou

KeynoteReliability in the Age of Abundance

CriteoWatch
SRE used to be about keeping systems alive under scarcity: limited machines, limited telemetry, limited deployment safety. Now the risk is abundance: too much cloud, too much telemetry, too many abstractions, too much automation and too many AI systems acting faster than humans can understand.... Read more

10:30

Coffee break

Main lobby

11:00

Piotr Zaniewski

Three Minutes Up, Thirty Seconds Down: Disposable GPU Clusters for AI Devs

vClusterWatch
AI developers want a GPU the way they want a kind cluster: kubectl apply and it's there, kubectl delete and it's gone, no console in between. The cloud default punishes that: provisioning latency, sticky nodes, per-tenant security setup. We give developers a kubectl context with an auto-provisioning GPU under it: a T4 or L4 in ~3 minutes, gone in ~30 seconds, zero infra to touch. Live demo: a dev creates a GPU-template tenant cluster, applies a pod, and a node auto-provisions in GCP. The pod runs, is deleted, the node disappears 30s later. Under the hood is a four-piece contract: tenant Kubernetes for isolated RBAC and node pools; Karpenter for auto-provisioning; an idle policy (consolidateAfter 30s) collapses nodes on idle, not on cluster delete; and a tunnel keeps the experience kubectl-native, not console-native. We close with the failure modes (Karpenter template/provider drift, NodeProvider creds silently disabled) and the GitOps boundary that keeps multi-team setups intact.... Read more

11:30

Amine Saboni

Scale your GPU workloads with Karpenter and Argo Workflows

Pruna AIWatch
Running ML workloads on Kubernetes can be a burden to manage, especially when the usages tend to diversify. Batch inference will involve different constraints as experimentation, evaluation or serving ML models for user-facing applications. * How to assign the correct GPUs for each usage ? * How to scale differently workloads with flexibility ? * How to efficiently scale down to 0, when no GPUs are needed ? To answer those questions, we will see in this talk how combine the standard orchestration approach of Argo Workflows with the autoscaling super powers of Karpenter, a tool developed by AWS and given back to the community. We'll dig into the difficulties of scaling GPUs and how Karpenter solves them, with a return of experience of Pruna's computing platform. ... Read more

12:00

Paul-Henri Pillet

Don't RTFM, use AI instead for your first Gatling tests!

GatlingWatch
AI has dramatically reduced the complexity of starting a new code project. Today, you don't read the docs first - you ask your coding agent. In this talk, we'll show how this applies to performance testing. How developers with zero Gatling experience can write realistic load tests in minutes, using AI as their entry point instead of documentation, and why this changes how teams adopt performance testing altogether.... Read more

12:30

Marcos Novelli Harispe

The Reliability Loop: Bridging Prevention and Incident Response with Ephemeral Environments

LoopStudioWatch
Unexpected production errors often come from a familiar place: shared environments that fail to accurately validate real-world conditions. Ephemeral Environments (EEs) offer a different approach. Instead of relying on fragile, long-lived shared environments, teams can validate each change in an isolated, production-like environment that includes code, infrastructure, and configuration. In this talk, we’ll explore how EEs reduce unexpected issues in production by improving quality validation, eliminating cross-team interference, enabling safer automated testing, and accelerating deployments. But prevention is only part of the story. We’ll also touch on how the same environments can be used after an incident to reproduce failures, reconstruct system state, and support deeper root cause analysis and post-mortems. This session focuses on how SRE teams can use an incremental and adaptable ephemeral strategy to connect prevention and incident response into a continuous reliability loop. ... Read more

13:00

Lunch & networking

Main lobby

14:00

Daniel Afonso

Incident Response Reimagined: Accelerating Resolution with AI Agents

PagerDutyWatch
It’s 3:45 AM, and you get paged. Still half asleep, you reach for your laptop. Turns out something is wrong, so you acknowledge the incident and embark on a journey to figure out what is happening, mobilize the right people, and eventually fix the issue. But what if we could make it easier? In this talk, I’ll show how AI agents can help you fast-track triage, coordinate responses, debug problems, and run fixes. To wrap up, I will show how we can go even further and move from merely reacting to incidents to proactively preventing them. In the spirit of shifting left, let’s understand how we can track future incidents our code may cause, right before we commit it!... Read more

14:30

Eugene Ngontang

Your compliance tool describes the platform. It doesn't run it

SpitzkopWatch
Every regulated org has the same scar: a beautiful governance tool - a CMDB, a compliance tracker, a project portal - that describes what production should look like, sitting next to a real platform that actually runs something else. Two sources of truth. Manual sync. Tickets. Drift. The compliance team trusts the tool; the SREs trust the cluster; nobody trusts the gap between them. This talk is a field report on closing that gap. Drawing on a teardown of a large enterprise platform and the sovereign platform we're building at SPITZKOP, I'll show how to make the control plane connected to the execution plane: GitOps as the single source of truth, policy-ascode (Kyverno) as the admission backstop, schema-first config generation (CUE) shifting compliance left, and a conformity engine that turns regulatory requirements (NIS2, DORA, EUCS) into executable, attested deployment steps — not Excel rows. You'll leave with a concrete pattern for making compliance a property of your pipeline, not a parallel universe.... Read more

15:00

Qinghui Xu

Building Streaming Infrastructure at Scale - Criteo's Journey with Apache Flink

CriteoWatch
As organizations process ever larger volumes of real time data, building a reliable streaming platform becomes a significant engineering challenge. In this session, Qinghui Xu shares Criteo’s experience operating Apache Flink at scale, covering the architectural decisions, operational lessons, and platform practices that enable large scale stream processing in production. Attendees will gain insight into the challenges of performance, reliability, and observability in distributed streaming systems, along with practical takeaways for teams running critical data workloads on Apache Flink. ... Read more

15:30

Francois Le Pape

Ship Fast, Break Nothing: Auto-Rollback with Cloudflare workers

ContentfulWatch
In this session, I’ll walk through a production rollout system that combines versioned deployments, gradual traffic ramps, error-budget-aware gates, and automatic rollback on the Cloudflare Developer Platform. We’ll cover the hard parts: defining trustworthy signals, preventing flappy rollbacks, and handling emergency bypass paths safely. You’ll leave with concrete patterns to increase deploy frequency while reducing risk and MTTR.... Read more

16:00

Networking & sponsor crawl

Main lobby

16:30

Solvik Blum

kube-image-keeper: keeping workloads alive when registries fail

EnixWatch
kube-image-keeper [kuik, https://github.com/enix/kube-image-keeper](https://github.com/enix/kube-image-keeper) makes sure that your workloads keep starting even when the upstream registry is down, Docker Hub rate-limits you, or an image tag gets deleted. It’s an open-source Kubernetes image resiliency solution that provides transparent image routing across multiple registries, as well as selective caching and replication of the images used by your workloads into local or external registries. On 100+ production clusters it has saved us countless incidents during registry outages and rate-limit storms, and made node autoscaling actually reliable. The talk covers the concrete usecases it solves, what we got wrong in v1, why we rewrote it as v2, and the edge cases that bit us along the way: mutating webhooks, garbage collection, and keeping the cache honest at scale.... Read more

17:00

Luckas Bosch

Are IT Certifications Actually Worth It?

Red HatWatch
IT certifications are everywhere in today’s industry, especially around cloud, Kubernetes, Linux, security, and DevOps. Some engineers see them as valuable learning tools and career accelerators, while others consider them disconnected from real production work. In this talk, I will share a practical and balanced perspective based on both certification experience and real-world production environments. We will discuss what certifications are actually good for, where they fall short, and how engineers can use them effectively without relying on them blindly.... Read more

17:30

Rodrigue NDE

FinOps Patterns for Kubernetes in Production - From Cost Blindness to -20%

kladrivaWatch
FinOps Patterns for Kubernetes in Production Running Kubernetes on AWS doesn't automatically mean running it efficiently. Most teams focus on stability and delivery and discover their cloud bill three months later, when the damage is done. In this talk, I share concrete lessons from optimizing EKS infrastructure costs in production: what to measure, where the waste actually hides, and which levers move the needle, without breaking anything. We'll cover: 1. Why standard cloud cost tools (Cost Explorer, Kubecost) only give you half the picture. The underestimated impact of Karpenter, instance diversification, and Spot strategy on real workloads. 2. How right-sizing Kubernetes requests saved more than any Reserved Instance commitment. 3. The FinOps discipline that stuck vs. the optimizations that quietly regressed No vendor pitch. No theoretical framework. Just patterns, numbers, and honest trade-offs from the field including a -20% reduction on a production EKS cluster. You'll leave with a checklist you can apply to your own cluster next week.... Read more

18:00

Brice Le Roux

The Art of Being Sober in an Age of Gluttony

OCTO TechnologyWatch
Modern systems are built in a culture of technological excess: infinite scaling, endless tooling, growing abstraction layers, and permanent optimization. For decades, the industry relied on Moore’s Law and continuous hardware growth. But as software complexity keeps increasing, SREs increasingly face the hidden costs of this gluttony: cognitive overload, operational fragility, rising cloud costs, operational entropy, and systems that nobody fully understands anymore. Drawing from real-world experience in SRE, GreenOps, and FinOps initiatives, this talk explores operational sobriety as an engineering discipline. Inspired in part by the “erooM” perspective, improving systems through software efficiency rather than endless hardware expansion, we will discuss how reducing unnecessary complexity can help build systems that remain understandable, resilient, and sustainable over time. The future of reliability may depend less on endless accumulation than on rediscovering the culture of economy, efficiency, and resilience that originally shaped engineering itself.... Read more

18:30

Walid Mansia

The SRE Nightmare Nobody Talks About

CANAL+ GroupWatch
AI agents are no longer just chatbots. They now execute code, access infrastructure, call APIs, and trigger real production actions. What works perfectly in a demo can quickly become a reliability and security nightmare at scale. In this talk, Walid Mansia explores the hidden operational challenges behind autonomous agents, MCP servers, and LLM systems: infinite loops, hallucinated actions, exploding token costs, broken automations, and observability gaps traditional SRE tooling cannot explain. Walid will also share practical patterns for safely operating AI systems in production, including sandboxing, execution limits, least-privilege access, and AI-native observability.... Read more

19:00

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room
09:00 Keynote: Production is a wilderness. Treat it like one.
Poone Mokari • ewake
09:30 Keynote: Decoupling Observability for Incident Response at Scale
Peter Marshall • Imply
10:00 Keynote: Reliability in the Age of Abundance
Maxime Brugidou • Criteo
10:30 Coffee break
11:00 Three Minutes Up, Thirty Seconds Down: Disposable GPU Clusters for AI Devs
Piotr Zaniewski • vCluster
11:30 Scale your GPU workloads with Karpenter and Argo Workflows
Amine Saboni • Pruna AI
12:00 Don't RTFM, use AI instead for your first Gatling tests!
Paul-Henri Pillet • Gatling
12:30 The Reliability Loop: Bridging Prevention and Incident Response with Ephemeral Environments
Marcos Novelli Harispe • LoopStudio
13:00 Lunch & networking
14:00 Incident Response Reimagined: Accelerating Resolution with AI Agents
Daniel Afonso • PagerDuty
14:30 Your compliance tool describes the platform. It doesn't run it
Eugene Ngontang • Spitzkop
15:00 Building Streaming Infrastructure at Scale - Criteo's Journey with Apache Flink
Qinghui Xu • Criteo
15:30 Ship Fast, Break Nothing: Auto-Rollback with Cloudflare workers
Francois Le Pape • Contentful
16:00 Networking & sponsor crawl
16:30 kube-image-keeper: keeping workloads alive when registries fail
Solvik Blum • Enix
17:00 Are IT Certifications Actually Worth It?
Luckas Bosch • Red Hat
17:30 FinOps Patterns for Kubernetes in Production - From Cost Blindness to -20%
Rodrigue NDE • kladriva
18:00 The Art of Being Sober in an Age of Gluttony
Brice Le Roux • OCTO Technology
18:30 The SRE Nightmare Nobody Talks About
Walid Mansia • CANAL+ Group
19:00 Wrap up

Speakers

Amine Saboni
Pruna AI
Brice Le Roux
OCTO Technology
Daniel Afonso
PagerDuty
Eugene Ngontang
Spitzkop
Francois Le Pape
Contentful
Luckas Bosch
Red Hat
Marcos Novelli Harispe
LoopStudio
Maxime Brugidou
Criteo
Paul-Henri Pillet
Gatling
Peter Marshall
Imply
Piotr Zaniewski
vCluster
Poone Mokari
ewake
Qinghui Xu
Criteo
Rodrigue NDE
kladriva
Solvik Blum
Enix
Walid Mansia
CANAL+ Group

Venue

Criteo

32 Rue Blanche
75009 Paris, France

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one