SREday

Site Reliability, DevOps and Cloud

March 12, 2026 London, UK

1
Day
30+
Speakers
3
Tracks
150+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

AuthZed, AWS, ClickHouse, Coralogix, Dynatrace, Ewake.ai, Gatling, Harness, ilert, Ilustre Municipalidad de Santo Domingo, Imply, ING, ING Germany, ING Netherland, Longhorn Games, MetalBear, Metoro, NOFire AI, Okta, Phoebe, RunWhen, Seamflow, Supercharged, Typeform, WSO2

Topics so far:

This is a past event, what's next?

Schedule

March 12, 2026 • 3 parallel tracks • 9AM - 8PM • London, in-person
view as table
screen 1 • Track 1

09:00

Peter Marshall

KeynoteDecoupled Observability - an architecture for scalable detection and investigation

ImplyWatch
Teams are being asked to retain more data, investigate problems and incidents further back in time, and respond faster — all while controlling costs. The challenge isn't a lack of tools; it's the tightly coupled architectures. In this session, we'll explore why decoupling the data layer from the interaction layer is critical. Using a real-world use case, we'll see how traditional approaches force teams to trade off visibility, performance, or usability at the worst possible moment. We'll break down how a decoupled architecture allows teams to run continuous detections on recent data while elastically scaling across large historical datasets, without duplicating data or abandoning existing workflows. The result is greater flexibility, better cost control, and faster incident response — even as data volumes and retention requirements continue to grow.... Read more

09:30

Dewan Ahmed

KeynoteSecure by Default: Building Confidence in AI-Driven Delivery

HarnessWatch
The fastest way to break trust in DevSecOps is to automate insecurity at scale. As AI takes a central role in our pipelines, it is time to rethink what "secure by default" really means. In this keynote, Dewan Ahmed will challenge the audience to look beyond vulnerability scanners and compliance gates. He will share a vision for intelligent security by design, where native intelligence within the delivery platform detects not only vulnerable code but risky delivery behavior such as misconfigured environments, suspicious artifact provenance, and drift between source and runtime. You will walk away with a framework for balancing automation with human oversight and examples from Harness’ work on building verifiable, auditable, AI-native delivery systems. In the new world of DevSecOps, safety is not a step; it is an outcome we continuously learn to improve.... Read more

10:00

Matt Henderson

KeynoteThe Immune System for Software: Lessons from Biology

PhoebeWatch
Most reliability work today is still centred on reactive troubleshooting: diagnosing a multitude of alerts, pulling large groups into incidents, and engineers scrambling to understand what's happening. To truly change that pattern, we need systems that can predict and prevent failures before they occur. Biological immune systems offer a powerful blueprint for how software can defend and ultimately heal itself. This talk introduces a framework for thinking about "software immunity", highlights the gaps in today's observability and includes some "under the hood" details on designing AI agents for reliability. Innate immunity comes from built-in defences like testing, feature flags, auto-scaling, and circuit breakers — mechanisms that provide immediate, general protection. Adaptive immunity, meanwhile, emerges from AI agents that learn from new data, refine their understanding of system behaviour, and apply those lessons to predict and pre-emptively fix failures. We'll break down the key ingredients for trustworthy AI agents in reliability and beyond: transparent reasoning rather than opaque black-box outputs; strong control mechanisms and guardrails; a governed data layer for effective data access; continuous learning from each execution cycle; and graduated autonomy — from suggestions, to human-in-the-loop actions, to fully automated remediation.... Read more

10:30

Tyler Hannan

KeynoteDo Metrics Matter?

ClickHouseWatch
When I get paged, I open the metrics dashboard. That hasn’t changed. Metrics are still the fastest way to get a rough sense of whether a system is unhealthy, especially when you’re dealing with known failure modes and issues you can reasonably anticipate ahead of time. But with the increase of automated tooling and exploratory analysis, we are faced with a question...do metrics matter?... Read more

11:00

Birol Yildiz

KeynoteWhen Incidents Fix Themselves: AI SRE in action

ilertWatch
The next evolution of incident response isn’t faster alerts, it’s autonomous resolution. Join ilert CEO Birol Yildiz as he shows how AI SRE agents now diagnose and remediate outages without waking anyone up. Learn how these systems combine observability data, deployment context, and code intelligence to restore services in minutes and hand over clean incident reports instead of 3 a.m. pages.... Read more

11:30

Paging You for Pizza - Powered by incident.io !

Main lobby

12:30

Deniz Yalcin & William Ravensbergen

Reliability starts at the source: Why Customer Data is your most underrated SRE dependency

ING Netherland & ING GermanyWatch
Reliability is often framed as an infrastructure problem, but at Banking, it begins with data. Customer data is a core SRE dependency that drives data quality, fraud prevention, knowing your customer, and user trust. When upstream data is delayed, malformed, duplicated, or semantically inconsistent, no amount of autoscaling or incident response can preserve a seamless experience. This talk explores why customer data quality, lineage and observability must be treated as first class reliability concerns. By shifting reliability thinking left toward data consistency, validation, and monitoring at ingestion, fintech teams can reduce incident blast radius, improve MTTR, and protect customer confidence where it matters most: at the source.... Read more

13:00

Adriana Villela

Observability is a Team Sport!

DynatraceWatch
Observability may have been a newcomer few years ago, but it's safe to say that it's a pretty well-established part of our tech lives today. But as organizations embrace Observability, they run the risk of creating Yet Another Silo, in much the same way that DevOps created more silos, instead of breaking them. Reliability can’t happen without Observability, and Observability itself must be looked at holistically. It isn't the responsibility of one single team, but instead weaves its way into multiple teams. In this talk, Adriana dives into the roles and responsibilities that Development, Quality Assurance, CI/CD, and SRE teams when it comes to contributing to the Observability story. Spoiler alert - it's not as straightforward as you might think! She will also show how an “Observability team”, if not designed and rolled out properly, can take away an organization’s collective responsibility for Observability, and dilute the promise of Observability.... Read more

13:30

Heather Thacker

Choose Your Weapon: The Performance Testing Arsenal

GatlingWatch
Your app works great on your laptop, in the dev environment. Then production hits 10x expected traffic during a marketing campaign and everything falls apart. Or maybe not, instead six months of data accumulates, causing the response times to be painfully slow. Load testing, stress testing, soak testing, and spike testing, they all sound similar, but address completely different problems, and most teams only do one, if at all. This talk breaks down the essential types of performance testing every app and web developer should understand. Learn when to use each approach, what problems they uncover, and how to integrate them into your development workflow without drowning in complexity. We'll cover real-world scenarios where each testing type saves production systems, helping you choose the right weapon for your performance battles. What you'll learn: - The differences between load, stress, soak, and spike testing, and when each matters - Which performance testing types reveal which production problems, before they happen - How to integrate performance testing into CI/CD without slowing down development and deployments - Practical criteria for deciding which tests your application actually needs... Read more

14:00

Tasmia Niazi

From Learner to Leader: My Journey into SRE and Why Reliability Is a Culture, Not a Role

Explore how SRE is more than tools and dashboards: it is a mindset, a culture, and a way of thinking that transforms teams. Attendees will learn how adopting an SRE state of mind empowers individuals, improves engineering outcomes, and builds high-trust, resilient organisations.... Read more

14:30

Networking and sponsor crawl

Main lobby

15:00

Kyle Forster, Lydia Thomas, Anusha Gundala, Sravan E & Sonu Kumar Singh

Demo: Let any engineer on any team handle any issue anywhere in your stack

RunWhen & ex-Google SRE
Production incidents and operational toil have been steadily rising for years, but the worst is ahead. The drivers are clear - i) growing complexity of modern tech stacks, ii) shrinking ratios of SRE/DevOps/QA engineers per developer, and iii) the freight train of ultra-fast AI coding. Before weekly releases with tens of thousands of lines of code become the new normal, SRE leaders need to think differently about how issues and incidents are handled through their software development lifecyle. In today's session, Kyle will show the PR that broke his team's traditional SDLC, and talk about ways teams are transitioning from legacy models of SRE to modern models. Hint: get any engineer on any team ready to handle any issue anywhere in the tech stack in any environment. In parallel, Lydia will drop in to a production environment she has never seen before and prioritize issues, drive to root cause, file detailed tickets and resolve incidents in <30 minutes using tools built with RunWhen's AI SRE platform.... Read more

15:30

Rob Charlwood

Stranger Teams: Keeping Demogorgons out of production

SuperchargedWatch
Build SRE teams like 80s kids on bikes! With high trust, shared missions, honesty, and the ability to tackle monsters together. Learn why your SRE might be broken and how to fix it with psychological safety, teamwork, and better practices.... Read more

16:00

Sohan Maheshwar

Surviving the Dual-Write Problem in Distributed Systems

AuthZedWatch
Dual writes are one of the most common sources of data inconsistency in distributed systems. This talk demystifies why dual writes cannot be made safe inside a single process, then walks through modern architectural patterns that teams use in production to eliminate or mitigate the problem.... Read more

16:30

Jake Page

Honey, the Audience Broke My App: Reproduce & Fix Live in Kubernetes with mirrord

MetalBearWatch
A room of engineers breaks my live app and we fix it together. They scan a QR code which triggers a real failure in a Kubernetes staging environment. Together we reproduce, trace, patch, and validate the fix with kubectl debug and mirrord. A real breakage repaired live.... Read more

17:00

Happy Hour by Imply - grab a beer!

Main lobby

17:30

_ClickHouse Meetup - After SREday

18:00

Adriana Villela - ClickHouse Meetup

18:30

Dale McDiarmid - ClickHouse Meetup

19:00

Chris Battarbee - ClickHouse Meetup

19:30

Rory Crispin - ClickHouse Meetup

20:00

Wrap up

Scan each other's QR codes & head to a nearby pub!
screen 2 • Track 2

11:30

Paging You for Pizza - Powered by incident.io !

Main lobby

12:30

William Mendes

SRE Management is a Hard Job. That’s Why You Should Do It Like an Engineer.

CoralogixWatch
Management feels messy, but it’s just another complex system, full of incidents, dependencies, and feedback loops. In this talk, you’ll learn how to apply engineering principles to leadership: observability, reliability, and iterative improvement for people instead of servers.... Read more

13:00

Christian Melendez

Two Autoscalers Walk into a Cluster: KEDA & Karpenter on Day-2 Duty

Two autoscalers enter, one cluster wins! Master KEDA + Karpenter Day-2 magic to dynamically scale pods AND nodes. Cut waste, handle spikes, save $$$. Walk away with configs that make your cluster be as efficient and cost-optimized as possible.... Read more

13:30

Ehsan Khodadadi

AISRE: It’s Time for AI Site Reliability Engineering

Your API monitoring was green. Dashboards calm. Then a quiet spike: cost per task up 40%, grounded answer rate down 8%, and users start regenerating responses twice as often. Infra metrics say “all good” , but the model silently shifted behavior after a prompt tweak plus a vendor embedding update. Non-AI-adopted SRE doesn’t page you here. ​AISRE would. As AI-powered systems move into production, many teams discover that traditional Site Reliability Engineering metrics​,latency, availability, and error rates​ are no longer sufficient to describe real system health. AI isn’t just predictable APIs anymore. We’re shipping probabilistic systems: prompts → retrieval → model decoding → agents → filters → feedback loops. Every layer can drift independently… and still return a 200 OK. In this talk, ​I introduce AI Site Reliability Engineering (AISRE): an extension of SRE principles tailored specifically for AI-driven systems. ​I explore how reliability must expand to include semantic correctness, grounding quality, safe tool execution, economic efficiency, and controlled behavioral drift.... Read more

14:00

Yury Lysak

SRE for pricing systems in Wealth & Fintech: Data‑Driven Solutions.

IndependentWatch
How to use simple but powerful data concepts to design smarter, more transparent pricing for wealth and fintech products without over engineering.... Read more

14:30

Networking and sponsor crawl

Main lobby

15:00

Andy Kuszyk

MCP is the new REST: making MCP our new API

TypeformWatch
Typeform is fast becoming an AI-native platform. This means we're transforming our REST API into a platform that LLMs can easily discover and operate. In this talk, I'll share the architectural patterns we're using to achieve this, as well as real life examples of how we're making MCP our new API.... Read more

15:30

Amila Mahaarachchi

Building Abstractions That Matter: A Developer Platform on Kubernetes

WSO2Watch
Kubernetes gives us powerful primitives, but a collection of primitives does not automatically become a developer platform. Platform teams are often left stitching together CI systems, GitOps, portals, observability stacks, and policy engines, only to expose the resulting complexity to developers. This talk shares the architectural thinking behind OpenChoreo, an open-source, modular platform that introduces higher-level abstractions on top of Kubernetes and other CNCF projects. I'll walk through how OpenChoreo separates concerns using control, data, CI, and observability planes, and how this approach reduces developer cognitive load while preserving strong governance for platform engineers. You'll learn how to design meaningful abstractions, when to hide and when to expose Kubernetes primitives, and how to balance developer experience with platform control. OpenChoreo brings together development workflows, a Backstage-powered portal, CI/CD, GitOps, and observability without turning Kubernetes into a developer-facing API. If you're building or struggling with an internal developer platform, this talk is for you.... Read more

16:00

Juan Pablo Vidal Araya

From ransomware hostage to leader in digital government: the rebirth of Santo Domingo

Ilustre Municipalidad de Santo DomingoWatch
A Chilean municipality was brought to a standstill by a ransomware attack and emerged as a benchmark in digital government. This talk reveals how Santo Domingo turned a cyber crisis into a national model for cybersecurity and digital transformation.... Read more

16:30

Nicolas Takashi

Reverse-Engineering PromQL Usage: A Proxy’s Tale

CoralogixWatch
Prometheus shows metrics but hides how queries are used. We built Prom Analytics Proxy to reveal who runs which PromQL queries, what slows them down, and which dashboards overload the backend, all without touching Prometheus, giving teams real visibility and control.... Read more

17:00

Happy Hour by Imply - grab a beer!

Main lobby

17:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
screen 3 • Track 3

11:30

Paging You for Pizza - Powered by incident.io !

Main lobby

12:30

Goran Minov

Zero Infrastructure, Zero Phishing: Building a Serverless Security Framework on GCP

OktaWatch
This is an intermediate talk suitable for backend developers and cloud architects. It includes a code walkthrough of the Cloud Run Functions and a data analysis segment comparing the solution's effectiveness.... Read more

13:00

Jagadeesh Devaraj

Catch Me If You Can: Hunting Misconfigurations Before They Break Prod

Outages don’t start in production—they start with a misconfiguration no one noticed. Join Jagadeesh Devaraj (JD) at SREday as he reveals how teams can detect misconfigs, drift, and risky changes beforethey become incidents. Discover a practical, platform‑agnostic approach using Policy‑as‑Code, CI/CD sensors, and Closed Control Loops (CCL) to turn early signals into action. This session distills real lessons, emerging patterns, and powerful strategies to help engineers build reliability by default—without slowing delivery. If you care about catching issues earlier, reducing blast radius, and preventing the next “worked‑on‑my‑machine” disaster, this talk is for you. Come learn how to hunt misconfigurations before they hunt you.... Read more

13:30

Israel Blancas

From Data to Diagnosis: Leveraging Observability for Application Success

CoralogixWatch
This talk combines multiple open source projects and show how some real environments are environments are observed on production: OpenTelemetry, Prometheus, Jaeger, Istio, Kiali and Kubernetes. What else can you ask for?... Read more

14:00

Charles Weir

Overcome Disasters Using Cyber Continuity

Lancaster UniversityWatch
Major software disasters may be almost inevitable, but organisations and ecosystems can survive them. Learn how cyber continuity techniques can prepare your systems to limit damage and support recovery.... Read more

14:30

Networking and sponsor crawl

Main lobby

15:00

Spiros Economakis

How AI is redefining SRE and Customer Experience

NOFire AIWatch
AI is reshaping SRE, but the real opportunity isn’t faster RCA, it’s building proactive, customer-centered reliability. This talk explores how AI and culture together move us from firefighting to foresight.... Read more

15:30

Poone Mokari

Listen to Production the Way It Deserves

Ewake.aiWatch
Production today is messy. There’s noise, complexity, and a constant stream of change. And while we’ve come a long way with observability, it still leans heavily on human foresight. Logs, metrics, alerts, they’re all things we had to think of ahead of time. But when we don’t? That’s where blind spots are born. Ambient agents try to shift that model. These are always-on, proactive teammates who don’t wait for a prompt. They listen to everything happening in production. They surface things we’d likely miss. In this talk, we’ll dive into what it takes to bring an ambient agent into your stack, how it listens, learns, and acts, and why this might just be the layer of intelligence your system’s been missing.... Read more

16:00

Ozan Kasikci

Zero-Downtime Updates: How to Evolve Fast Without Breaking Everything

Longhorn GamesWatch
How do you ship multiple times a week without your users noticing — except for improvements? This talk explores modern zero-downtime deployment strategies and observability patterns that let teams move fast while keeping uptime close to 100%.... Read more

16:30

Dmitrii Iniutin

Beyond Terraform: Building Production Infrastructure with General-Purpose Languages

SeamflowWatch
Infrastructure as Code has transformed how we manage systems — but the tools we use define how far we can scale. While Terraform remains the industry standard, its domain-specific language limits flexibility and maintainability at scale. In this session, Dmitry shares lessons from building production infrastructure at two fast-scaling startups — Fuse and Seamflow — using AWS CDK and Pulumi, frameworks that leverage general-purpose programming languages for infrastructure management. He’ll cover what made these tools more powerful, how they improved testing and reusability, and what teams should know before adopting them.... Read more

17:00

Happy Hour by Imply - grab a beer!

Main lobby

17:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time screen 1 screen 2 screen 3
09:00 KeynoteDecoupled Observability - an architecture for scalable detection and investigation
Peter Marshall • Imply
09:30 KeynoteSecure by Default: Building Confidence in AI-Driven Delivery
Dewan Ahmed • Harness
10:00 KeynoteThe Immune System for Software: Lessons from Biology
Matt Henderson • Phoebe
10:30 KeynoteDo Metrics Matter?
Tyler Hannan • ClickHouse
11:00 KeynoteWhen Incidents Fix Themselves: AI SRE in action
Birol Yildiz • ilert
11:30 Paging You for Pizza - Powered by incident.io !
12:30 Reliability starts at the source: Why Customer Data is your most underrated SRE dependency
Deniz Yalcin & William Ravensbergen • ING Netherland & ING Germany
SRE Management is a Hard Job. That’s Why You Should Do It Like an Engineer.
William Mendes • Coralogix
Zero Infrastructure, Zero Phishing: Building a Serverless Security Framework on GCP
Goran Minov • Okta
13:00 Observability is a Team Sport!
Adriana Villela • Dynatrace
Two Autoscalers Walk into a Cluster: KEDA & Karpenter on Day-2 Duty
Christian Melendez • AWS
Catch Me If You Can: Hunting Misconfigurations Before They Break Prod
Jagadeesh Devaraj • ING
13:30 Choose Your Weapon: The Performance Testing Arsenal
Heather Thacker • Gatling
AISRE: It’s Time for AI Site Reliability Engineering
Ehsan Khodadadi • ING
From Data to Diagnosis: Leveraging Observability for Application Success
Israel Blancas • Coralogix
14:00 From Learner to Leader: My Journey into SRE and Why Reliability Is a Culture, Not a Role
Tasmia Niazi
SRE for pricing systems in Wealth & Fintech: Data‑Driven Solutions.
Yury Lysak • Independent
Overcome Disasters Using Cyber Continuity
Charles Weir • Lancaster University
14:30 Networking and sponsor crawl
15:00 Demo: Let any engineer on any team handle any issue anywhere in your stack
Kyle Forster, Lydia Thomas, Anusha Gundala, Sravan E & Sonu Kumar Singh • RunWhen & ex-Google SRE
MCP is the new REST: making MCP our new API
Andy Kuszyk • Typeform
How AI is redefining SRE and Customer Experience
Spiros Economakis • NOFire AI
15:30 Stranger Teams: Keeping Demogorgons out of production
Rob Charlwood • Supercharged
Building Abstractions That Matter: A Developer Platform on Kubernetes
Amila Mahaarachchi • WSO2
Listen to Production the Way It Deserves
Poone Mokari • Ewake.ai
16:00 Surviving the Dual-Write Problem in Distributed Systems
Sohan Maheshwar • AuthZed
From ransomware hostage to leader in digital government: the rebirth of Santo Domingo
Juan Pablo Vidal Araya • Ilustre Municipalidad de Santo Domingo
Zero-Downtime Updates: How to Evolve Fast Without Breaking Everything
Ozan Kasikci • Longhorn Games
16:30 Honey, the Audience Broke My App: Reproduce & Fix Live in Kubernetes with mirrord
Jake Page • MetalBear
Reverse-Engineering PromQL Usage: A Proxy’s Tale
Nicolas Takashi • Coralogix
Beyond Terraform: Building Production Infrastructure with General-Purpose Languages
Dmitrii Iniutin • Seamflow
17:00 Happy Hour by Imply - grab a beer!
17:30 Meet & Greet
_ClickHouse Meetup - After SREday
Wrap up Wrap up
18:00 Uncovered: The Hard Truth About OpenTelemetry's Vendor Neutrality
Adriana Villela - ClickHouse Meetup • Dynatrace
18:30 Observability Updates in ClickHouse
Dale McDiarmid - ClickHouse Meetup • ClickHouse
19:00 How we built an SRE Agent with Clickhouse and eBPF
Chris Battarbee - ClickHouse Meetup • Metoro
19:30 Running Grafana at Scale
Rory Crispin - ClickHouse Meetup • ClickHouse
20:00 Wrap up

Speakers

_ClickHouse Meetup - After SREday
Adriana Villela
Dynatrace
Adriana Villela - ClickHouse Meetup
Dynatrace
Amila Mahaarachchi
WSO2
Andy Kuszyk
Typeform
Birol Yildiz
ilert
Charles Weir
Lancaster University
Chris Battarbee - ClickHouse Meetup
Metoro
Christian Melendez
AWS
Dale McDiarmid - ClickHouse Meetup
ClickHouse
Deniz Yalcin
& William Ravensbergen
ING Netherland & ING Germany
Dewan Ahmed
Harness
Dmitrii Iniutin
Seamflow
Ehsan Khodadadi
ING
Goran Minov
Okta
Heather Thacker
Gatling
Israel Blancas
Coralogix
Jagadeesh Devaraj
ING
Jake Page
MetalBear
Juan Pablo Vidal Araya
Ilustre Municipalidad de Santo Domingo
Kyle Forster,
Lydia Thomas,
Anusha Gundala,
Sravan E
& Sonu Kumar Singh
RunWhen & ex-Google SRE
Matt Henderson
Phoebe
Nicolas Takashi
Coralogix
Ozan Kasikci
Longhorn Games
Peter Marshall
Imply
Poone Mokari
Ewake.ai
Rob Charlwood
Supercharged
Rory Crispin - ClickHouse Meetup
ClickHouse
Sohan Maheshwar
AuthZed
Spiros Economakis
NOFire AI
Tasmia Niazi
Tyler Hannan
ClickHouse
William Mendes
Coralogix
Yury Lysak
Independent

Venue

Everyman Canary Wharf

Crossrail Place,
Canary Wharf,
E14 5AR, London, UK
Level -2

Tube access
Jubilee, Elizabeth and DLR lines: Canary Wharf station

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one