SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.
Teams are being asked to retain more data, investigate problems and incidents further back in time, and respond faster — all while controlling costs. The challenge isn't a lack of tools; it's the tightly coupled architectures.
In this session, we'll explore why decoupling the data layer from the interaction layer is critical. Using a real-world use case, we'll see how traditional approaches force teams to trade off visibility, performance, or usability at the worst possible moment.
We'll break down how a decoupled architecture allows teams to run continuous detections on recent data while elastically scaling across large historical datasets, without duplicating data or abandoning existing workflows.
The result is greater flexibility, better cost control, and faster incident response — even as data volumes and retention requirements continue to grow.... Read more
The fastest way to break trust in DevSecOps is to automate insecurity at scale. As AI takes a central role in our pipelines, it is time to rethink what "secure by default" really means.
In this keynote, Dewan Ahmed will challenge the audience to look beyond vulnerability scanners and compliance gates. He will share a vision for intelligent security by design, where native intelligence within the delivery platform detects not only vulnerable code but risky delivery behavior such as misconfigured environments, suspicious artifact provenance, and drift between source and runtime.
You will walk away with a framework for balancing automation with human oversight and examples from Harness’ work on building verifiable, auditable, AI-native delivery systems. In the new world of DevSecOps, safety is not a step; it is an outcome we continuously learn to improve.... Read more
Most reliability work today is still centred on reactive troubleshooting: diagnosing a multitude of alerts, pulling large groups into incidents, and engineers scrambling to understand what's happening. To truly change that pattern, we need systems that can predict and prevent failures before they occur. Biological immune systems offer a powerful blueprint for how software can defend and ultimately heal itself.
This talk introduces a framework for thinking about "software immunity", highlights the gaps in today's observability and includes some "under the hood" details on designing AI agents for reliability.
Innate immunity comes from built-in defences like testing, feature flags, auto-scaling, and circuit breakers — mechanisms that provide immediate, general protection. Adaptive immunity, meanwhile, emerges from AI agents that learn from new data, refine their understanding of system behaviour, and apply those lessons to predict and pre-emptively fix failures.
We'll break down the key ingredients for trustworthy AI agents in reliability and beyond: transparent reasoning rather than opaque black-box outputs; strong control mechanisms and guardrails; a governed data layer for effective data access; continuous learning from each execution cycle; and graduated autonomy — from suggestions, to human-in-the-loop actions, to fully automated remediation.... Read more
When I get paged, I open the metrics dashboard. That hasn’t changed. Metrics are still the fastest way to get a rough sense of whether a system is unhealthy, especially when you’re dealing with known failure modes and issues you can reasonably anticipate ahead of time. But with the increase of automated tooling and exploratory analysis, we are faced with a question...do metrics matter?... Read more
The next evolution of incident response isn’t faster alerts, it’s autonomous resolution. Join ilert CEO Birol Yildiz as he shows how AI SRE agents now diagnose and remediate outages without waking anyone up. Learn how these systems combine observability data, deployment context, and code intelligence to restore services in minutes and hand over clean incident reports instead of 3 a.m. pages.... Read more
Reliability is often framed as an infrastructure problem, but at Banking, it begins with data. Customer data is a core SRE dependency that drives data quality, fraud prevention, knowing your customer, and user trust. When upstream data is delayed, malformed, duplicated, or semantically inconsistent, no amount of autoscaling or incident response can preserve a seamless experience.
This talk explores why customer data quality, lineage and observability must be treated as first class reliability concerns. By shifting reliability thinking left toward data consistency, validation, and monitoring at ingestion, fintech teams can reduce incident blast radius, improve MTTR, and protect customer confidence where it matters most: at the source.... Read more
Observability may have been a newcomer few years ago, but it's safe to say that it's a pretty well-established part of our tech lives today. But as organizations embrace Observability, they run the risk of creating Yet Another Silo, in much the same way that DevOps created more silos, instead of breaking them.
Reliability can’t happen without Observability, and Observability itself must be looked at holistically. It isn't the responsibility of one single team, but instead weaves its way into multiple teams. In this talk, Adriana dives into the roles and responsibilities that Development, Quality Assurance, CI/CD, and SRE teams when it comes to contributing to the Observability story. Spoiler alert - it's not as straightforward as you might think!
She will also show how an “Observability team”, if not designed and rolled out properly, can take away an organization’s collective responsibility for Observability, and dilute the promise of Observability.... Read more
Your app works great on your laptop, in the dev environment. Then production hits 10x expected traffic during a marketing campaign and everything falls apart. Or maybe not, instead six months of data accumulates, causing the response times to be painfully slow. Load testing, stress testing, soak testing, and spike testing, they all sound similar, but address completely different problems, and most teams only do one, if at all.
This talk breaks down the essential types of performance testing every app and web developer should understand. Learn when to use each approach, what problems they uncover, and how to integrate them into your development workflow without drowning in complexity. We'll cover real-world scenarios where each testing type saves production systems, helping you choose the right weapon for your performance battles.
What you'll learn:
- The differences between load, stress, soak, and spike testing, and when each matters
- Which performance testing types reveal which production problems, before they happen
- How to integrate performance testing into CI/CD without slowing down development and deployments
- Practical criteria for deciding which tests your application actually needs... Read more
Explore how SRE is more than tools and dashboards: it is a mindset, a culture, and a way of thinking that transforms teams. Attendees will learn how adopting an SRE state of mind empowers individuals, improves engineering outcomes, and builds high-trust, resilient organisations.... Read more
14:30
Networking and sponsor crawl
Main lobby
15:00
Kyle Forster, Lydia Thomas, Anusha Gundala, Sravan E & Sonu Kumar Singh
Production incidents and operational toil have been steadily rising for years, but the worst is ahead. The drivers are clear - i) growing complexity of modern tech stacks, ii) shrinking ratios of SRE/DevOps/QA engineers per developer, and iii) the freight train of ultra-fast AI coding.
Before weekly releases with tens of thousands of lines of code become the new normal, SRE leaders need to think differently about how issues and incidents are handled through their software development lifecyle.
In today's session, Kyle will show the PR that broke his team's traditional SDLC, and talk about ways teams are transitioning from legacy models of SRE to modern models. Hint: get any engineer on any team ready to handle any issue anywhere in the tech stack in any environment. In parallel, Lydia will drop in to a production environment she has never seen before and prioritize issues, drive to root cause, file detailed tickets and resolve incidents in <30 minutes using tools built with RunWhen's AI SRE platform.... Read more
Build SRE teams like 80s kids on bikes! With high trust, shared missions, honesty, and the ability to tackle monsters together. Learn why your SRE might be broken and how to fix it with psychological safety, teamwork, and better practices.... Read more
Dual writes are one of the most common sources of data inconsistency in distributed systems. This talk demystifies why dual writes cannot be made safe inside a single process, then walks through modern architectural patterns that teams use in production to eliminate or mitigate the problem.... Read more
A room of engineers breaks my live app and we fix it together. They scan a QR code which triggers a real failure in a Kubernetes staging environment. Together we reproduce, trace, patch, and validate the fix with kubectl debug and mirrord. A real breakage repaired live.... Read more
Management feels messy, but it’s just another complex system, full of incidents, dependencies, and feedback loops. In this talk, you’ll learn how to apply engineering principles to leadership: observability, reliability, and iterative improvement for people instead of servers.... Read more
Two autoscalers enter, one cluster wins! Master KEDA + Karpenter Day-2 magic to dynamically scale pods AND nodes. Cut waste, handle spikes, save $$$. Walk away with configs that make your cluster be as efficient and cost-optimized as possible.... Read more
Your API monitoring was green. Dashboards calm. Then a quiet spike: cost per task up 40%, grounded answer rate down 8%, and users start regenerating responses twice as often. Infra metrics say “all good” , but the model silently shifted behavior after a prompt tweak plus a vendor embedding update. Non-AI-adopted SRE doesn’t page you here. AISRE would.
As AI-powered systems move into production, many teams discover that traditional Site Reliability Engineering metrics,latency, availability, and error rates are no longer sufficient to describe real system health.
AI isn’t just predictable APIs anymore. We’re shipping probabilistic systems: prompts → retrieval → model decoding → agents → filters → feedback loops. Every layer can drift independently… and still return a 200 OK.
In this talk, I introduce AI Site Reliability Engineering (AISRE): an extension of SRE principles tailored specifically for AI-driven systems. I explore how reliability must expand to include semantic correctness, grounding quality, safe tool execution, economic efficiency, and controlled behavioral drift.... Read more
How to use simple but powerful data concepts to design smarter, more transparent pricing for wealth and fintech products without over engineering.... Read more
Typeform is fast becoming an AI-native platform. This means we're transforming our REST API into a platform that LLMs can easily discover and operate. In this talk, I'll share the architectural patterns we're using to achieve this, as well as real life examples of how we're making MCP our new API.... Read more
Kubernetes gives us powerful primitives, but a collection of primitives does not automatically become a developer platform. Platform teams are often left stitching together CI systems, GitOps, portals, observability stacks, and policy engines, only to expose the resulting complexity to developers.
This talk shares the architectural thinking behind OpenChoreo, an open-source, modular platform that introduces higher-level abstractions on top of Kubernetes and other CNCF projects. I'll walk through how OpenChoreo separates concerns using control, data, CI, and observability planes, and how this approach reduces developer cognitive load while preserving strong governance for platform engineers.
You'll learn how to design meaningful abstractions, when to hide and when to expose Kubernetes primitives, and how to balance developer experience with platform control. OpenChoreo brings together development workflows, a Backstage-powered portal, CI/CD, GitOps, and observability without turning Kubernetes into a developer-facing API.
If you're building or struggling with an internal developer platform, this talk is for you.... Read more
A Chilean municipality was brought to a standstill by a ransomware attack and emerged as a benchmark in digital government. This talk reveals how Santo Domingo turned a cyber crisis into a national model for cybersecurity and digital transformation.... Read more
Prometheus shows metrics but hides how queries are used. We built Prom Analytics Proxy to reveal who runs which PromQL queries, what slows them down, and which dashboards overload the backend, all without touching Prometheus, giving teams real visibility and control.... Read more
17:00
Happy Hour by Imply - grab a beer!
Main lobby
17:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
This is an intermediate talk suitable for backend developers and cloud architects. It includes a code walkthrough of the Cloud Run Functions and a data analysis segment comparing the solution's effectiveness.... Read more
Outages don’t start in production—they start with a misconfiguration no one noticed.
Join Jagadeesh Devaraj (JD) at SREday as he reveals how teams can detect misconfigs, drift, and risky changes beforethey become incidents.
Discover a practical, platform‑agnostic approach using Policy‑as‑Code, CI/CD sensors, and Closed Control Loops (CCL) to turn early signals into action. This session distills real lessons, emerging patterns, and powerful strategies to help engineers build reliability by default—without slowing delivery.
If you care about catching issues earlier, reducing blast radius, and preventing the next “worked‑on‑my‑machine” disaster, this talk is for you.
Come learn how to hunt misconfigurations before they hunt you.... Read more
This talk combines multiple open source projects and show how some real environments are environments are observed on production: OpenTelemetry, Prometheus, Jaeger, Istio, Kiali and Kubernetes. What else can you ask for?... Read more
Major software disasters may be almost inevitable, but organisations and ecosystems can survive them. Learn how cyber continuity techniques can prepare your systems to limit damage and support recovery.... Read more
AI is reshaping SRE, but the real opportunity isn’t faster RCA, it’s building proactive, customer-centered reliability. This talk explores how AI and culture together move us from firefighting to foresight.... Read more
Production today is messy. There’s noise, complexity, and a constant stream of change. And while we’ve come a long way with observability, it still leans heavily on human foresight. Logs, metrics, alerts, they’re all things we had to think of ahead of time. But when we don’t? That’s where blind spots are born.
Ambient agents try to shift that model. These are always-on, proactive teammates who don’t wait for a prompt. They listen to everything happening in production. They surface things we’d likely miss.
In this talk, we’ll dive into what it takes to bring an ambient agent into your stack, how it listens, learns, and acts, and why this might just be the layer of intelligence your system’s been missing.... Read more
How do you ship multiple times a week without your users noticing — except for improvements? This talk explores modern zero-downtime deployment strategies and observability patterns that let teams move fast while keeping uptime close to 100%.... Read more
Infrastructure as Code has transformed how we manage systems — but the tools we use define how far we can scale. While Terraform remains the industry standard, its domain-specific language limits flexibility and maintainability at scale.
In this session, Dmitry shares lessons from building production infrastructure at two fast-scaling startups — Fuse and Seamflow — using AWS CDK and Pulumi, frameworks that leverage general-purpose programming languages for infrastructure management. He’ll cover what made these tools more powerful, how they improved testing and reusability, and what teams should know before adopting them.... Read more
17:00
Happy Hour by Imply - grab a beer!
Main lobby
17:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
Crossrail Place,
Canary Wharf,
E14 5AR, London, UK
Level -2
Tube access
Jubilee, Elizabeth and DLR lines: Canary Wharf station
Sponsors & Partners
Want to become a sponsor? Get in touch!
Deniz Yalcin & William Ravensbergen
ING Netherland & ING Germany
Reliability starts at the source: Why Customer Data is your most underrated SRE dependency
Abstract
Reliability is often framed as an infrastructure problem, but at Banking, it begins with data. Customer data is a core SRE dependency that drives data quality, fraud prevention, knowing your customer, and user trust. When upstream data is delayed, malformed, duplicated, or semantically inconsistent, no amount of autoscaling or incident response can preserve a seamless experience.
This talk explores why customer data quality, lineage and observability must be treated as first class reliability concerns. By shifting reliability thinking left toward data consistency, validation, and monitoring at ingestion, fintech teams can reduce incident blast radius, improve MTTR, and protect customer confidence where it matters most: at the source.
Bio
Deniz Yalcin is head of Customer Data at ING Germany, leading teams responsible for customer data processes and platform integrity. With a background in SRE and IT service management, coaching, and large scale organizational change, she works at the intersection of technology, governance, and leadership.
William Ravensbergen is a Developer Advocate at ING, where he focuses on building and nurturing a global SRE minded community. With a long background in DevOps, SRE, and IT operations, he brings practical experience from running large scale banking platforms into community building, advocacy, and reliability culture.
Adriana Villela
Dynatrace
Observability is a Team Sport!
Abstract
Observability may have been a newcomer few years ago, but it's safe to say that it's a pretty well-established part of our tech lives today. But as organizations embrace Observability, they run the risk of creating Yet Another Silo, in much the same way that DevOps created more silos, instead of breaking them.
Reliability can’t happen without Observability, and Observability itself must be looked at holistically. It isn't the responsibility of one single team, but instead weaves its way into multiple teams. In this talk, Adriana dives into the roles and responsibilities that Development, Quality Assurance, CI/CD, and SRE teams when it comes to contributing to the Observability story. Spoiler alert - it's not as straightforward as you might think!
She will also show how an “Observability team”, if not designed and rolled out properly, can take away an organization’s collective responsibility for Observability, and dilute the promise of Observability.
Bio
Adriana Villela is a Principal Developer Advocate at Dynatrace, where she focuses on cloud native technologies, observability, and helping developers build and operate reliable systems at scale. She is a CNCF Ambassador and a maintainer for the OpenTelemetry End User SIG, actively contributing to the open source ecosystem and community education.
In addition to her advocacy work, Adriana hosts the Geeking Out Podcast, where she explores technical topics and career stories with practitioners across the industry. She also writes regularly on her personal Medium blog, sharing practical insights drawn from hands-on experience. Based in Toronto, she works remotely and remains deeply engaged with global developer communities.
Heather Thacker
Gatling
Choose Your Weapon: The Performance Testing Arsenal
Abstract
Your app works great on your laptop, in the dev environment. Then production hits 10x expected traffic during a marketing campaign and everything falls apart. Or maybe not, instead six months of data accumulates, causing the response times to be painfully slow. Load testing, stress testing, soak testing, and spike testing, they all sound similar, but address completely different problems, and most teams only do one, if at all.
This talk breaks down the essential types of performance testing every app and web developer should understand. Learn when to use each approach, what problems they uncover, and how to integrate them into your development workflow without drowning in complexity. We'll cover real-world scenarios where each testing type saves production systems, helping you choose the right weapon for your performance battles.
What you'll learn:
- The differences between load, stress, soak, and spike testing, and when each matters
- Which performance testing types reveal which production problems, before they happen
- How to integrate performance testing into CI/CD without slowing down development and deployments
- Practical criteria for deciding which tests your application actually needs
Bio
Developer Advocate with a background in software engineering.
Tasmia Niazi
From Learner to Leader: My Journey into SRE and Why Reliability Is a Culture, Not a Role
Abstract
Explore how SRE is more than tools and dashboards: it is a mindset, a culture, and a way of thinking that transforms teams. Attendees will learn how adopting an SRE state of mind empowers individuals, improves engineering outcomes, and builds high-trust, resilient organisations.
Bio
Tasmia Niazi is a passionate and dedicated professional with a background in technology and managing projects from initiation to post implementation support. Tasmia has successfully navigated her career while pursuing higher education, mentoring and managing family life. She recently completed her Masters Apprenticeship in Software Engineering all while balancing the challenges of pregnancy raising a toddler and working. With a keen interest in innovation, Tasmia thrives in fast paced environments contributing to the tech community, speaking at Hackathons, mentoring juniors to arranging platform away days. She is a strong advocate for the empowerment of women in technology and continuously inspires those around her with resilience and passion. Mentoring, raising awareness for women in technology and being an Ambassador are the most aspiring roles Tasmia plays.
Kyle Forster, Lydia Thomas, Anusha Gundala, Sravan E & Sonu Kumar Singh
RunWhen & ex-Google SRE
Demo: Let any engineer on any team handle any issue anywhere in your stack
Abstract
Production incidents and operational toil have been steadily rising for years, but the worst is ahead. The drivers are clear - i) growing complexity of modern tech stacks, ii) shrinking ratios of SRE/DevOps/QA engineers per developer, and iii) the freight train of ultra-fast AI coding.
Before weekly releases with tens of thousands of lines of code become the new normal, SRE leaders need to think differently about how issues and incidents are handled through their software development lifecyle.
In today's session, Kyle will show the PR that broke his team's traditional SDLC, and talk about ways teams are transitioning from legacy models of SRE to modern models. Hint: get any engineer on any team ready to handle any issue anywhere in the tech stack in any environment. In parallel, Lydia will drop in to a production environment she has never seen before and prioritize issues, drive to root cause, file detailed tickets and resolve incidents in <30 minutes using tools built with RunWhen's AI SRE platform.
Bio
Kyle Forster is the Founder of RunWhen, a company pioneering the use of AI for troubleshooting mission critical applications in dev, test and production environments. Prior to RunWhen, Kyle was the Senior Director for Product Management in Google's Kubernetes team. He is a second time founder, having previously started Big Switch Networks, a pioneer in Software Defined Networking (acquired by Arista Networks). Kyle started his career at Cisco after doing an MBA and MS in Computer Science from Stanford University and an MSE in Electrical Engineering from Princeton University. He currently holds six patents in wireless and software defined networking technologies lives with his wife, three children and two dogs in the greater London area where he recently settled after relocating from Silicon Valley.
Lydia Thomas is a London-based entrepreneur and technology founder. She is the founder of Smash Hit Technologies Ltd and The Keep Us Company Ltd, and previously worked at Google. Lydia is also a member of the Associates’ Committee of The Worshipful Company of International Bankers. Her work sits at the intersection of technology, business, and innovation, with a focus on building and scaling impactful companies.
Anusha Gundala is a Senior Site Reliability Engineer at Sky and a TechWomen100 Awards 2025 winner. She specializes in cloud, Kubernetes, and observability, with a focus on improving reliability, security, and standardization across large scale distributed systems. Her work centers on Google Cloud Platform and modern monitoring stacks including Grafana, where she leads initiatives to enhance system visibility and operational performance. Prior to Sky, she held roles at iomart and Wipro, building experience across cloud infrastructure, monitoring, and operations engineering.
Sravan E is a Solutions Architect specializing in observability, site reliability engineering, and operational resilience with over 18 years of experience. He focuses on designing and implementing systems that improve performance and ensure stability across complex environments. His work includes developing observability architectures, defining SLIs and SLOs, and leveraging tools such as Splunk ITSI, Dynatrace, Datadog, and Elastic. Sravan has held roles across organizations including Schain Technologies, dunnhumby, HSBC, Fiserv, and Accenture, delivering resilient data platforms and advancing AIOps driven automation strategies.
Sonu Kumar Singh is a Cloud Engineering Specialist and Senior Site Reliability Engineer at Lloyds Banking Group with over 15 years of experience designing scalable, secure, and self healing cloud systems. He specializes in Kubernetes, Terraform, and multi cloud environments across Azure and Google Cloud Platform, with a focus on building reliable and automated infrastructure. His work emphasizes applying software engineering principles to solve operational challenges and improve system performance. Prior to Lloyds, he held roles across engineering and DevOps functions, contributing to cloud architecture, CI/CD pipelines, and distributed system resilience.
Rob Charlwood
Supercharged
Stranger Teams: Keeping Demogorgons out of production
Abstract
Build SRE teams like 80s kids on bikes! With high trust, shared missions, honesty, and the ability to tackle monsters together. Learn why your SRE might be broken and how to fix it with psychological safety, teamwork, and better practices.
With nearly 20 years of engineering expertise spanning infrastructure, DevOps, SRE, software, and web development, Rob excels in designing scalable, secure, and high-performance cloud infrastructure. His portfolio boasts delivering exceptional projects for renowned organisations, including John Lewis, Lloyds Bank, Deloitte, BBC, HMRC, Google, The FA Premier League, FIFA, UEFA, Major League Soccer, Shell, and Mars.
A seasoned cloud practitioner, Rob has deep experience with GCP, AWS, and Azure, as well as mastery of Kubernetes, Terraform, CI/CD pipelines, and programming languages like Go, Python, and JavaScript. Beyond technical prowess, he thrives in fast-paced environments, solving complex problems with creativity, adaptability, and a collaborative spirit.
Passionate about knowledge sharing, Rob actively contributes to the Open Source Community, co-organises the Django Bristol and Bath Users Group (DBBUG), and regularly speaks at local and industry events, including Golang Bristol meet ups.
Sohan Maheshwar
AuthZed
Surviving the Dual-Write Problem in Distributed Systems
Abstract
Dual writes are one of the most common sources of data inconsistency in distributed systems. This talk demystifies why dual writes cannot be made safe inside a single process, then walks through modern architectural patterns that teams use in production to eliminate or mitigate the problem.
Bio
Sohan is a Lead Developer Advocate at AuthZed, based in the Netherlands. He started his career as a developer building mobile apps and has been living in the cloud since 2013, in companies such as Amazon, Fermyon and Gupshup. He is also an O' Reilly author, having created a course on Cloud Concepts for Everyone.
He has always been interested in emerging technologies and how it shapes the world around us.
Jake Page
MetalBear
Honey, the Audience Broke My App: Reproduce & Fix Live in Kubernetes with mirrord
Abstract
A room of engineers breaks my live app and we fix it together. They scan a QR code which triggers a real failure in a Kubernetes staging environment. Together we reproduce, trace, patch, and validate the fix with kubectl debug and mirrord. A real breakage repaired live.
Bio
I'm Jake a DevOps engineer turned DevRel.
Over the last 5 years I have been heavily focused on the world of Cloud Native Dev tooling, from cloud FinOps, Packaging and Software Delivery, not exclusively but many times in the context of Kubernetes clusters.
Having transitioned from a previous career as a high school teacher, any chance I get to speak in front of a crowd on topics that I'm passionate about I try to take. I'm a Lisbon resident and love to frequent the local meetup scene. So if you see me around, don't be a stranger and let's chat.
_ClickHouse Meetup - After SREday
Meet & Greet
Abstract
Bio
After SREday wraps up, join us for ClickHouse Meetup, exclusively in screen 1!
Adriana Villela - ClickHouse Meetup
Dynatrace
Uncovered: The Hard Truth About OpenTelemetry's Vendor Neutrality
Abstract
Bio
After SREday wraps up, join us for ClickHouse Meetup, exclusively in screen 1!
Dale McDiarmid - ClickHouse Meetup
ClickHouse
Observability Updates in ClickHouse
Abstract
Bio
After SREday wraps up, join us for ClickHouse Meetup, exclusively in screen 1!
Chris Battarbee - ClickHouse Meetup
Metoro
How we built an SRE Agent with Clickhouse and eBPF
Abstract
Bio
After SREday wraps up, join us for ClickHouse Meetup, exclusively in screen 1!
Rory Crispin - ClickHouse Meetup
ClickHouse
Running Grafana at Scale
Abstract
Bio
After SREday wraps up, join us for ClickHouse Meetup, exclusively in screen 1!
William Mendes
Coralogix
SRE Management is a Hard Job. That’s Why You Should Do It Like an Engineer.
Abstract
Management feels messy, but it’s just another complex system, full of incidents, dependencies, and feedback loops. In this talk, you’ll learn how to apply engineering principles to leadership: observability, reliability, and iterative improvement for people instead of servers.
Bio
William Mendes is an Engineering Leader with over 15 years of experience designing, scaling, and leading high-performance systems and teams.
Christian Melendez
AWS
Two Autoscalers Walk into a Cluster: KEDA & Karpenter on Day-2 Duty
Abstract
Two autoscalers enter, one cluster wins! Master KEDA + Karpenter Day-2 magic to dynamically scale pods AND nodes. Cut waste, handle spikes, save $$$. Walk away with configs that make your cluster be as efficient and cost-optimized as possible.
Bio
Christian Melendez is Principal Specialist Solutions Architect and EMEA Lead for Compute at AWS, with a strong background in Kubernetes platform engineering. Author of the Kubernetes Autoscaling book. He has been working with Kubernetes since 2017, helping large enterprises—including telecommunications, airline, and ride-hailing companies—optimize their workloads. Christian is the creator of the Karpenter Blueprints project and an active contributor to autoscaling solutions in the cloud-native space. He frequently delivers talks and workshops on Karpenter and Kubernetes optimization strategies.
Ehsan Khodadadi
ING
AISRE: It’s Time for AI Site Reliability Engineering
Abstract
Your API monitoring was green. Dashboards calm. Then a quiet spike: cost per task up 40%, grounded answer rate down 8%, and users start regenerating responses twice as often. Infra metrics say “all good” , but the model silently shifted behavior after a prompt tweak plus a vendor embedding update. Non-AI-adopted SRE doesn’t page you here. AISRE would.
As AI-powered systems move into production, many teams discover that traditional Site Reliability Engineering metrics,latency, availability, and error rates are no longer sufficient to describe real system health.
AI isn’t just predictable APIs anymore. We’re shipping probabilistic systems: prompts → retrieval → model decoding → agents → filters → feedback loops. Every layer can drift independently… and still return a 200 OK.
In this talk, I introduce AI Site Reliability Engineering (AISRE): an extension of SRE principles tailored specifically for AI-driven systems. I explore how reliability must expand to include semantic correctness, grounding quality, safe tool execution, economic efficiency, and controlled behavioral drift.
Bio
Ehsan Khodadadi is a Senior Site Reliability Engineer at ING, with extensive experience leading and building reliability practices across large-scale systems. Before rejoining ING, he led the Site Reliability Engineering team at LeasePlan, where he focused on system stability, team management, and operational excellence. His background includes roles at Techspire and BMW Group, where he combined deep technical expertise in DevOps and Linux systems with a pragmatic, hands-on approach to problem solving. Ehsan is known for creating strong engineering teams and improving service reliability through thoughtful automation and collaboration
Yury Lysak
Independent
SRE for pricing systems in Wealth & Fintech: Data‑Driven Solutions.
Abstract
How to use simple but powerful data concepts to design smarter, more transparent pricing for wealth and fintech products without over engineering.
Bio
Seasoned leader with 15+ years of experience driving growth, innovation, and operational excellence across tech, telecom, fintech and other sectors. Head of Strategic Projects at Lionsoul Global, leading the launch of investment and lending products on the platform and managing strategic partnerships.
Andy Kuszyk
Typeform
MCP is the new REST: making MCP our new API
Abstract
Typeform is fast becoming an AI-native platform. This means we're transforming our REST API into a platform that LLMs can easily discover and operate. In this talk, I'll share the architectural patterns we're using to achieve this, as well as real life examples of how we're making MCP our new API.
Bio
Andy Kuszyk is a Staff Engineer at Typeform, an occasional blog author, and an Emacs enthusiast! He's mostly worked in backend systems with the likes of Go, Terraform, and Kubernetes, but also enjoys thinking about the way teams work, communicate, and transfer knowledge.
Amila Mahaarachchi
WSO2
Building Abstractions That Matter: A Developer Platform on Kubernetes
Abstract
Kubernetes gives us powerful primitives, but a collection of primitives does not automatically become a developer platform. Platform teams are often left stitching together CI systems, GitOps, portals, observability stacks, and policy engines, only to expose the resulting complexity to developers.
This talk shares the architectural thinking behind OpenChoreo, an open-source, modular platform that introduces higher-level abstractions on top of Kubernetes and other CNCF projects. I'll walk through how OpenChoreo separates concerns using control, data, CI, and observability planes, and how this approach reduces developer cognitive load while preserving strong governance for platform engineers.
You'll learn how to design meaningful abstractions, when to hide and when to expose Kubernetes primitives, and how to balance developer experience with platform control. OpenChoreo brings together development workflows, a Backstage-powered portal, CI/CD, GitOps, and observability without turning Kubernetes into a developer-facing API.
If you're building or struggling with an internal developer platform, this talk is for you.
Bio
Amila heads the engineering team at WSO2. With over 15 years in the software industry—14 of them at WSO2—he has played a key role in shaping cloud-based solutions. Amila has led multiple teams, including WSO2 Cloud and Choreo, driving innovation in software development, deployment, and cloud operations. His extensive expertise spans the entire software lifecycle, from development to scalable cloud infrastructure.
Juan Pablo Vidal Araya
Ilustre Municipalidad de Santo Domingo
From ransomware hostage to leader in digital government: the rebirth of Santo Domingo
Abstract
A Chilean municipality was brought to a standstill by a ransomware attack and emerged as a benchmark in digital government. This talk reveals how Santo Domingo turned a cyber crisis into a national model for cybersecurity and digital transformation.
Bio
I am Juan Pablo Vidal, a professional passionate about technology, innovation, and management. For more than seven years, I have supported public and private organizations in turning ideas into real, measurable services, always with a clear focus on digital transformation, cybersecurity, regulatory compliance, and personal data protection.
I am motivated by identifying opportunities for improvement, optimizing processes, and designing value propositions that simplify people’s lives. I enjoy both strategy, aligning objectives, resources, and reference frameworks, and execution, because I believe a project only truly matters when it moves from paper into practice and delivers concrete results.
In every challenge, I bring together technological curiosity, rigor in compliance, and a user-centered perspective. In this way, from the first spark of an idea to the day the solution is fully operational, I strive to ensure that innovation is secure, efficient, and above all, useful to those who need it.
Nicolas Takashi
Coralogix
Reverse-Engineering PromQL Usage: A Proxy’s Tale
Abstract
Prometheus shows metrics but hides how queries are used. We built Prom Analytics Proxy to reveal who runs which PromQL queries, what slows them down, and which dashboards overload the backend, all without touching Prometheus, giving teams real visibility and control.
Bio
Nicolas is a Software Engineer with a Platform Engineer role at Coralogix. He's mostly interested in topics related to the observability ecosystem, as well as Kubernetes and distributed systems. He is also an open-source contributor to projects such as Prometheus Operator, Perses, and OpenTelemetry.
Goran Minov
Okta
Zero Infrastructure, Zero Phishing: Building a Serverless Security Framework on GCP
Abstract
This is an intermediate talk suitable for backend developers and cloud architects. It includes a code walkthrough of the Cloud Run Functions and a data analysis segment comparing the solution's effectiveness.
Bio
Goran Minov is a Cyber Security Architect and 8x Certified Professional, holding key Google Cloud credentials including Professional Cloud Architect, DevOps, and Database Engineer. Specialising in Identity & Access Management (IAM), he currently serves as a Senior Technical Account Manager at Okta, where he helps customers automate and secure their identity infrastructure.
A passionate community leader, Goran has been an Organiser for GDG London since 2019, helping grow one of the largest developer communities in the UK. He is also the Founder of The Cloud Circuit, an independent event series connecting London’s cloud community, and a Co-founder of The Android Circuit. Beyond organising, he actively mentors the next generation of developers through the GDG Academy.
As a speaker, Goran regularly shares his expertise on Cloud Security and Zero Trust Architecture. He has delivered workshops and talks at major venues including Droidcon London and DevFest.
Jagadeesh Devaraj
ING
Catch Me If You Can: Hunting Misconfigurations Before They Break Prod
Abstract
Outages don’t start in production—they start with a misconfiguration no one noticed.
Join Jagadeesh Devaraj (JD) at SREday as he reveals how teams can detect misconfigs, drift, and risky changes beforethey become incidents.
Discover a practical, platform‑agnostic approach using Policy‑as‑Code, CI/CD sensors, and Closed Control Loops (CCL) to turn early signals into action. This session distills real lessons, emerging patterns, and powerful strategies to help engineers build reliability by default—without slowing delivery.
If you care about catching issues earlier, reducing blast radius, and preventing the next “worked‑on‑my‑machine” disaster, this talk is for you.
Come learn how to hunt misconfigurations before they hunt you.
Bio
Jagadeesh Devaraj is an Engineering Lead for SRE and Engineering Excellence at ING, with over 17 years of experience in cloud architecture and platform engineering. He has led large scale resilience, security, and cloud modernization initiatives across financial services, technology, and global enterprises.
Israel Blancas
Coralogix
From Data to Diagnosis: Leveraging Observability for Application Success
Abstract
This talk combines multiple open source projects and show how some real environments are environments are observed on production: OpenTelemetry, Prometheus, Jaeger, Istio, Kiali and Kubernetes. What else can you ask for?
Bio
Software Engineer @ Coralogix doing observability stuff. Google Developer Expert in Google Cloud.
Charles Weir
Lancaster University
Overcome Disasters Using Cyber Continuity
Abstract
Major software disasters may be almost inevitable, but organisations and ecosystems can survive them. Learn how cyber continuity techniques can prepare your systems to limit damage and support recovery.
Bio
I am passionate about improving the effectiveness of software development teams. From working with me, more than a hundred project teams have become more productive, been more reliable, and had more fun.
As mentor, technical lead, manager, author or consultant, I deliver. Whether the need is for team agreement, technical change management, software architecture or improved processes, my engaging approach and positive drive get the results needed.
I helped introduce object-oriented and agile methods to the UK; was technical lead for the world’s first smartphone; and led a company providing outsourced app development for fifteen years. That company, Penrillian, was praised both for effective delivery of superb software and for being a great place to work. My book, Small Memory Software, is the only general purpose guide for people working with software memory limitations.
I am currently based near England's Lake District, and working with one of Britain's top academic Software Security teams. We’re creating techniques to help software developers to create the secure software we need for the 21st century.
Spiros Economakis
NOFire AI
How AI is redefining SRE and Customer Experience
Abstract
AI is reshaping SRE, but the real opportunity isn’t faster RCA, it’s building proactive, customer-centered reliability. This talk explores how AI and culture together move us from firefighting to foresight.
Bio
Spiros Economakis is the founder and CEO of NOFire.ai, where he’s rethinking how AI can help engineering teams build reliable systems that don’t just recover faster but fail less often.
He’s spent over fifteen years as an engineer and leader across startups and large-scale platforms, bridging the gap between DevOps, SRE, and AI. His writing on Reliability Engineering has influenced how teams think about ownership, causality, and the future of observability.
Spiros speaks and writes with a rare blend of technical depth and human perspective always focused on one thing: how teams can understand systems deeply enough to build with confidence.
Poone Mokari
Ewake.ai
Listen to Production the Way It Deserves
Abstract
Production today is messy. There’s noise, complexity, and a constant stream of change. And while we’ve come a long way with observability, it still leans heavily on human foresight. Logs, metrics, alerts, they’re all things we had to think of ahead of time. But when we don’t? That’s where blind spots are born.
Ambient agents try to shift that model. These are always-on, proactive teammates who don’t wait for a prompt. They listen to everything happening in production. They surface things we’d likely miss.
In this talk, we’ll dive into what it takes to bring an ambient agent into your stack, how it listens, learns, and acts, and why this might just be the layer of intelligence your system’s been missing.
Bio
Pooné Mokari is the CEO and co-founder of Ewake.ai, an AI Reliability Teammate on a mission to bring real peace of mind to engineering teams. Drawing on her experience as an SRE at Criteo, she founded Ewake to offer engineers their dream teammate, which investigates issues reactively and watches production proactively. Throughout her career, she was active as a speaker in different tech conferences, such as Devoxx Belgium and Devoxx France. She’s also been engaged in mentoring women in tech.
Ozan Kasikci
Longhorn Games
Zero-Downtime Updates: How to Evolve Fast Without Breaking Everything
Abstract
How do you ship multiple times a week without your users noticing — except for improvements? This talk explores modern zero-downtime deployment strategies and observability patterns that let teams move fast while keeping uptime close to 100%.
Bio
I'm an entrepreneur and software engineer with more than 12 years of experience. I worked in four different startups before co‑founding Longhorn Games where we build mobile games. My background spans backend development and DevOps, and I've built CI/CD pipelines, test automation systems, and cloud infrastructure for servers, databases, and networks at various startups. This experience gives me deep insight into startup culture and the technical challenges companies face as they grow.
Dmitrii Iniutin
Seamflow
Beyond Terraform: Building Production Infrastructure with General-Purpose Languages
Abstract
Infrastructure as Code has transformed how we manage systems — but the tools we use define how far we can scale. While Terraform remains the industry standard, its domain-specific language limits flexibility and maintainability at scale.
In this session, Dmitry shares lessons from building production infrastructure at two fast-scaling startups — Fuse and Seamflow — using AWS CDK and Pulumi, frameworks that leverage general-purpose programming languages for infrastructure management. He’ll cover what made these tools more powerful, how they improved testing and reusability, and what teams should know before adopting them.
Bio
Dmitrii Iniutin is a founding engineer at Seamflow, where he works across the full engineering stack to build and scale core product systems. Before that, he was a lead engineer at Fuse Energy, joining as one of the earliest technical hires and helping the company grow into a global renewable-energy unicorn. He built critical components including the trading desk, billing engine, APIs, notification platform, and the entire infrastructure backbone. Previously, Dmitrii led engineering at hicebank, rapidly advancing to leadership and assembling a new team while shaping the system architecture. He began his career at Yandex, developing large-scale video streaming infrastructure, and earlier interned at JetBrains.
Peter Marshall
Imply
KeynoteDecoupled Observability - an architecture for scalable detection and investigation
Abstract
Teams are being asked to retain more data, investigate problems and incidents further back in time, and respond faster — all while controlling costs. The challenge isn't a lack of tools; it's the tightly coupled architectures.
In this session, we'll explore why decoupling the data layer from the interaction layer is critical. Using a real-world use case, we'll see how traditional approaches force teams to trade off visibility, performance, or usability at the worst possible moment.
We'll break down how a decoupled architecture allows teams to run continuous detections on recent data while elastically scaling across large historical datasets, without duplicating data or abandoning existing workflows.
The result is greater flexibility, better cost control, and faster incident response — even as data volumes and retention requirements continue to grow.
Bio
Peter Marshall is a technology leader and community builder with a background in developer relations, data architecture, and digital transformation. As Director of Developer Relations at Imply, he leads programs that grow and engage the global Apache Druid community through education, support, and events. With experience across startups, enterprises, and the public sector, Peter brings a blend of technical expertise and strategic vision to help organizations connect with developers and drive impact through open source.
Dewan Ahmed
Harness
KeynoteSecure by Default: Building Confidence in AI-Driven Delivery
Abstract
The fastest way to break trust in DevSecOps is to automate insecurity at scale. As AI takes a central role in our pipelines, it is time to rethink what "secure by default" really means.
In this keynote, Dewan Ahmed will challenge the audience to look beyond vulnerability scanners and compliance gates. He will share a vision for intelligent security by design, where native intelligence within the delivery platform detects not only vulnerable code but risky delivery behavior such as misconfigured environments, suspicious artifact provenance, and drift between source and runtime.
You will walk away with a framework for balancing automation with human oversight and examples from Harness’ work on building verifiable, auditable, AI-native delivery systems. In the new world of DevSecOps, safety is not a step; it is an outcome we continuously learn to improve.
Bio
Dewan Ahmed is a Principal Developer Advocate at Harness and a Governing Board General Member Representative at the Continuous Delivery Foundation. He focuses on DevRel and content strategy across CI/CD, DevOps, and open source, with deep expertise in software supply chain security and developer experience.
Matt Henderson
Phoebe
KeynoteThe Immune System for Software: Lessons from Biology
Abstract
Most reliability work today is still centred on reactive troubleshooting: diagnosing a multitude of alerts, pulling large groups into incidents, and engineers scrambling to understand what's happening. To truly change that pattern, we need systems that can predict and prevent failures before they occur. Biological immune systems offer a powerful blueprint for how software can defend and ultimately heal itself.
This talk introduces a framework for thinking about "software immunity", highlights the gaps in today's observability and includes some "under the hood" details on designing AI agents for reliability.
Innate immunity comes from built-in defences like testing, feature flags, auto-scaling, and circuit breakers — mechanisms that provide immediate, general protection. Adaptive immunity, meanwhile, emerges from AI agents that learn from new data, refine their understanding of system behaviour, and apply those lessons to predict and pre-emptively fix failures.
We'll break down the key ingredients for trustworthy AI agents in reliability and beyond: transparent reasoning rather than opaque black-box outputs; strong control mechanisms and guardrails; a governed data layer for effective data access; continuous learning from each execution cycle; and graduated autonomy — from suggestions, to human-in-the-loop actions, to fully automated remediation.
Bio
Matt Henderson is the co-founder and CEO of Phoebe, an AI agent platform for software reliability. Previously, he was the CEO of Stripe Europe, and led the company’s international operations across product and engineering. Earlier in his career, Matt was a product director at Amazon and Google, and co-founded the ML analytics startup Rangespan (acquired by Google). He is also an angel investor in over 100 startups.
Tyler Hannan
ClickHouse
KeynoteDo Metrics Matter?
Abstract
When I get paged, I open the metrics dashboard. That hasn’t changed. Metrics are still the fastest way to get a rough sense of whether a system is unhealthy, especially when you’re dealing with known failure modes and issues you can reasonably anticipate ahead of time. But with the increase of automated tooling and exploratory analysis, we are faced with a question...do metrics matter?
Bio
Tyler Hannan is Senior Director of Developer Advocacy at ClickHouse and a longtime leader in developer relations, community, and product storytelling. He specializes in translating complex technical systems into clear narratives that connect engineers, products, and business strategy.
Birol Yildiz
ilert
KeynoteWhen Incidents Fix Themselves: AI SRE in action
Abstract
The next evolution of incident response isn’t faster alerts, it’s autonomous resolution. Join ilert CEO Birol Yildiz as he shows how AI SRE agents now diagnose and remediate outages without waking anyone up. Learn how these systems combine observability data, deployment context, and code intelligence to restore services in minutes and hand over clean incident reports instead of 3 a.m. pages.
Bio
Birol Yildiz is the Co-founder and CEO of ilert, adeptly steering the company with a rare combination of technical and product expertise. His prior experience includes a significant role as Chief Product Owner for Big Data products at REWE Digital. With a strong foundation in computer science, Birol bridges the gap between developer and product strategist, constantly striving to innovate and provide customer-centric solutions at ilert.