SREday

Site Reliability, DevOps and Cloud

November 7, 2025 ING Cedar, Amsterdam, Netherlands

1
Day
25+
Speakers
2
Tracks
200+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

adjoe, Booking.com, Catchpoint, ClickHouse, Datadog, Funambol, Imply, ING, Inter Cars, LinearB, Lumigo, MeloMar IT, Microsoft, OGD ict-diensten, PagerDuty, SolarWinds, Tanzu, VictoriaMetrics

Topics so far:
Cloud Infrastructure
...and more

This is a past event, what's next?

Schedule

November 7, 2025 2 parallel tracks 9:30AM - 6PM Amsterdam, in-person
view as table
main room • Track 1

09:30

Guido Smit

KeynoteThe accountability problem

ING
When everyone owns reliability, does anyone truly feel accountable? Lets explore why extreme ownership is the cornerstone of resilient systems.... Read more

10:00

Peter Marshall

KeynoteWhat Observability Can Learn From BI: Decoupling for Speed, Scale, and Flexibility

Imply
Today’s observability platforms are often vertically integrated—binding data storage, query, and visualization layers into a single stack. This tight coupling drives up costs, makes integrations painful, and slows teams down. But it doesn’t have to be this way. In this talk, we’ll explore how SRE teams can benefit from a more modular approach to observability—one inspired by the evolution of Business Intelligence. Just as BI stacks evolved to separate ETL, data warehouses, and dashboards, observability stacks can be designed around clear boundaries: interoperable tools, technology-neutral query layers, and plug-and-play storage. You’ll learn why decoupled observability architecture is essential for cost control, agility, and tool flexibility—and how to move toward a stack that meets the real-world needs of today’s SRE teams.... Read more

10:30

Tasmia Niazi, Ehsan Khodadadi, Roheel I & Bas van Alphen

Panel Discussion

Discover the reliability tech stack and practices behind top European institutions.... Read more

11:00

Coffee break

Main lobby

11:30

Lukasz Groszkowski

SRE at Scale: Keeping e‑Commerce Alive Across 20+ Countries

Inter Cars
A deep dive into how we ensure reliability and availability in a complex distributed environment, with lessons learned from real-world incidents and rollouts.... Read more

12:00

Diana Todea

Cutting Through Metrics Cardinality Noise with VictoriaMetrics

VictoriaMetrics
In high-scale environments, metrics cardinality isn’t just a resource concern, it’s an architectural challenge. Left unchecked, it can impact performance, query latency, and even system stability. This talk takes a deep technical dive into how VictoriaMetrics enables advanced observability practices with a strong focus on cardinality management. I’ll explore how to design efficient and scalable scrape configurations using Prometheus-compatible jobs and exporters, optimize your label strategies, and use built-in cardinality analysis tools within VictoriaMetrics to identify and mitigate high-cardinality patterns early. The session also covers integration with Grafana open source for visualizing metrics in a way that supports signal clarity and operational response, as well as setting up practical alerting strategies that minimize noise while ensuring fast issue detection. By the end of the talk, the audience will understand the engineering trade-offs of high-cardinality metrics and how to detect them, how VictoriaMetrics handles storage and querying at scale, how to build resilient and low-overhead scrape configs with Prometheus compatibility and how to use Grafana open source to highlight cardinality hot spots and improve alert signal quality.... Read more

12:30

Yishai Beeri

Your AI Code Reviews Are Missing the Point (And How to Fix It)

LinearB
Most AI code review implementations focus on the wrong metrics—counting comments generated or code accepted rather than measuring developer velocity and code quality improvements. The real value lies in intelligent context integration and organizational learning at scale. This talk examines successful AI code review deployments across engineering organizations, revealing four critical success factors: seamless integration with your existing development context (codebase, tickets, architectural decisions), specialized review guidelines that scale from 5 to 50,000 repositories, comprehensive observability to understand where AI adds value versus where it creates noise, and strategic human-AI collaboration patterns. We'll explore real case studies of teams who've moved beyond basic linting to utilizing AI code review that understands business logic, catches architectural anti-patterns, and actually accelerates PR cycles. You'll learn when AI reviews shine, when they don't, and how to measure the difference. This isn't about replacing human reviewers—it's about building an intelligent system that amplifies human expertise and reduces cognitive load where it matters most.... Read more

13:00

Lunch & networking

Main lobby

14:00

Dale McDiarmid

How We Built ClickStack - an open source, open telemetry native Observability stack

ClickHouse
Modern observability is built on a flawed foundation: three siloed pillars - logs, metrics, and traces - each powered by different engines with separate query models, storage formats, and operational costs. Users are forced to manually correlate across systems, accept duplication, or pay high SaaS bills. But what if observability is just a data problem? One that needs a general-purpose solution instead of purpose-built compromises? This talk argues that true observability requires fast, high-cardinality queries over unsampled data at scale and at low cost. Traditional search and metrics engines were not designed for this, but column-oriented databases are. We introduce ClickStack, a fully open source, OpenTelemetry-native observability stack built around ClickHouse. It provides fast, flexible querying and efficient storage, enabling real-time visibility without compromise.... Read more

14:30

Michael Cote

Platform Engineering and AI - Two Buzzwords Finally Meet!

Tanzu
Two Buzzwords Finally Meet! How should we manage AI in large organizations? What are the "services" developers need to add AI to enterprise apps, and what role do platform engineers take? Where do data scientists fit in? How does MCP, A2A, and whatever the latest AI API is fit in? There are so many questions because the field of managing AI in enterprises barely exists. We need to figure it out quickly however, before we see a repeat of a past patterns, like shadow AI and unmanageable apps in production. Based on real-world examples, this talk will go over what we currently know about using platform engineering to help developers use AI.... Read more

15:00

Alessandro Vozza

The USB-c for your Copilot: Securing MCP Servers with API Management

Microsoft
The Model Context Protocol (MCP) has seen a meteoric rise to become the de-facto connecting mycelium for GenAI models and applications. In this session, we will explore the critical role of API management in securing MCP servers. Learn how to leverage API management as the "USB-c" connector for your Copilot, ensuring seamless and secure integration. We will cover best practices, real-world examples, and practical tips to enhance your server security and streamline operations. Join us to discover how API management can be the key to unlocking robust security for your MCP servers.... Read more

15:30

Networking & sponsor crawl

Main lobby

16:00

Kevin van der Vlist

From Incident Response to Preventive Mitigation: Leveraging CodeQL and LLMs at Scale

ING
After recovering from incidents, we must ensure that effective mitigation measures are in place to prevent similar issues in the future. However, clearly expressing the characteristics of past problems and identifying similar occurrences across large codebases is a significant challenge, especially at the scale of ING. We present our approach to encoding known and ING-specific (anti)patterns as CodeQL queries. These queries are used to assess both individual and multiple repositories, enabling us to pinpoint risks and reach out directly to the responsible teams, rather than broadcasting generic warnings to everyone. A key challenge is the volume of findings: a single query can yield hundreds of results, overwhelming engineering teams if not managed carefully. To address this, we define a process that not only assesses findings but also generates actionable patches and contextual advice. By integrating Large Language Models (LLMs), we combine the specific context of each code finding with the surrounding source code and high-level problem descriptions, resulting in tailored, practical recommendations for engineers. We translate incident learnings and ING-specific anti-patterns into CodeQL queries, making them reusable and scalable. For each finding, we use LLMs to generate context-aware advice and, where possible, code patches. This ensures recommendations are both actionable and relevant to the code owner. By combining static analysis with LLM-driven advice, we turn past mistakes into valuable lessons, foster professional growth, and continuously improve our engineering community. We believe this approach is broadly applicable to other large organizations seeking to scale preventive mitigation.... Read more

16:30

Daniel Afonso

Full Service Ownership & The Lifecycle of a Service

PagerDuty
Services are the backbone of our systems. They are the pieces that make up our businesses—whether they are literal microservices or functional components of a traditional application, we can’t do the computer thing without services. When it comes to a service in your company or organization, who’s responsible for it? The cast of characters involved in the lifecycle of a service is more than just software engineers. It can include program managers, product owners, sustainability teams (SREs/operations engineers), and business stakeholders, just to name a few. Topics covered in this talk include: – Defining what a service means to you and your organization – Roles in service ownership – What are you observing about your service? – How you want a team to respond to a service outage or incident – Managing your service in production – Tuning your service – Understanding how the service impacts the business... Read more

17:00

Marius Kimmina

From Spot Ocean to Karpenter - One Year Later

adjoe
Over One year of operating Karpenter in production across multiple Kubernetes Clusters has taught us at adjoe many valuable lessons that we want to share in this talk. We will talk about how we handled the migration process, why we build a custom controller to handle broken nodes as well as how we have optimized our nodepools to avoid too frequent disruptions and much more... Read more

17:30

Leo Visser

Secure your cloud automation

OGD ict-diensten
So you decided to automate tasks in your Azure cloud platform? Are you sure this automation is secured properly and not causing any more attack surfaces for malicious actors? In this talk I will cover different ways of automating your cloud platform and risks involved with these techniques. I will go over topics like authentication, secrets, auditability and more. I’ll also show situation I encountered in the field and how these were fixed to provide more confidentiality, integrity and availability. So if you want to learn to automate your could platform with a zero trust policy, then join this talk! Afterward you will know common mistakes and how to avoid them while also being capable of setting up proper automation in the cloud!... Read more

18:00

Wrap up

Scan each other's QR codes & head to a nearby pub!
UPSTAIRS room • Track 2

11:00

Coffee break

Main lobby

11:30

Leon Adato

Observability tools don't suck, you just have too many

Catchpoint
Does your organization collect observability solutions like they're pokemon cards? You're not alone. This happens teams think it's easier to buy their own tool than share. In this talk, we'll cover both WHY it's bad, and WHAT you can do.... Read more

12:00

Joris Bonnefoy

From experimentation to continuous verification: how to benefit from the entire spectrum of Chaos Engineering

Datadog
Chaos Engineering is often misunderstood as simply “breaking things on purpose.” This talk challenges that perception and repositions Chaos Engineering as a critical pillar of reliability and resilience engineering. Rather than focusing on failure injection alone, we explore how to leverage existing knowledge, validate known truths, and foster confidence in complex systems. In the first part, we deconstruct common myths around Chaos Engineering and reframe its core principles. Learn how aligning chaos practices with reliability goals can transform the way your organization perceives and applies these techniques—by emphasizing structured validation over blind experimentation. The second part brings theory into practice with a hands-on framework that reimagines chaos experiments as self-feeding, iterative loops—mirroring the scientific method. We introduce the concept of continuous verification, drawing parallels with integration testing, and show how Chaos Engineering can seamlessly integrate into the Software Development Lifecycle (SDLC) through a shift-left approach. The session wraps up with a visual framework for implementing a sustainable Chaos Engineering strategy, including how to evolve gamedays into repeatable, hypothesis-driven validations that scale with system changes. Whether you're just exploring Chaos Engineering or looking to mature your reliability strategy, this talk will leave you with actionable insights, a modernized mindset, and a clear path to operational resilience.... Read more

12:30

Maxim Schepelin

How to set SLOs, drive improvements, and make friends with business stakeholders

Booking.com
We, tech people, have internalized the concept of reliability so deeply that we don't need an explanation for why it's bad to have services failing in production. It doesn't matter what your software is doing — whether it controls train schedules, allows people to make money transfers, or serves funny pictures online. We have invented an entire language to talk about reliability. We say things like, "This service runs with three nines availability," "Mean-Time-To-Repair for the website is 25 minutes," or "This month, we consumed 80% of our error budget." And yet, it's a common struggle to convince business stakeholders to prioritize technical improvements. In the past seven years, I've been responsible for products and services used by millions of people. I have been on call for years and have spent many nights resolving production incidents. From that experience, I learned how to make reliability a shared priority and prove, with data, the value of technical improvements. In this session, I'll share techniques you can use to improve the reliability of your services, convince business stakeholders of the importance of reliability, and drive positive change in your organization.... Read more

13:00

Lunch & networking

Main lobby

14:00

Marcel Koert

It’s Not the Tools - It’s Us: How Human Biases Undermine Reliability

MeloMar IT
In the world of DevOps and SRE, we pride ourselves on automation, observability, and engineering excellence. But even the most sophisticated infrastructure can be derailed by something far more human: our own brains. This talk explores the invisible enemies of reliability — cognitive biases that affect how engineers make decisions, solve problems, and respond to incidents. From confirmation bias during outages to groupthink in postmortems, we’ll unpack six specific mental shortcuts that silently undermine your systems. Each bias is brought to life with real-world engineering examples, memorable visuals, and practical strategies you can apply immediately. You’ll hear relatable war stories — like how anchoring on a bad hypothesis prolonged a major incident, or how optimism bias led to a Friday deployment that ended in disaster. More than a lecture, this is a wake-up call. We don’t need more YAML linters or dashboards. We need to debug ourselves. By the end of the session, attendees will: • Understand how specific biases impact reliability and team performance • Recognize symptoms of biased thinking in real-world DevOps workflows • Learn actionable techniques to create more resilient, self-aware engineering cultures This talk is ideal for DevOps engineers, SREs, team leads, and reliability advocates looking to take their incident management and system design to the next level — not by adding more tools, but by thinking more clearly.... Read more

14:30

Renato Losio

Nobody Ever Got Fired for Implementing Multi-AZ

Funambol
Using multiple Availability Zones (AZs) is often seen as essential for building resilient and highly available cloud systems. This is true, until it is not. While Multi-AZ is a proven architectural choice, there are important drawbacks to consider and common assumptions that don’t always hold up. There are also different ways to implement it, each with its own trade-offs. In this session, we will explore the Multi-AZ journey, separating fact from fiction. We will discuss when Multi-AZ truly helps, the costs involved, and some surprising side effects that are often overlooked.... Read more

15:00

Yan Cui

Patterns and Practices for Building Resilient AWS Serverless Applications

Lumigo
Lambda provides multi-AZ support out of the box, but even then, things can still go wrong in production. Region-wide outages and performance degradations can render your applications non-responsive. And what if you're dealing with downstream systems that aren't as scalable as your system and can't handle the load you put on them? The bottom line is that many things can go wrong, and they often do at the worst of times. The goal of building resilient systems is not to prevent failures, but to design systems that can withstand them. In this talk, we will examine several practices and architectural patterns that can help you build more resilient serverless architectures, such as multi-region design, employing DLQs and surge queues and cell-based architectures. We'll also explore how chaos experiments can help us identify failure modes before they happen in production.... Read more

15:30

Networking & sponsor crawl

Main lobby

16:00

Simone Romani

Gain confidence in your code with mutation testing

ING
One of the best ways to assess if code is resilient against bugs is to break it on purpose and see how it reacts. The reaction should be a failure in the tests. If there is no reaction, it means that the tests are not effective enough, meaning their assertions are broad and imprecise. Mutation testing comes to the rescue for this specific challenge. This methodology changes the source code and then runs the unit tests against the mutated codebase. The generated report helps the engineer find where the weak spots in the tests are. In this talk we will cover the theory behind this methodology, followed by a live demo where code which could be described as "100% tested" would still be subject to bugs and how its related tests can be improved. This approach will offer a way for engineers to gain confidence in their code and especially in their tests. With a high test strength, source code will not only be strong but also malleable to modifications, with the safe guardrails of unit tests protecting them from introducing bugs. The audience will learn how to write more efficient and resilient tests due to the mutation tests giving them a different perspective on their code quality, compared to the normal tests. Mutation testing will also drive better production code, following the principle of Test Driven Design.... Read more

16:30

Moustafa Aboelnaga

Strategic Observability, How to Design the Right Stack for Your Organization.

SolarWinds
Designing your observability stack the right way requires clarity of purpose and organizational alignment. This session will introduce a strategic framework for setting your observability direction: defining goals, identifying data priorities, and choosing technologies that serve your reliability vision. We’ll explore maturity stages, integration strategies, and the cultural shifts needed to make observability a shared responsibility. It will be a theoretical, high-level session for SRE and DevOps leaders who want to think before they build more than a practical one.... Read more

17:00

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room UPSTAIRS room
09:30 KeynoteThe accountability problem
Guido Smit • ING
10:00 KeynoteWhat Observability Can Learn From BI: Decoupling for Speed, Scale, and Flexibility
Peter Marshall • Imply
10:30 Panel Discussion
Tasmia Niazi, Ehsan Khodadadi, Roheel I & Bas van Alphen
11:00 Coffee break
11:30 SRE at Scale: Keeping e‑Commerce Alive Across 20+ Countries
Lukasz Groszkowski • Inter Cars
Observability tools don't suck, you just have too many
Leon Adato • Catchpoint
12:00 Cutting Through Metrics Cardinality Noise with VictoriaMetrics
Diana Todea • VictoriaMetrics
From experimentation to continuous verification: how to benefit from the entire spectrum of Chaos Engineering
Joris Bonnefoy • Datadog
12:30 Your AI Code Reviews Are Missing the Point (And How to Fix It)
Yishai Beeri • LinearB
How to set SLOs, drive improvements, and make friends with business stakeholders
Maxim Schepelin • Booking.com
13:00 Lunch & networking
14:00 How We Built ClickStack - an open source, open telemetry native Observability stack
Dale McDiarmid • ClickHouse
It’s Not the Tools - It’s Us: How Human Biases Undermine Reliability
Marcel Koert • MeloMar IT
14:30 Platform Engineering and AI - Two Buzzwords Finally Meet!
Michael Cote • Tanzu
Nobody Ever Got Fired for Implementing Multi-AZ
Renato Losio • Funambol
15:00 The USB-c for your Copilot: Securing MCP Servers with API Management
Alessandro Vozza • Microsoft
Patterns and Practices for Building Resilient AWS Serverless Applications
Yan Cui • Lumigo
15:30 Networking & sponsor crawl
16:00 From Incident Response to Preventive Mitigation: Leveraging CodeQL and LLMs at Scale
Kevin van der Vlist • ING
Gain confidence in your code with mutation testing
Simone Romani • ING
16:30 Full Service Ownership & The Lifecycle of a Service
Daniel Afonso • PagerDuty
Strategic Observability, How to Design the Right Stack for Your Organization.
Moustafa Aboelnaga • SolarWinds
17:00 From Spot Ocean to Karpenter - One Year Later
Marius Kimmina • adjoe
Wrap up
17:30 Secure your cloud automation
Leo Visser • OGD ict-diensten
18:00 Wrap up

Speakers

Alessandro Vozza
Microsoft
Dale McDiarmid
ClickHouse
Daniel Afonso
PagerDuty
Diana Todea
VictoriaMetrics
Guido Smit
ING
Joris Bonnefoy
Datadog
Kevin van der Vlist
ING
Leo Visser
OGD ict-diensten
Leon Adato
Catchpoint
Lukasz Groszkowski
Inter Cars
Marcel Koert
MeloMar IT
Marius Kimmina
adjoe
Maxim Schepelin
Booking.com
Michael Cote
Tanzu
Moustafa Aboelnaga
SolarWinds
Peter Marshall
Imply
Renato Losio
Funambol
Simone Romani
ING
Tasmia Niazi,
Ehsan Khodadadi,
Roheel I
& Bas van Alphen
Yan Cui
Lumigo
Yishai Beeri
LinearB

Venue

ING Cedar - Hosting Sponsor

Bijlmerdreef 106
1102 CT Amsterdam, Netherlands

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one