SREday

Site Reliability, DevOps and Cloud

October 28, 2025 Ilert, Cologne, Germany

1
Day
12+
Speakers
1
Track
50+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

Booking.com, Cisco, ClickHouse, Hyground, ilert, Paymenttools, VictoriaMetrics

Topics so far:

We'll announce more talks soon, stay tuned!

This is a past event, what's next?

Schedule

October 28, 2025 single track 10AM - 3PM Cologne, in-person
view as table
main room • Track 1

10:00

Birol Yildiz

KeynoteWhen Incidents Fix Themselves: AI SRE in action

ilert
The next evolution of incident response isn’t faster alerts, it’s autonomous resolution. Join ilert CEO Birol Yildiz as he shows how AI SRE agents now diagnose and remediate outages without waking anyone up. Learn how these systems combine observability data, deployment context, and code intelligence to restore services in minutes and hand over clean incident reports instead of 3 a.m. pages.... Read more

10:30

Coffee break

Main lobby

11:00

Jannes Wolf & Dominik Rehbock

Agents in Production: The Best Ways to Shoot Yourself in the Foot

Hyground
This is a tongue in cheek talk about the pitfalls of letting AI Agents interact with highly critical systems and sensitive data. With new capabilities also come new vulnerabilities. With help from OWASP LLM Top 10 we’ll walk through three real-world failure scenarios: * Agents making catastrophic but 'innocent' mistakes * Customer Data leaking into AI Lab training data * Prompt injections enabling deliberate exfiltration... Read more

11:30

Maxim Schepelin

How to set SLOs, drive improvements, and make friends with business stakeholders

Booking.com
We, tech people, have internalized the concept of reliability so deeply that we don't need an explanation for why it's bad to have services failing in production. It doesn't matter what your software is doing — whether it controls train schedules, allows people to make money transfers, or serves funny pictures online. We have invented an entire language to talk about reliability. We say things like, "This service runs with three nines availability," "Mean-Time-To-Repair for the website is 25 minutes," or "This month, we consumed 80% of our error budget." And yet, it's a common struggle to convince business stakeholders to prioritize technical improvements. In the past seven years, I've been responsible for products and services used by millions of people. I have been on call for years and have spent many nights resolving production incidents. From that experience, I learned how to make reliability a shared priority and prove, with data, the value of technical improvements. In this session, I'll share techniques you can use to improve the reliability of your services, convince business stakeholders of the importance of reliability, and drive positive change in your organization.... Read more

12:00

Vlad Seliverstov

How We Built ClickStack - an open source, open telemetry native Observability stack

ClickHouse
Modern observability is built on a flawed foundation: three siloed pillars - logs, metrics, and traces - each powered by different engines with separate query models, storage formats, and operational costs. Users are forced to manually correlate across systems, accept duplication, or pay high SaaS bills. But what if observability is just a data problem? One that needs a general-purpose solution instead of purpose-built compromises? This talk argues that true observability requires fast, high-cardinality queries over unsampled data at scale and at low cost. Traditional search and metrics engines were not designed for this, but column-oriented databases are. We introduce ClickStack, a fully open source, OpenTelemetry-native observability stack built around ClickHouse. It provides fast, flexible querying and efficient storage, enabling real-time visibility without compromise.... Read more

12:30

Rajat Gupta

Predict, Don’t Page: AI for Incident Forecasting and Safe Auto-Remediation

Paymenttools
Modern SRE teams drown in alerts that show up after damage is done. This talk is a practical blueprint for going proactive: predicting incidents before they page you, and triggering safe, reversible auto-remediation.... Read more

13:00

Lunch & networking

Main lobby

14:00

Jose Gomez-Selles

Modern Observability with OpenTelemetry in C++

VictoriaMetrics
Do you still think that traces are those debug level logs to enable when something goes wrong? Do you think that this modern observability thingy is just another buzzword for other languages like javascript or go? Do you believe that "std::cout" is the best way to debug everything? Well, I do agree up to some point. But it does not scale that easily. For modern, distributed systems, you will need to adopt modern observability. The path seems too complicated. But you are not alone! In this talk we'll demystify how to simply create metrics and distributed traces from a bunch of logs by using the OpenTelemetry standard and libraries. Metrics included too! After describing the basics, Jose will also share an example he built to monitor many transactions in production without impacting performance. After this talk, you will be able to talk with higher confidence about OpenTelemetry, Jaeger, distributed tracing, Prometheus or VictoriaMetrics.... Read more

14:30

Ricard Bejarano

Speeding up Terraform caching with OverlayFS

Cisco
The Terraform plugin cache, unfortunately, does not support concurrent Terraform inits. This is a massive efficiency and performance loss for those who use Terraform at a big enough scale, since we're left with the following choices: - (A) Disable caching, and download gigabytes of providers off the Terraform Registry on every single init. - (B) Serialize all terraform init runs so they can safely share the cache, reducing the Terraform pipeline's throughput (and infuriating developers, been there, done that). We didn't like our options here, so we got creative. One day, everything clicked: OverlayFS! OverlayFS is a Linux filesystem which combines the contents of multiple read-only directories and one writable layer on top, into a single volume, effectively implementing a sort of read-through cache within your filesystem. If you've ever used a live linux distribution, this is what those use. Containers use OverlayFS. CoreOS used OverlayFS to provide ephemeral /etc directories. You rarely see such a low-level tool this high up the stack, so what is doing here? We tweaked our workflow so that our plugin cache is mounted once per init using OverlayFS, which makes the centralized cache read-only (and thus, concurrent-safe) and gives Terraform a writable overlay where it can write whatever providers were missing. We then added a final, non-blocking step to feed back those new providers to the central cache (in a serial fashion using filesystem locking), effectively implementing a write-back cache. After testing it on the lab, we promoted it to our production pipeline. After a couple rounds of plans to fill up the cache, we saw stunning results. Plan times dropped by -61%. Our concurrent plan capacity 10x'd because we were no longer saturating both the network (pulling) and the disk (writing providers). And after all, we didn't add much complexity to the setup anyway. We're using long-standing, reliable, kernel tech (overlayfs and flock) to address a shortcoming of a higher level tool like Terraform.... Read more

15:00

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room
10:00 Keynote: When Incidents Fix Themselves: AI SRE in action
Birol Yildiz • ilert
10:30 Coffee break
11:00 Agents in Production: The Best Ways to Shoot Yourself in the Foot
Jannes Wolf & Dominik Rehbock • Hyground
11:30 How to set SLOs, drive improvements, and make friends with business stakeholders
Maxim Schepelin • Booking.com
12:00 How We Built ClickStack - an open source, open telemetry native Observability stack
Vlad Seliverstov • ClickHouse
12:30 Predict, Don’t Page: AI for Incident Forecasting and Safe Auto-Remediation
Rajat Gupta • Paymenttools
13:00 Lunch & networking
14:00 Modern Observability with OpenTelemetry in C++
Jose Gomez-Selles • VictoriaMetrics
14:30 Speeding up Terraform caching with OverlayFS
Ricard Bejarano • Cisco
15:00 Wrap up

Speakers

Birol Yildiz
ilert
Jannes Wolf
& Dominik Rehbock
Hyground
Jose Gomez-Selles
VictoriaMetrics
Maxim Schepelin
Booking.com
Rajat Gupta
Paymenttools
Ricard Bejarano
Cisco
Vlad Seliverstov
ClickHouse

Venue

The offices of Ilert.com

Bayenstraße 65,
50678 Köln, Germany

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one