SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.
Adyen, Andersen Lab, AWS, BIMcollab, Booking.com, Cisco, Dash0, DXC Technology, Funambol, Imubit, ING, Krateo, lomo.dev, LoopStudio, Microsoft, NCR Voyix, Netdata, OGD ICT-diensten, PagerDuty, Plumelo, Rabobank, Reliability Engineering Lab, Reliaburger, Spare Cores, SQUER, Uber, Varnish Software, vCluster Labs, VendoQ, VictoriaMetrics
Bijlmerdreef 106
1102 CT Amsterdam, Netherlands



We’ve all been there. It’s 3 a.m., an alert just fired, and you’re staring at a dashboard. The P99 latency is spiking, but the average looks fine. The CPU is high, but the load average is normal. What’s actually going on? The hard truth is that most of our dashboards are subtle liars. They lie by averaging percentiles, hiding the real user pain. They lie with poorly chosen time windows that miss the crucial spike. They lie because we're forced to guess: is this a counter or a gauge? Can I sum this? Is a line graph even the right way to look at this? We expect every engineer on our team to also be a part-time statistician, and at 3 a.m., that’s a recipe for disaster. In this talk, I'll share stories of how misleading dashboards have sent teams down the wrong rabbit hole during critical incidents. We'll dissect the common lies and half-truths our monitoring tells us every day. Then, we'll talk about the alternative: building "opinionated" observability. It's about baking our expertise and best practices directly into our tools, so the right chart appears by default, the alert already has context, and the system guides you to the right questions. This isn't just about making prettier graphs. It's about giving back our most valuable resource time, and letting engineers get back to being engineers, not full-time data analysts. What you'll get from this talk: How to spot the 3 most common ways your dashboards are probably lying to you. A simple mental model for choosing the right visualization for the job (and knowing when a line graph is a terrible choice). Practical ways to start baking your team's expertise and "tribal knowledge" into your monitoring with smart defaults. How to argue for, or build, "opinionated" tools that stop wasting your team's time and lead to faster, more accurate incident response.
Ralph Meijer is a Software Engineer and technical leader with over 20 years of experience building scalable distributed systems, real-time communications platforms, event-driven architectures, and observability solutions. Having served in technical leadership roles across both fast-paced startups and global multinationals, he specializes in connecting complex systems and bridging the gap between engineering teams and cross-functional partners in product, design, and legal. Ralph balances security, reliability, privacy, performance, and maintainability across every stage of system design. Deeply committed to open standards and collaborative engineering, he actively contributes to Open Source communities and serves as Chair of the XMPP Standards Foundation.
Dashboards were invented for a world without AI. They encode assumptions about human cognition, scanning, pattern recognition, correlation, that AI now handles better. This contrarian talk argues that the future of operational observability isn't better dashboards, but AI agents that render dashboards obsolete for the decisions that matter. We'll examine: The dashboard's original sin: Assumptions that worked for small systems but fail at scale What AI-native observability actually looks like: Concrete examples from production systems The transition path: Pragmatic stages from AI augmentation to dashboard sunset The human role in AI-native operations: Why this isn't about replacing humans Based on three years of building AI observability systems, including conversational interfaces that replace manual investigation and ML anomaly detection achieving 10⁻³⁶ false positive rates, this talk challenges you to reconsider whether dashboards are helping or hindering your operational excellence.
Costa Tsaousis is the Founder and CEO of Netdata. Since 1995, Costa has been actively working on internet-related startups. He has been a co-founder and C-level executive of many successful projects, including Internet Service Providers, Cloud Hosting Providers, and Fintech startups. With a passion for innovation and open-source, he now leads Netdata, a monitoring solution aiming to simplify and modernize infrastructure observability for all of us.
during this session we will show Self-Service DevOps Assistant (SDA). SDA is a cloud-agnostic Al assistant built for DevOps, SRE, and platform engineering teams. It connects your cloud infrastructure, CI/CD pipelines, issue trackers, code repositories, and knowledge base into a single interface - and answers engineering questions in natural language, based on what is actually happening in your environment right now
In my current role at Andersen, I am responsible for partnership with major Cloud Services Providers and Neo-cloud to build products and solutions at the intersection of AI, Data and Cloud technologies
Do you trust your AI coding assistant? What if I told you that attackers have found ways to manipulate it and attack your code? With everyone now using AI coding assistants it’s time to look at the risks! During this talk I’ll show you several new techniques attackers are already using. This will range from hidden messages (ASCII smuggling) to abusing mistyping and characters that look the same (typosquatting). I will also show how an LLM can make mistakes when generating code (hallucinations). Did you know that a smart attacker can abuse this too? When you join this talk, you’ll learn how to spot hidden text in your instruction file and prompts. I will also explain how to set up a trusted dependency repository to prevent the malicious code from entering your production environment!
Active in the IT industry since 2012, Leo Visser is a Subject Matter Expert for Azure and AI Foundry at OGD. He advises organizations on AI, automation, and cloud architecture, combining tools like PowerShell, Power Platform, Azure Logic Apps, and native Azure services. His work strongly focuses on security, sustainability, and long-term value. Leo is a Microsoft PowerShell MVP.
How do you plan for unplanned incidents? You practice with Chaos Engineering. Strong incident response doesn’t just happen, you have to build the skills and train your team. Practicing for major incidents gives your team insight into how your applications will behave when something goes wrong as well as how the team will interact to solve problems. Combining your Incident Response practices with Chaos Engineering roots your response practice in real-world scenarios, helping your team build confidence.
Daniel Afonso is a Senior Developer Advocate at PagerDuty, SolidJS DX team member, Instructor at Egghead.io, and Author of State Management with React Query. Daniel has a full-stack background, having worked with different languages and frameworks on various projects from IoT to Fraud Detection. He is passionate about learning and teaching and has spoken at multiple conferences around the world about topics he loves. In his free time, when he's not learning new technologies or writing about them, he's probably reading comics or watching superhero movies and shows.
AI development has a cluster problem: laptop Kubernetes like kind can't host real GPUs, so engineers juggle Docker Compose, full clusters, and scripts that drift from production. We show a pattern where one laptop kubectl context spans local Docker workers and a cloud NVIDIA T4, joined by a single curl over a WireGuard tunnel, so the same workload runs locally and on cloud GPU without two separate setups. The idea is Kubernetes-native: any Linux host (KVM, GCE, bare metal) becomes a node by pasting a one-line curl that installs kubelet and connects over the tunnel, in one tenant cluster. We demo it live: laptop Docker workers, a local KVM node, and a GCE T4 in a single context. We pull llama3.2 onto the T4, deploy a service that talks to it, expose it through a built-in LoadBalancer, and pause/resume the cluster like a laptop. We close with the failure modes (tunnel reconnect after sleep, egress cost on model pulls, GPU eviction) and the GitOps that keeps the topology declarative.
An active contributor to open source, content creator on Medium and YouTube, focusing on practical, scalable solutions in cloud-native environments. DevOps and Platform Engineering practitioner and advocate.
Ephemeral Environments (EEs) replace long-lived shared environments with on-demand, isolated infrastructure. They improve reliability and developer velocity, but they also fundamentally change how infrastructure costs behave. If every pull request spins up its own environment, costs can quickly become unpredictable. What used to be a fixed, always-on expense turns into a dynamic system driven by engineering activity. In this talk, we explore the FinOps implications of adopting Ephemeral Environments from an SRE perspective. Instead of focusing on specific tools or optimizations, we’ll examine how teams rethink cost control when infrastructure becomes short-lived and event-driven. We’ll discuss the trade-offs between environment fidelity and cost, how lifecycle-driven infrastructure changes spending patterns, and how reducing shared environment failures can offset infrastructure growth. This session is about understanding the cost model behind Ephemeral Environments, and how to adopt them without losing financial control.
Marcos Novelli Harispe is a Fullstack Engineer and ORT University Lecturer focused on scalable AWS cloud systems. He specializes in event-driven architectures, Data Lakes, and infrastructure as code (Terraform), with extensive experience migrating production microservices to serverless. Equal parts hands-on engineer and educator, Marcos is passionate about reliability, observability, and mentoring the next generation of tech talent.
Picture this: you start a new role, eager to learn and contribute with your ideas! Your next task is to get familiar with the database setup, and then you start encountering these massive PostgreSQL databases — 100TB, 200TB, 300TB... And you start questioning yourself: how do you backup (and restore) a +100TB database? And how about HA? Performance? Vacuum? It should work the same way as for a 100GB database, right? Well, maybe not exactly. Blog posts and best practice guides make PostgreSQL seem straightforward—until you push it to its limits. At extreme scale, you will find yourself questioning the most fundamental assumptions about how PostgreSQL works. Over the last years, my team at Adyen has been exploring the boundaries of what PostgreSQL can do, and today I will share our findings with you (at least the ones I can!).
Teresa has worked with databases for almost 10 years. After several years as an Oracle DBA, she transitioned to PostgreSQL and never looked back! Extensibility and the Community are the two main aspects she finds fascinating about PostgreSQL. Teresa is also part of the PGDay Lowlands organization. Teresa started her career far away from databases, as a Civil Engineer. So, if you've always wanted to know how tunnels and bridges are built, just ask! In her free time, Teresa enjoys spending time outdoors (mainly admiring rocks), hiking, and cooking.
AI agents are moving from advisory copilots to operational actors capable of calling APIs, changing infrastructure, triggering workflows and making decisions across production environments. Once that happens, the central reliability question is no longer simply whether the model produces a good answer, but who authorized the agent to act, under which identity, with what scope, and how that authority can be observed, constrained and revoked. This session examines the production architecture required for autonomous operations: non-human identity lifecycle and explicit ownership, delegated authorization and least privilege across tools and services, trust boundaries outside the model layer, auditable action trails, human escalation paths, deterministic safety gates and circuit breakers. We will also look at failure containment: what happens when an agent performs a technically valid action that is operationally wrong, and how SRE teams can limit blast radius before autonomy becomes systemic risk. The goal is a practical, product-independent framework for increasing useful autonomy without allowing operational, security and economic risk to increase at the same rate.
Bajric Sanel, Ph.D. is an economist, technologist and senior software and systems engineer with more than 20 years of hands-on experience across software architecture, production infrastructure, AI, cloud/DevOps, cybersecurity, IoT and complex systems. His work focuses on the intersection of production AI, distributed systems, operational reliability, security and technology economics - particularly what happens when autonomous systems move from experimentation into real production environments. He leads technology and delivery across Bright Moon Digital, Studio387 and VendoQ, combining engineering practice with research in economics and technical sciences.
Choosing a cloud instance type for a task is still largely a heuristic exercise. While public pricing and hardware specifications are available, they are fragmented, inconsistently structured, and challenging to compare across (or even within) cloud providers -- especially once real workload performance is taken into account. In this talk, we present Spare Cores Navigator, a Python-queryable and open-source benchmark dataset that covers thousands of cloud server types from multiple vendors, with standardized performance and cost-efficiency metrics. We demonstrate how instance selection can be expressed as a simple data query, e.g. filtering by workload characteristics, hardware or compliance constraints, and budget, then ranking candidates by price-performance.
Gergely Daroczi, PhD, has been a passionate open-source package developer for two decades. With over 15 years in the fintech, adtech, healthtech, and other SaaS industries, he has expertise in data science and engineering, as well as cloud infrastructure, with a focus on building scalable data platforms. Gergely maintains a dozen open-source R and Python projects and organizes a tech meetup with 1,800 members in Hungary -- along with other open-source and data conferences.
Before your first container hits production on Kubernetes you install a distro, a CNI, an ingress, cert-manager, Prometheus, Loki and ArgoCD. None of them is your app. After ten years of running platform teams on Kubernetes, I built Reliaburger to put the lot into one Rust binary: scheduler, gossip, Raft, eBPF service discovery, ingress, mTLS, registry, metrics, logs, GitOps and chaos testing, driven by TOML. An app with health checks, ingress and autoscaling is twelve lines.
This is a hands-on session, so bring a laptop (macOS on Apple silicon or Linux; Windows is not supported yet). In 60 minutes you will:
1. Install Reliaburger with one command and start a node on your laptop. No container runtime, no root, no VMs: the built-in process runtime runs plain binaries.
2. Deploy an app from a twelve-line TOML file and watch it in the terminal dashboard and the web UI.
3. Ship a deliberately sick version and watch the health checker restart it, then roll back with one command.
4. Break it on purpose with the built-in fault injector: kill an instance, freeze another, and let relish wtf explain what just happened.
5. Read the manual and the source code from inside the binary. I will be next to you for the questions the manual does not answer.
6. Provide feedback!
It is 0.1.0. It will break. That is the point, and I want to hear about it.
Miko Pawlikowski has been running platform teams on top of Kubernetes since version 1.0, and has the scars to show for it. He is the author of Chaos Engineering (Manning) and of Reliaburger, a batteries-included container orchestrator written in Rust, along with the book Building Reliaburger that documents how every subsystem was designed and built.
Before your first container hits production on Kubernetes you install a distro, a CNI, an ingress, cert-manager, Prometheus, Loki and ArgoCD. None of them is your app. After ten years of running platform teams on Kubernetes, I built Reliaburger to put the lot into one Rust binary: scheduler, gossip, Raft, eBPF service discovery, ingress, mTLS, registry, metrics, logs, GitOps and chaos testing, driven by TOML. An app with health checks, ingress and autoscaling is twelve lines.
This is a hands-on session, so bring a laptop (macOS on Apple silicon or Linux; Windows is not supported yet). In 60 minutes you will:
1. Install Reliaburger with one command and start a node on your laptop. No container runtime, no root, no VMs: the built-in process runtime runs plain binaries.
2. Deploy an app from a twelve-line TOML file and watch it in the terminal dashboard and the web UI.
3. Ship a deliberately sick version and watch the health checker restart it, then roll back with one command.
4. Break it on purpose with the built-in fault injector: kill an instance, freeze another, and let relish wtf explain what just happened.
5. Read the manual and the source code from inside the binary. I will be next to you for the questions the manual does not answer.
6. Provide feedback!
It is 0.1.0. It will break. That is the point, and I want to hear about it.
Miko Pawlikowski has been running platform teams on top of Kubernetes since version 1.0, and has the scars to show for it. He is the author of Chaos Engineering (Manning) and of Reliaburger, a batteries-included container orchestrator written in Rust, along with the book Building Reliaburger that documents how every subsystem was designed and built.
Every engineer remembers their worst incident. Mine taught me something uncomfortable: our systems were fine within the hour. It was us, the engineers, that took all night. The people scrambled. Since then, one question keeps bugging me: why do we test everything except our own response? We chaos-test infrastructure and load-test services, but the people component remains untested. This talk is about closing that gap through rehearsal. We look through the lens of incident simulations, and I'll share what I've learned: what breaks first in a team under pressure (communication, not competence), why your most experienced engineer can be your biggest liability, and how a team that has practiced its worst day performs when the real one arrives.
Niek is a tech lead at Schuberg Philis. He designs, builds, and runs mission critical infrastructure solutions and is an expert in reliability engineering. Niek’s practical, no-nonsense style and engaging delivery help technical leaders and engineers effectively navigate the complex balance between reliability, security, and delivering business value.
Every SRE team knows their clusters are overprovisioned. But how much exactly, and where? I kept asking this question on our EKS clusters. kubectl top gives percentages, not dollars. AWS Cost Explorer sees instances, not pods. Every tool that connects the two wanted me to deploy Helm charts, agents, and dashboards. I just wanted a number. So I built Burn — an open-source CLI that reads your kubeconfig, fetches real-time pricing from AWS and Azure APIs, and gives you per-namespace cost breakdown in 30 seconds. No agent, no dashboard, no cluster changes. Running it on production, I found 33% idle capacity ($117/month on a 5-node cluster), a pod requesting 500m CPU but using 0.12m, and debug pods nobody remembered deploying. I deleted the waste that same day. In this talk I'll cover: - How Burn splits node cost into CPU and RAM using ratio-based pricing - Why P95 metrics matter more than averages for rightsizing - How we detect Ingress-based load balancers that other tools miss - Honest trade-offs of an agentless approach vs full platforms like Kubecost Attendees will leave knowing how to identify idle resources, understand Kubernetes cost allocation math, and evaluate the right level of cost tooling for their clusters.
Cloud/DevOps Engineer who enjoys solving complex problems and building efficient, scalable systems. My expertise includes Kubernetes, AWS, and CI/CD pipelines, with a strong focus on automating processes to simplify developers' workflows. Additionally, I maintain a supercomputer (HPC) at NEU IBM Center, ensuring it runs at peak performance for high-demand tasks. Creator of Burn, an open-source Kubernetes FinOps CLI.
Modern applications often outgrow simple API key or role-based access models, especially in multi-tenant and microservice environments. This talk explores how OpenFGA enables fine-grained, relationship-based authorization to model complex access patterns such as “user X can access resource Y because they belong to team Z.” We’ll walk through real-world scenarios where traditional RBAC breaks down, demonstrate how OpenFGA implements Google Zanzibar–inspired authorization, and show how to integrate it into cloud-native platforms and APIs. Attendees will leave with practical patterns, architecture guidance, and lessons learned from implementing authorization as a standalone service in modern platform architectures.
Ankit Asthana is a Senior Cloud & Platform Architect Engineer at SQUER, with over 13 years of experience across AI infrastructure, DevOps, SRE, and platform engineering. He specializes in building secure, scalable cloud-native systems on Kubernetes and AWS, and works at the intersection of platform engineering and AI-driven workloads. Ankit is an active community contributor and regularly speaks at meetups on topics such as developer platforms, policy-as-code, and modern authorization systems. Dastan Oryngali is a Senior Software Developer with over eight years of experience building robust backend systems, scalable microservices, and cloud-based solutions. Throughout his career, he has driven key engineering initiatives across diverse industries, including fintech, consulting, and compliance. Specialized in Java, Spring Boot, and React, Dastan focuses on crafting high-performance architecture and seamless full-stack applications.
Modern operations teams are overwhelmed by dashboards, alerts, logs, vulnerability reports, documentation, and operational data. The challenge is no longer collecting information—it's helping engineers make the right decision quickly. In this session, I'll share the engineering journey behind AskPeter, an AI-powered engineering assistant designed to help production teams investigate incidents, understand system behaviour, analyse vulnerabilities, and retrieve operational knowledge using natural language. Rather than replacing engineers, AskPeter acts as an engineering teammate, combining observability data, operational telemetry, documentation, and AI reasoning to reduce investigation time and improve decision-making. I'll walk through the architecture, design decisions, challenges, security considerations, and lessons learned while introducing AI into real production engineering workflows. This isn't another AI demo—it's an honest engineering story about building AI that engineers can actually trust.
Caner Şahan is a Principal Solution Architect with over 15 years of experience designing enterprise platforms for retail and mission-critical operations. His work focuses on Enterprise Architecture, AI-powered Operations, Observability, DevOps transformation, and Platform Engineering. He is the creator of AskPeter, an AI engineering assistant that combines production telemetry, operational knowledge, and enterprise data to help engineers troubleshoot complex systems more efficiently. Caner is passionate about applying AI to solve real engineering challenges while keeping reliability, security, and operational excellence at the centre of every solution.
You've seen the demo. The agent read the alert, walked the traces, found the misconfigured deployment, and proposed a fix — all in ninety seconds, on stage. Now you have to decide whether it goes anywhere near your production cluster. The published evidence should make you cautious. On IBM's Kubernetes incident benchmark, the best model today scores 56% under a metric that awards zero for a single missed entity. On CUJBench, agents given the full toolset performed worse than agents restricted to browser evidence — 20% versus 28% — and one frontier model collapsed from 52% to 12% accuracy when handed more tools, burning 92 tool calls and 4 million tokens per failed run without ever submitting an answer. Across 1,675 root-cause analysis runs, hallucinated interpretation of data showed up in 71% of them, at every model tier. None of that means agents are useless. It means the demo told you almost nothing, and the questions you'd normally ask a vendor don't cover the failure modes that actually matter here. This talk is six questions, each grounded in what ITBench, AIOpsLab, SREGym and CUJBench actually measured — and each with a way to answer it on your own cluster rather than taking a number on faith. We'll cover why more context can make diagnosis worse, why an agent that retrieves the right evidence still names the wrong service, why a scenario catalog is not a benchmark, and what happens when an agent "resolves" an incident by deleting the fault injector that caused it. You'll leave with the six questions, and a harness design for answering them yourself.
Diego Braga is CTO at Krateo PlatformOps. Diego has spent the last 10 years building open-source architectures for customers, from embedded devices to large-scale distributed systems. Most recently he has been focused on the open cloud infrastructure space, and on emerging patterns for cloud-native applications.
What engineering leaders can learn from Site Reliability Engineering? At first glance, SRE feels like a deeply technical concern: uptime dashboards, infrastructure capacity, on-call rotations. But core SRE concepts are deeper than they seem. Definitions of service criticality, Service-Level Indicators, Service-Level Objectives, and error budgets are not just technical speak. These ideas can be applied far beyond production environments. This talk explores how core SRE concepts can become high-leverage leadership tools for shaping team culture, guiding prioritization, and driving meaningful business outcomes.
Maxim Schepelin is a solution architect at Booking.com and coauthor of Engineering Manager's Compass, a practical guide to the everyday realities of engineering leadership. Over fifteen years, he's built products used by millions of people every day—and picked up a thing or two about exiting both Vim and Emacs along the way.
Varnish is a well-known reverse open source HTTP caching proxy that accelerates websites, applications, files, and any other HTTP-base type of workload. The technology has been around for more than 15 years and powers millions of active websites. While Varnish is considered a stable and reliable piece of acceleration software, there hasn't been a lot of hype surrounding the project. The recent release of Varnish version 9 deserves a bit of hype. In this presentation, Thijs will explain the major new features of Varnish 9, which include TLS support, dynamic backends, OTEL support, and a wide range of exciting modules. He’ll show you how to use them to apply a more intelligent caching layer to your web stack. Thijs will also explain some major changes in the way the open source project is run, and will present some cloud-native additions to the project like Docker images, Helm Charts, and even a Varnish-powered Kubernetes Gateway Controller.
As the Technical Evangelist at Varnish Software, Thijs Feryn focuses on web performance, software scalability, and content delivery. He demonstrates content-driven and technical messaging through presentations, videos, books, blog posts, social media posts, podcasts, and other media. Thijs is a published author and wrote Getting Started with Varnish Cache and Varnish 6 by Example. As a public speaker, he has a track record of over 380 presentations in 26 different countries, where he is often praised for his energetic and engaging presentation style. As an evangelist, Thijs is also active in many open-source communities, most notably the Varnish and PHP community. He has contributed to various communities for over 15 years both technically and as an organizer and facilitator. Prior to joining Varnish Software, Thijs Feryn spent 15 years in the web hosting industry, tackling web performance and scalability issues on a daily basis and evangelizing these topics. For more information about Thijs’ past & upcoming presentations, please visit https://feryn.eu/speaking.
Every vendor now demos an agent that finds the root cause and opens the PR. Almost nobody talks about what that agent is reading — and that's the part that decides whether self-healing is real or a very expensive way to break production faster. Humans tolerate bad telemetry. We squint at an inconsistently named metric, we know which dashboard lies, we carry the tribal context that isn't in any span. An agent has none of that. It acts on exactly what the pipeline says, which makes the data contract — not the model — the binding constraint on autonomy. This talk makes the case that OpenTelemetry is that contract, and gets concrete about it. The GenAI semantic conventions (invoke_agent, execute_tool, gen_ai.*) now let you observe agents as first-class citizens rather than opaque HTTP calls. The same conventions, pointed the other way, are what let an agent reason across your estate without a bespoke integration per tool. Then the honest half: where this breaks today, what the conventions still don't cover, and a staged trust model for moving from suggest → propose → act, with a blast radius you can defend in a post-mortem.
Judith Redi is VP of Engineering at Dash0, an OpenTelemetry-native observability company. She leads the engineering organisation building the platform's ingest, storage, and analysis layers — which means she spends a lot of time looking at what real OTel pipelines do under load, and at what breaks when teams migrate onto them. Before moving into industry engineering leadership, Judith spent a decade in research: a PhD from the University of Genoa on machine learning for image quality assessment, postdoctoral work at Eurecom, and years as an Assistant Professor in TU Delft's Multimedia Computing group, with a Veni award and an IBM Faculty Fellowship along the way. She is a long-standing advocate for gender diversity in engineering.
Platform engineering is entering a new phase: infrastructure is no longer only provisioned through pipelines, templates, and tickets, but increasingly through agentic workflows that can understand intent, reason over platform standards, generate changes, and help operate production systems. In this session, we will explore how agentic approaches can reshape the platform engineering and SRE lifecycle, from infrastructure creation to incident response. Through a live demo, we will first use GitHub Copilot to provision cloud infrastructure by working with repository context, infrastructure-as-code patterns, and platform conventions. Then we will shift into operations and show how Azure SRE Agent can support proactive incident management through detection, diagnosis, response planning, and guided remediation. While the demo uses GitHub Copilot and Azure SRE Agent, the focus is not on a particular service or product. The real discussion is about reusable patterns: how to design agent-ready platforms, how to encode organizational standards as context, how to keep humans in control, and how to introduce safe autonomy into infrastructure and reliability workflows. Attendees will leave with a practical mental model for agentic platform engineering: where agents can help, where guardrails are essential, and how platform teams can evolve from ticket-driven enablement toward intent-driven, context-aware operations.
What I actually do : I draw boxes and lines and say the word "Kubernetes" a lot. Recently i'm prompting too much while letting Github Copilot do all the magic. https://www.dailydoseofghcp.com/ What LinkedIn says I do : After completing a Bachelor’s in Electronics Engineering and a Master’s in Computer Science, I started my career at HPE (then HP) working in pre-sales around x86 server technologies during the Alpha Server era. I later joined Intel, where I evangelized the tick-tock story of Xeon processors across CEE & MEA.Next came AWS, where I had the opportunity to work as AM, SA and TAM with several fast-growing companies including Opsgenie (now Atlassian), Gram Games (now Zynga), Trendyol (now Alibaba) and Yemeksepeti (now Delivery Hero). In 2019 I moved to the Netherlands as a Partner Solutions Architect, helping expand the AWS partner ecosystem across EMEA and later working directly with enterprise customers in the BENELUX region. Most recently, I joined Microsoft as an AI & Cloud Architect, working with Digital Natives and Startups across EMEA.
An AI agent writes good code when it knows what “good” means in your codebase. Usually, it does not. It cannot see the utility functions you already have, understand your architectural conventions, or recognize the patterns your team expects. As a result, it often recreates existing logic, introduces inconsistencies, and produces code that feels disconnected from the rest of the system. And because each session starts with little or no memory, yesterday’s correction is often forgotten today. That is not the model being bad. It is the model working without the context and guidance that a new colleague receives on day one. In this talk, I show how to combine quality gates and engineering practices into a harness around AI-assisted development, so that problems are caught by automated checks instead of by users. We build the harness step by step, examine where each practice helps and where it breaks down, and explore the trade-offs involved. Finally, we look at how to measure whether these techniques actually improve outcomes, using data and evidence rather than impressions.
Ario is a Lead Software Engineer and Tech Innovator with deep expertise in platform engineering, distributed systems, and cloud-native architectures. Specialist in Java, microservices, and middleware (Kafka, Spring, Kubernetes), with a focus on data engineering and AI integration. Proven track record of architecting scalable enterprise systems while leading cross-functional teams with empathy, agility, and continuous mentorship.
In March 2025, Cloudflare's R2 went down for an hour. Production was using an old secret version. The backend had rotated to the new one. Secret-config drift. One hour. Global outage. Because a secret and its config lived in different places, managed by different tools, versioned separately. We stopped treating them as special. Secrets are just config that happens to be encrypted. Same git repo. Same commit. Same deploy. Same rollback. Encrypted with sops-nix, baked into a NixOS image at build time — full OS, not just the app. Deployed via Incus, no SSH. Key pushed separately. Decrypted locally, delivered through systemd credentials. No separate secret server. No separate pipeline. No version mismatch. Rotation doesn't disappear — it gets better. A secret rotation is a git commit, a build, a deploy. You know exactly when it happened, who did it, and which systems have it. The audit trail is git history. No logs to tamper with, no API calls to trace. 30 minutes. How it works, where it hurts, and why this model gets stronger as your system gets more complex.
I run Plumelo, a consultancy working with NixOS, Terraform, and Incus in production. Based in Romania. Big believer in self-hosted and open source. Enjoy playing around with hardware, routers, and networking.
Teams spend a lot of time defining functional and non-functional requirements, writing acceptance criteria, and verifying behavior in controlled environments. That work matters, but it can also lead teams to optimize for test environments while real incidents emerge in production. Production is where systems meet latency, dependency failures, retries, partial outages, and degraded behavior. Most incidents do not come from one obviously broken service. They happen in the gaps between services, where assumptions were never made explicit and failure modes were never explored. In this talk, we give a practical introduction to chaos engineering as a way to close that gap. Not as a dramatic exercise in breaking everything, but as a disciplined way to learn how systems behave under stress. By running small, intentional experiments, teams can test assumptions, uncover weak spots before they become incidents, and better understand how their systems really behave in production. One of the biggest lessons is that reliability often improves before the first experiment even runs. As soon as teams start asking “what if?” questions together, they surface hidden assumptions and identify practical changes worth making immediately. Attendees will leave with a better understanding of why incidents emerge between systems, how chaos engineering connects to a Site Reliability Engineering (SRE) mindset, and how to start with safe, small experiments in their own environment.
Jacob is a leader who inspires others through his commitment to learning, sharing knowledge, and fostering growth. With a strong foundation as a software developer, team lead, and technical consultant, he brings real-world experience to every role he takes on. Jacob actively guides organizations through changes, such as adopting Team Topologies, and supports teams by mentoring developers, creating development teams, and leading workshops. His hands-on approach ensures that his ideas are practical and grounded in experience. Driven by a passion for helping people and teams succeed, Jacob focuses on building strong, collaborative environments. He combines technical expertise with leadership to advocate for better ways of working and continuous improvement.
The fact that your homelab services don't run on several availability zones is no excuse not to have relatively high availability numbers. Your Home Assistant instance or your ad-blocking DNS server can be just as important as any other service. In this talk we're going to provide some practical tips for improving your homelab's reliability for cheap.
Ricard is a Lead Site Reliability Engineer at Cisco ThousandEyes' SRE team. Outside of business hours you can find him working on his homelab, which is what inspired this talk. Ricard is writing a book about homelabbing, so go talk to him if you have a homelab!
After five years of managing serverless databases, I have learned that my rollercoaster journey is very similar to the CPU usage you dream of seeing in the console. This session shares five hard-earned lessons learned while working with so-called serverless databases.
Renato has extensive experience as a cloud architect, tech lead, and cloud services specialist. Currently, he lives in Berlin and works remotely as a principal cloud architect. His primary areas of interest include cloud services and relational databases. He is an editor at InfoQ and a recognized AWS Data Hero.
Most SRE tooling waits for something to break, and Azure SRE Agent changes that. In this talk I'll show how far agentic SRE has come as an AI teammate that doesn't just respond to the 3 AM page, but continuously watches your production environment, reasons over it against best practices and authoritative guidance via MCP, and acts to keep it healthy. We'll go under the hood on what the agent can genuinely do in production today, where the honest limits of autonomy are, and how a governance model built on RBAC and human-in-the-loop keeps you in control.
Albert Tanure is a Senior Cloud Solutions Architect at Microsoft, based in the Netherlands, where he helps global enterprises modernize legacy systems, strengthen cloud resilience, and adopt agentic AI. A Docker Captain and Microsoft MVP, he's the author of ASP.NET Core 9 Essentials (Packt) and Jornada Cloud Native, and an active community voice through talks, mentoring, and writing on cloud-native architecture, DevOps, and AI-assisted engineering.
AI assistants have made writing code faster than ever, but they also introduce major risks—frequently shipping secrets, SQL injections, and flawed Infrastructure as Code (IaC) templates by default. While security tools exist to catch these errors, most require uploading source code to third-party SaaS clouds. This practical talk introduces Lomo CodeSecurity, an open-source, local-first security scanner built to run entirely on localhost. We will explore how to wire a local security gate into your daily development loop by orchestrating eight major industry scanners (including Trivy, Gitleaks, and Checkov) and optionally triaging findings using a local Ollama LLM—with zero cloud dependencies. Attendees will walk away knowing what AI-generated code consistently gets wrong, how to build offline AI security tooling, and see a live demo that goes from zero to a populated findings dashboard in just 90 seconds.
Mladen Maric is a seasoned infrastructure expert with a deep background in large-scale operations. He spent years as a Senior Site Reliability Engineer (SRE) at Zscaler, operating the Zero Trust Exchange at a scale of hundreds of billions of transactions per day. Prior to that, he spent nine years at Juniper Networks building SDN/NFV automation and hybrid clouds for major global telecommunications providers, including Deutsche Telekom and Orange. He is currently building open-source, privacy-first security tools for modern engineering workflows.
AI services can look healthy at the infrastructure layer while users are already experiencing a reliability incident. A GPU may be allocated, Kubernetes may report every pod as ready, and the model endpoint may still suffer from queueing, slow time-to-first-token, tail-latency spikes, noisy-neighbor effects, capacity exhaustion, or degraded output paths. In this talk, I’ll show how to apply SRE thinking to production AI platforms: defining meaningful SLIs and SLOs across the platform, model-serving, and application layers; deciding who owns which failure mode across SRE, platform, and ML teams; and designing shared GPU environments that can degrade gracefully instead of failing unpredictably. I’ll cover practical signals such as queue time, saturation, time-to-first-token, tail latency, successful-inference rate, and cost per successful request, along with lessons from operating Kubernetes/OpenShift AI and multi-tenant GPU workloads. The goal is to give SREs a concrete framework for answering a deceptively simple question: when an AI service is “up,” how do you know it is actually reliable?
Luca Berton is a Production AI speaker and advisor, Docker Captain, former Red Hat engineer, and author of eight technical books. He has 15+ years of enterprise infrastructure experience and speaks internationally on Kubernetes, platform engineering, AI infrastructure, automation, and reliability, including KubeCon + CloudNativeCon Europe 2026 and Red Hat Summit 2026.
This session outlines a strategic shift from fragmented, team-owned deployment scripts to standardized, platform-driven Continuous Delivery (CD). The goal is to eliminate manual bottlenecks and local heroics, replacing them with a secure, scalable, and observable path to production. Core Takeaways The Shift: Delivery must evolve from localized craftsmanship into a standardized, policy-aware platform capability. The Maturity Ladder: Organizations progress through six stages: from basic standardization and policy enforcement to progressive rollouts, release trains, and ultimate zero-touch operations. Ownership Split: Product teams retain ownership of feature intent and domain context, while the central platform owns the repeatable, automated safety guardrails. Policy Centralization: Express delivery policy once within the platform instead of copying manual approval and scan thresholds across fragmented team pipelines. Integrated Security: Supply chain security and artifact governance are embedded directly into the delivery flow to block unsafe packages by default. Why It Matters Now? With the rise of AI-assisted engineering, the volume of code changes is exploding. Relying on human-gated releases creates massive organizational bottlenecks, while automating without platform guardrails exponentially increases production risk.
Anand is a dynamic IT leader specializing in AIOps, Observability, Test Engineering, and Release Management. Currently IT Area Lead at ING Bank and an international digital transformation consultant, he excels at leading global teams and leveraging data-driven insights to boost productivity. Anand holds a B.E. in Computer Science from the Army Institute of Technology, Pune.
Finding performance issues in modern software is like finding a needle in a haystack and intuition on where to look first is often wrong. APerf is an open source tool we have used many times to help with performance debugging by looking “wide” before going “deep”. This session will present the tool along with an performance regression example
Nati is a Solutions Architect with AWS. He delights in helping customers simplify complex systems, teaching them about the inner workings of cloud services and debugging annoying technical oddities. When he is not at his computer he is soldering electronic kits, tinkering with smaller computers and drumming on a Taiko.
Modern IT organizations don’t deliver value through isolated teams—they deliver it through connected flows. This session introduces seven end-to-end value streams that form the backbone of a high-performing digital operating model. From strategy and portfolio, from idea to deployment, from detection to correction, and vulnerability to fix, and more. Learn how connecting these flows reduces handoffs, eliminates bottlenecks, and accelerates business outcomes. Learn a practical framework for making your IT organization truly flow.
Rob Akershoek is a Senior IT Management and DevOps Architect at DXC Technology and Chair of the IT4IT Forum within The Open Group. Recognized as one of the Top 25 Thought Leaders of 2024 by HDI, Rob is a leading voice in modern IT operating models and digital transformation. He helps organizations evolve into lean, agile, and compliant digital service providers capable of thriving in complex hybrid cloud and multi-vendor environments. His expertise spans the design and implementation of end-to-end IT operating models, IT process design and governance, streamlining IT value streams, and implementing modern DevOps toolchains. Rob specializes in aligning disciplines such as Enterprise Architecture (TOGAF), Lean Portfolio Management, (Scaled) Agile and DevOps, Cloud Operations, ITSM, FinOps, and Risk & Compliance. He is also known for architecting integrated DevSecOps toolchains that connect CI/CD, ITSM, observability, and security into cohesive, automated ecosystems. A frequent international speaker and published author, Rob is the lead author of the IT4IT Management Guide and co-author of the IT4IT Standard, contributing significantly to the advancement of IT management practices worldwide. Rob is a regular invited speaker at international conferences, webinars, and community events, where he shares his ideas on modern IT management and digital transformation. He has spoken at events organized by itSMF, DevOps communities, SMFS, SITS, ServiceSpace, and The Open Group.
How we are transforming our monitoring strategy: Moving from legacy, SNMP/Syslog-based monolithic tools to a modern, open-source distributed observability platform focused on end-user experience rather than just device health
Giovanni Pepe is a Sr Engineering Manager at Uber, currently responsible for the Corp Network and SRE teams. Razvan Mihai Cicu is a Senior Infrastructure Engineer with Uber, based in Amsterdam. He specializes in network automation and observability.
Earlier this year, I wrapped up a six-month migration of seven Python data-ingestion Lambdas into services. Most of the code moved unchanged into an environment where concurrent deliveries shared a filesystem, memory space, and process. I’ll use five cases to show what changed in practice. Duplicate deliveries collided in /tmp. Six SQLAlchemy engines created per request eventually filled a 4 GiB process. Lambda’s async invocation retries disappeared when the same work moved behind RabbitMQ. A blocking database call caused liveness probes to fail. A DLQ replay kept taking its pod down. In each case, application code depended on behavior supplied by the runtime. Filesystem isolation, process lifetime, delivery semantics, network topology, and resource boundaries formed the runtime contract around our code. The services had a runtime contract of their own. We had to identify which guarantees carried over and which responsibilities now belonged in the application. I’ll finish with a checklist engineers can use to find those dependencies before moving a workload between managed runtimes.
Denys Zhak is a software engineer with nine years of experience in backend engineering and distributed systems, focused on production reliability. At Imubit, he builds cloud and on-prem software for industrial machine-learning deployments. He led the six-month migration of seven Python ingestion Lambdas into services described in this talk. Outside work, he contributes to open source. Much of that work has been on Astral’s Ruff and ty. He writes about production engineering at denyszhak.com.
Databases for logs are usually forced to pick a side: build heavy inverted indexes for the ingested logs and pay for them on every write, or ingest raw data cheaply and pay with slow queries later. Is there a better approach? Yes - to make brute-force scanning so selective that most data is never read. This approach is taken by VictoriaLogs. This talk follows a log entry through the VictoriaLogs engine: how it is ingested with no upfront schema, how it lands in an immutable LSM-like compressed column-oriented storage, and how queries run the same structure in reverse, applying a stack of pruning mechanisms based on time, stream labels, and bloom filters so that only the data that can actually match is ever read. Beyond one system's internals, this is a talk about a design stance: when you specialize in one kind of data and know exactly how it is laid out and how it will be consumed, you can be smarter about what you build, sometimes discarding indexes entirely, and avoiding reading most of the data at all. You leave with a concrete set of patterns you can apply to your own storage problems, whether the data is logs or something else entirely.
Aliaksandr is a co-founder and the principal architect of VictoriaMetrics. He is also a well-known author of the popular performance-oriented libraries: fasthttp, fastcache and quicktemplate. Prior to VictoriaMetrics, Aliaksandr held CTO and Architect roles with adtech companies serving high volumes of traffic. He holds a Master’s Degree in Computer Software Engineering. He decided to found VictoriaMetrics after experiencing the shortcomings of all available time series databases and observability solutions.
Before your first container hits production on Kubernetes you install a distro, a CNI, an ingress, cert-manager, Prometheus, Grafana, Loki, ArgoCD and Harbor. None of them is your app. 57% of Kubernetes users run more than 11 separate components, the control plane is a pet that needs restoring from backup at 3am, and "why can't A talk to B?" sends you through DNS, Endpoints, kube-proxy, iptables, CNI logs and NetworkPolicy. I have been evaluating, running and being paged for Kubernetes since 1.0 in late 2015, and I kept notes. This year I stopped complaining and built the thing: Reliaburger, a container orchestrator in Rust that ships scheduling, gossip membership, Raft, eBPF service discovery, ingress, mTLS, an image registry, metrics, logs, GitOps and chaos testing in one binary, driven by TOML. An app with health checks, ingress and autoscaling is twelve lines. There is no overlay network, no CNI and no kube-proxy: a name, a virtual IP and about 390 lines of C in the kernel. The talk covers the four scars and what Reliaburger does about each: apps instead of pods, a control plane that is elected rather than installed, SWIM gossip on one slide, and one connection from web to redis step by step. It also covers what it cost to build with Claude and Codex (200 hours and about a thousand pounds for roughly 220k lines of Rust), the rules that kept the models from doubling the codebase every week, the design we got wrong, and what Reliaburger deliberately does not do. Then a live demo, if the demo gods allow. It is 0.1.0, free and Apache 2.0. Download it, break it, tell me.
Miko Pawlikowski has been running platform teams on top of Kubernetes since version 1.0, and has the scars to show for it. He is the author of Chaos Engineering (Manning) and of Reliaburger, a batteries-included container orchestrator written in Rust, along with the book Building Reliaburger that documents how every subsystem was designed and built.









































