SREday

Site Reliability, DevOps and Cloud

October 22, 2026 Google for Startups Campus, Warsaw, Poland

1
Day
20+
Speakers
2
Tracks
100+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

Akamai Technologies, Aram Meem, Beewise, Canonical, Cisco, EPAM Systems, Google, Hyland, IN Groupe, Inter Cars S.A., LTIMindtree, Netflix, Novo Nordisk, Quesma, Replika, VictoriaMetrics

Topics so far:

Event Starts In:

Tickets

Schedule

October 22, 2026 • 2 parallel tracks • 9AM - 4:30PM • Warsaw, in-person
view as table
main room • Track 1

09:00

Riccardo Carlesso

Stop Grepping, Start Reasoning: Skill-Based Agentic SRE 🐉

Google
Let’s be honest: debugging production at scale is soul-crushing manual labor. At Google, we decided making thousands of SREs act like human search engines in 2026 was a bug, not a feature. Enter the Agentic SRE Extension — a skill-based framework that lives in your terminal and actually knows its way around Kubernetes. In this high-energy deep dive, I’m showing you exactly how we use modular skills and MCP to chain diagnostic tools, analyze metric regressions, and execute safe mitigations (Rollbacks, Throttling) without the "deleted production" anxiety. We’ll dissect the "Outage Investigator" agent's logic loop, see it draft a technical postmortem in seconds, and discuss why we built this as a portable framework whose skills can be easily leveraged by any modern AI harness. You’ll leave with the code to wire up your own Kubernetes stack and let the agent's skills do the heavy lifting while you drink espresso. No fluff, no 101s. Just AI agents, specialized skills, and less toil.... Read more

09:30

Coffee break

Main lobby

10:00

Joshua Borkowski-Clark

Incident Management at Google

Google
This presentation provides a high level overview of Incident Management, the practice of responding to an incident in a structured way. It talks about the general incident lifecycle, the Incident Command System framework, and some general how-to tips on applying it. The talk is aimed at those who are typically on call, or are responsible for resolving a incident when things go wrong. By the end of the presentation, the audience should have a practical understanding of how to manage a incident.... Read more

10:30

Yernar Kenzhetayev

On-call health: The Human Side of Reliability

Netflix
Every Engineer knows the feeling: your phone pages at 3am, your heart rate spikes before you've even read the alert, and you're debugging a production issue half-asleep. We treat this as normal. It shouldn't be. This talk is about on-call health. I'll talk about the different shapes on-call can take and why the structure itself can make burnout better or worse. I'll get into the psychological toll, the anxiety of carrying a pager, the burnout that builds quietly over months, and what actually helps: real training, runbooks people trust, and practice before the real incident hits.... Read more

11:00

Aliaksandr Valialkin

How to Create a Database Optimized for Petabytes of Logs and Wide Events

VictoriaMetrics
Databases for logs are usually forced to pick a side: build heavy inverted indexes for the ingested logs and pay for them on every write, or ingest raw data cheaply and pay with slow queries later. Is there a better approach? Yes - to make brute-force scanning so selective that most data is never read. This approach is taken by VictoriaLogs. This talk follows a log entry through the VictoriaLogs engine: how it is ingested with no upfront schema, how it lands in an immutable LSM-like compressed column-oriented storage, and how queries run the same structure in reverse, applying a stack of pruning mechanisms based on time, stream labels, and bloom filters so that only the data that can actually match is ever read. Beyond one system's internals, this is a talk about a design stance: when you specialize in one kind of data and know exactly how it is laid out and how it will be consumed, you can be smarter about what you build, sometimes discarding indexes entirely, and avoiding reading most of the data at all. You leave with a concrete set of patterns you can apply to your own storage problems, whether the data is logs or something else entirely.... Read more

11:30

Kuba Rachwalski

Observability Isn't a Tool Problem. It's an Org Problem.

Cisco
Most organizations invest in an observability solution and assume the results will follow. Six months later the tool is deployed, but alert noise is worse than before, nobody outside the platform team logs in, and the number of incidents hasn't changed. The missing piece is never the technology: it's the operating model around it. I lead the team that has helped several hundred large enterprises across EMEA make their observability actually work, and I'll share what the successful implementations did differently: how they structured the team, how they standardized, and how they proved the value to the business. You'll leave with a checklist of what to get right first.... Read more

12:00

Lunch & networking

Main lobby

13:00

Przemyslaw Hejman

The Agent Did What!? A Flight Recorder for AI Agents

Quesma
Your AI agent changed the code, ran a few commands, and declared victory. Can you reconstruct what happened? Keeping agent trajectories gives you a record of the prompts, tool calls, and outputs behind the result. Analysing those records can help you understand failed tasks, find recurring mistakes, improve how your team uses agents, and investigate suspected misuse. If you only start collecting after something goes wrong, the evidence may already be gone. This talk will explore what you can learn from stored trajectories and why they’re worth keeping even before you know which questions you’ll need to ask. I’ll also share lessons from building Quesma Shipper, our open-source trajectory collector: dealing with scattered session data, changing formats, and transcripts full of secrets. We’ll discuss how to preserve useful records without creating an unnecessary store of sensitive data, and where trajectories alone fall short as evidence.... Read more

13:30

Ali Ogun

Mission (A)Impossible: AWS DevOps Agent Meets AFT

Hyland
WS Control Tower Account Factory for Terraform (AFT) brings the flexibility of Terraform and the GitOps to governed AWS account provisioning. But AFT is a black box and there's no single dashboard to explain pipeline failures. In this talk, we follow a failure back to its source and ask the real question of the new SRE era: could AWS DevOps Agent actually investigate and trace it home? The answer reveals something important about AI operations.... Read more

14:00

Hanna Novikava

Your Observability Bill Is a Design Problem, Not a Vendor Problem

SRE/DevOps Engineer
Observability costs tend to grow faster than traffic, and the usual response, renegotiating with the vendor or cutting retention, treats the symptom rather than the cause. The real driver is almost always a small set of design decisions: unbounded label cardinality, logs carrying work that belongs to metrics, and instrumentation added by default rather than by intent. This talk breaks down where the money actually goes in a modern telemetry pipeline, from the moment a signal is emitted to the moment someone queries it, and shows which changes cut cost without cutting visibility. Drawing on patterns common across platform teams, it covers which changes pay off, which turn out not to be worth the effort, and which trade-offs are worth accepting deliberately. It also covers the point at which sampling stops being safe.... Read more

14:30

Piotr Wojcikowski

"Response Code: Burnout" — How to Survive On-Call in SRE

Inter Cars S.A.
At Inter Cars, our SRE team keeps e-Catalog running — a B2B e-commerce platform serving customers across 20+ European markets. Engineers work on-call shifts 24/7/365, picking up alerts and responding to malfunctions whenever they hit, day or night. This talk is a practical, honest look at what on-call actually feels like: what causes the frustration and stress, and what — if anything — actually helps. Drawing on real experience running a critical B2B platform, we'll dig into whether burnout can really be managed at all, or whether we just get better at living with it.... Read more

15:00

Networking & sponsor crawl

Main lobby

15:30

Mateusz Kulewicz

Observing Coding Agents

Canonical
A significant shift in the software engineering industry is happening right now. More and more people are using agents in their workflows to write code - either under close supervision or completely autonomously. Pull requests are written, reviewed, and sometimes even merged without any human involvement. New tools and models are released every week, and there's lots of interest in completely autonomous agents such as OpenClaw and Hermes. As an observability engineer, my main focus is on making sure I know what's going on in my systems. How to observe coding agents? Do traditional observability principles and techniques still matter in this new stochastic paradigm? In this talk, I will show you how to keep track what your agents are doing and some of the practices that are forming in the industry as we speak - and whether you can use the tools you already have. If you know the basic observability terms, you are the right person to be in the audience.... Read more

16:00

Ilia Matytcin

SRE at Scale: How We Standardized Reliability Across 20+ Teams

EPAM Systems
Scaling SRE practices from a single team to an entire engineering organization of 20+ teams is rarely just a technical challenge — it is a cultural one. In this talk, I will share a real-world case study of how we rolled out standardized SLOs, unified observability dashboards, and actionable alerting across dozens of autonomous delivery teams without creating a central bottleneck. I will cover what worked, what didn't, and the hard-earned lessons from fighting alert fatigue, defining meaningful SLOs, and getting teams to truly own the reliability of their services. Attendees will leave with a practical playbook for driving a reliability-first culture at scale — and a list of pitfalls to avoid along the way.... Read more

16:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
secondary room • Track 2

09:30

Coffee break

Main lobby

10:00

​Kirill Solovei

Centralized cross-account EKS observability

Replika
Complex systems require extensive monitoring and observability. Systems as complex as Kubernetes clusters have so many moving parts that sometimes it's a task and a half just to configure their monitoring properly. This talk is a deep dive into cross-account observability for multiple EKS clusters, exploring various implementation options, outline the pros and cons of each approach, and explanation of one of them in close detail. Whether you're an aspiring engineer seeking best-practice advice, a seasoned professional ready to disagree with everything, or a manager looking for ways to optimize costs -- this talk might be just right for you.... Read more

10:30

Mary Adeseluka

Beyond Alerts: Building an Observability Strategy That Drives SRE Action

LTIMindtree
Today's SRE teams have a lot of data coming in, but just having more information doesn't always mean they can improve system reliability. When observability is all about dashboards, alerts, and increasing numbers of signals, engineers may find themselves spending too much time dealing with noise rather than making systems better. This discussion explores how to move beyond alerts and build an observability strategy designed around action. We'll explore how to link metrics, logs, traces, events, and business context to the real decisions SREs need to make like spotting new issues, understanding how they affect things, speeding up how we handle incidents, and stopping them from happening again. The aim is not to gather more data, but to make observability a real part of how teams work. This helps them find issues quicker, learn more about what's happening, make decisions with confidence, and keep improving how reliable their systems are.... Read more

11:00

Michal Gawron

Beyond the Golden Image: Continuous Compliance for a Hybrid Server Fleet

Novo Nordisk
A golden image gives you a compliant server on the day it is deployed. It does not keep it compliant for the next three years - patches, agent versions, manual changes and ownership all drift on independent clocks, and at fleet scale that stops being a technical problem and becomes an operating-model problem. This talk presents a lifecycle architecture that treats image build, onboarding, desired-state assessment, controlled remediation and centralised patching as a single control loop across hybrid Windows and Linux estates, and shows where the boundary between each layer belongs. I will cover why assessment and remediation need different risk models, why patching is a separate control plane from configuration, and why a machine that has stopped reporting is a bigger problem than a machine reporting non-compliance.... Read more

11:30

Marek Smigielski

Is Istio Service Mesh Still Worth the Money in the AI Era?

IN Groupe
Do you have a **service mesh** running in your cluster, or are you considering **Istio** and wondering whether the benefits justify another layer of complexity? As **AI** speeds up prototyping and application changes, **SREs** have more to do than ever. Let's examine where a **service mesh helps**, and when it simply gives you **more infrastructure to maintain**. We’ll look **under the hood**, then explore how its **observability capabilities** help you understand **service dependencies** and **troubleshoot traffic**, including calls leaving your cluster. We’ll also examine where **Istio fits into rapid experimentation**: **A/B testing**, **blue-green deployments**, **shadow traffic**, and recovering when a new component misbehaves. If you wonder **what a service mesh is**, or if you are simply **lost with your cluster traffic**, this session is for you!... Read more

12:00

Lunch & networking

Main lobby

13:00

Hubert Poznanski

Faster, Smarter, Fragile: The human cost of AI Boom

Senior DevOps Engineer
AI has made engineering teams genuinely faster - but the bill is coming, and it will be paid in more than money. This talk looks at what happens when AI budgets get cut in half: workflows with AI stitched into their spine, engineers who have never debugged without an assistant, and codebases full of generated code that no one truly owns. Skill atrophy is technical debt stored in people, and it has no refactoring. Finally, we'll cover how platform teams can build AI as a detachable dependency rather than a foundation - including one concrete practice you can take back to your team.... Read more

13:30

Artur Polek

The Black Box Strategy: Building SRE Infrastructure for Unknown Applications

Akamai Technologies
When I joined a team of 8 developers, their Linode Kubernetes Engine infrastructure was built through UI clicks, manual helm deployments, and YAML manifests scattered across environments - and I didn't understand what their application actually did.... Read more

14:00

Oleksandr Zhyhalo

Observability for IoT Fleets Without Going Bankrupt

Beewise
Running observability for devices in the field is a different problem from running it for services on your PaaS. Power, bandwidth and update cadence are fixed by physics, and each one rules out tooling that would be the obvious pick anywhere else. This talk works through the decisions an IoT fleet forces on you: what to collect, where to process it when every byte is metered, how to index without labels, and why storage architecture—not query speed—determines your long-term costs.... Read more

14:30

Alex Dainiak

Incremental Refactoring of Legacy Systems Under Load

Aram Meem
Replacing or refactoring a legacy system while it handles live high-throughput traffic is akin to replacing an airplane engine in mid-flight. In high-load environments where even minutes of downtime result in direct revenue loss, full rewrites are often too risky, making incremental refactoring the only viable path. I will provide real-world engineering strategies for safely modernizing legacy monolithic services under active load without breaching SLAs or risking data integrity. I am going to cover practical execution patterns: isolating legacy boundaries, leveraging the Strangler Fig pattern, deploying feature flags, and executing shadow reads and dual-writes for data storage migrations. Special emphasis is placed on observability and SRE guardrails — how to structure telemetry, automated circuit breakers, and canary deployments to catch performance regressions instantly during transition phases. The audience will gain hard-earned insights into balancing architectural evolution, system stability, and continuous delivery.... Read more

15:00

Networking & sponsor crawl

Main lobby

15:30

Felix Masomera

When Reliability Becomes the Problem: The SRE Traps Nobody Talks About

Site Reliability Engineer
SRE is about making systems reliable, but our efforts to improve reliability can sometimes introduce complexity of their own. This talk explores common SRE traps, including alert fatigue, over-automation, excessive tooling, overengineering, and reliance on “hero” engineers. Through practical examples and real-world lessons, we’ll examine how these challenges can affect the reliability and operability of our systems and how to build systems that are not only reliable, but also maintainable, understandable, and easier to operate.... Read more

16:00

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room secondary room
09:00 Stop Grepping, Start Reasoning: Skill-Based Agentic SRE 🐉
Riccardo Carlesso • Google
09:30 Coffee break
10:00 Incident Management at Google
Joshua Borkowski-Clark • Google
Centralized cross-account EKS observability
​Kirill Solovei • Replika
10:30 On-call health: The Human Side of Reliability
Yernar Kenzhetayev • Netflix
Beyond Alerts: Building an Observability Strategy That Drives SRE Action
Mary Adeseluka • LTIMindtree
11:00 How to Create a Database Optimized for Petabytes of Logs and Wide Events
Aliaksandr Valialkin • VictoriaMetrics
Beyond the Golden Image: Continuous Compliance for a Hybrid Server Fleet
Michal Gawron • Novo Nordisk
11:30 Observability Isn't a Tool Problem. It's an Org Problem.
Kuba Rachwalski • Cisco
Is Istio Service Mesh Still Worth the Money in the AI Era?
Marek Smigielski • IN Groupe
12:00 Lunch & networking
13:00 The Agent Did What!? A Flight Recorder for AI Agents
Przemyslaw Hejman • Quesma
Faster, Smarter, Fragile: The human cost of AI Boom
Hubert Poznanski • Senior DevOps Engineer
13:30 Mission (A)Impossible: AWS DevOps Agent Meets AFT
Ali Ogun • Hyland
The Black Box Strategy: Building SRE Infrastructure for Unknown Applications
Artur Polek • Akamai Technologies
14:00 Your Observability Bill Is a Design Problem, Not a Vendor Problem
Hanna Novikava • SRE/DevOps Engineer
Observability for IoT Fleets Without Going Bankrupt
Oleksandr Zhyhalo • Beewise
14:30 "Response Code: Burnout" — How to Survive On-Call in SRE
Piotr Wojcikowski • Inter Cars S.A.
Incremental Refactoring of Legacy Systems Under Load
Alex Dainiak • Aram Meem
15:00 Networking & sponsor crawl
15:30 Observing Coding Agents
Mateusz Kulewicz • Canonical
When Reliability Becomes the Problem: The SRE Traps Nobody Talks About
Felix Masomera • Site Reliability Engineer
16:00 SRE at Scale: How We Standardized Reliability Across 20+ Teams
Ilia Matytcin • EPAM Systems
Wrap up
16:30 Wrap up

Speakers

Alex Dainiak
Aram Meem
Ali Ogun
Hyland
Aliaksandr Valialkin
VictoriaMetrics
Artur Polek
Akamai Technologies
Felix Masomera
Site Reliability Engineer
Hanna Novikava
SRE/DevOps Engineer
Hubert Poznanski
Senior DevOps Engineer
Ilia Matytcin
EPAM Systems
Joshua Borkowski-Clark
Google
Kuba Rachwalski
Cisco
Marek Smigielski
IN Groupe
Mary Adeseluka
LTIMindtree
Mateusz Kulewicz
Canonical
Michal Gawron
Novo Nordisk
Oleksandr Zhyhalo
Beewise
Piotr Wojcikowski
Inter Cars S.A.
Przemyslaw Hejman
Quesma
Riccardo Carlesso
Google
Yernar Kenzhetayev
Netflix
​Kirill Solovei
Replika

Venue

Google for Startups Campus - Warsaw

Centrum Praskie Koneser, Plac Konesera 10
03-736 Warsaw, Poland

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one