SREday

Site Reliability, DevOps and Cloud

June 12, 2025 Cologne, Germany

1
Day
12+
Speakers
1
Track
50+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

Accenture, Electrolux Group, Funambol, Grafana Labs, huettermann.net, ilert, ING Hubs Romania, PagerDuty, Pegasystems, VMware / Pivotal

Topics so far:

This is a past event, what's next?

Schedule

June 12, 2025 single track 10AM - 5PM Cologne, in-person
view as table
main room • Track 1

10:00

Birol Yildiz

KeynoteAI Agents for Incident Management

ilert
AI agents are transforming incident management by automating detection, diagnosis, and response. This keynote explores how intelligent systems can reduce downtime, enhance decision-making, and streamline operations, offering a glimpse into the future of autonomous incident resolution.... Read more

10:30

Coffee break

Main lobby

11:00

Kat Gaines

Staying in your lane: incident response for leaders

PagerDuty
During incident response, every second counts. As a leader, you're expected to have all the answers, and that pressure can make it very tempting to play the hero and take control of the response process. While "I'll just do it myself" may seem faster on occasion, this instinct can hinder your team's effectiveness and actually slow down resolution. We'll explore how you can evolve from a command-and-control manager into a genuinely empowering leader. You'll learn how to maintain composure under pressure while creating space for your team to shine. We'll cover essential frameworks, usable by new managers and individual contributors building incident response processes from scratch as well as experienced leaders looking to level up their approach. Key takeaways will include: - How to recognize when you may be over-managing during an incident… and how to empower your team to help you recognize these moments. - Building team confidence and ownership in crisis situations. - Clear guidelines for when to step in versus step back. - Establishing and maintaining clear communication channels. Whether you're establishing your first incident response playbook, fine-tuning an existing process, or managing up to help your leaders help you do your best work, you'll leave equipped to help your team handle incidents more effectively – and help you sleep better at night knowing they've got this.... Read more

11:30

Michael Cote

Platform Engineering for Private Cloud

VMware / Pivotal
_Platform engineering_ is the art of building and managing the infrastructure that powers your applications: a mix of cloud, a handful of DevOps, a pinch of SRE, and a thick glaze of product management. While it’s “nothing new,” many organizations are just starting to practice it—and for good reason. But what happens when your platform is running on private cloud? With around 50% of enterprise apps still running on private clouds, platform engineering for private platforms is surprisingly under-discussed. This talk dives into real-world examples and stories from large organizations tackling this challenge. The hurdles often lie in adapting platform engineering to existing IT stacks and processes—most organizations can’t simply start from scratch, nor would they want to abandon what’s currently driving revenue. If you support your organization’s apps, platform engineering is something you’ll probably be doing soon. Come learn how your peers are navigating these challenges and share your own experiences.... Read more

12:00

Dr. Michael Hüttermann

DevOps XXL: How to Scale-up in Large Setups with SRE?

huettermann.net
This session derives Dos and Donts from very large setups. It aims to draw a thin line between DevOps and SRE, defines those two concepts and offers practical guidance how to leverage both. We discuss how to bring together the communities of DevOps and SRE. This interactive session is based on both, good practices from the field and insights from academic research.... Read more

12:30

Alina Astapovich

Automating SRE Operations with Multi-Agent AI: InfraAssistant Approach

Electrolux Group
SRE teams often face challenges with a high volume of routine tasks and requests, making it difficult to focus on critical, high-priority issues. At Electrolux, we faced the same challenge, which led us to develop InfraAssistant —a multi-agent AI-powered solution designed to automate key operational tasks such as infrastructure management, user onboarding, and responding to internal requests. This shift reduced our manual workload and significantly improved operational efficiency. InfraAssistant is built on specialized agents that coordinate to autonomously manage complex tasks, reducing the need for continuous manual involvement. This session will cover the design and orchestration of these agents, showcasing how InfraAssistant helps SRE teams by automating day-to-day operations, minimizing repetitive tasks, and enhancing the management of complex infrastructure.... Read more

13:00

Lunch & networking

Main lobby

14:00

Andrada Raducanu

From Code to Cluster: Orchestrating 10,000+ Kubernetes deployments with 1 pipeline

ING Hubs Romania
There is a sea of tools one can use for the critical phase of Deployment during your SDLC. To keep our environment secure and reliable, ING chose to work with Kubernetes and Azure DevOps. In this talk, we will share the success story of how 1200 in-house developed APIs reached 10 000+ Production deployments in half a year, using one single pipeline. In order to stay in control, we use Open Policy Agent. To ensure the reliability and the resilience of the APIs, we use tools like: QuotaAutoscaler (ING open source CRD) and HorizontalPodAustoscaler, native rollback mechanisms with Helm, automatic certificates using CertManager and Prometheus monitoring. The pipeline deploys code in Azure Kubernetes Service and on-prem Kubernetes clusters. This solution was built as a platform, designed to be agnostic to the target system, reducing the cognitive load on the teams and allowing them to focus on the application development. We call this The Kingsroad.... Read more

14:30

Agnieszka Welian

How to tame chaos effectively?

Pegasystems
Imagine a self-healing system that handles surprises, letting you sleep peacefully. If that sounds appealing, chaos engineering could be the answer. Trusted by Netflix, LinkedIn, Google, and Facebook, it's key for business resilience. In this session, we'll explore its history, learn how to apply its principles to stress-test applications, and review tools for fault injection in real-world scenarios.... Read more

15:00

Bohdan Pohorilets

Cloud infrastructure with AWS Cloud Development Kit

Freelance SRE
This hands-on workshop guides engineers in developing cloud infrastructure using AWS CDK to make a web application public. Participants will learn how to spin up, deploy, and tear down cloud resources with a few console commands. Key technologies include AWS (CDK, S3, Lambda, RDS) and Ruby on Rails with PostgreSQL. Attendees need AWS and GitHub accounts, CLI tools, and specific runtime environments. The session balances interactive deployments with insights into cloud architecture and best practices. *[The Github Repo is available here](https://github.com/bpohoriletz/prototype)*... Read more

15:30

Renato Losio

Things Fall Apart: Navigating Managed Databases for Over a Decade as a Non-DBA

Funambol
Learn to navigate database benchmarks wisely! From crashing managed instances to skyrocketing storage costs, I'll share hard-earned lessons from a decade of managing production databases on the cloud without DBA expertise.... Read more

16:00

Michele Dodic

From Firefighting to Zero-Touch: How AIOps Empowers SREs to Build Resilient, Business-Driven Platforms

Accenture
Explore how AIOps, combined with Observability and Chaos Engineering, is empowering SREs to transform IT operations into a strategic driver of business value. This session delves into strategies for achieving end-to-end visibility and resilience—from infrastructure and applications to business processes. Learn how advanced analytics, AI and automation can optimize the user journey, enhance customer experiences, and drive measurable business impact. Through real-world case studies and a zero-touch AIOps demo powered by GenAI, gain actionable insights to predict failures, address vulnerabilities, and ensure seamless system performance across the entire operational stack.... Read more

16:30

Syed Usman Ahmad

From Text to Visuals: Observe your GitHub repos efficiently

Grafana Labs
Whether you are just starting your journey as a DevOps Engineer or already an expert SRE, you need to use Git (e.g. GitHub, GtlLab) to maintain your code and track issues, PRs, commits, etc. You start with a few repositories but then it becomes difficult to get visibility into what's going on in a complex app when it's scattered across 20 different GitHub repos. In this talk, we will demonstrate an example of how to monitor your GitHub repo using the GitHub Data source plugin that allows you to represent data e.g. open/close issues, Pull n Merge requests, and other various statuses visually in Grafana dashboards and gives a dynamic and interactive view to monitoring for a more adaptable and user-centric dashboard experience either for your personal projects or for the entire team. Later, we see more advanced features to get better observability. It will be an introduction to the Grafana Dashboards, and also an excellent opportunity to learn more about utilizing advanced visualization options to effectively represent complex data. Join us to learn more about Grafana dashboards, community contributions and share your feedback and suggestions!... Read more

17:00

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room
10:00 Keynote: AI Agents for Incident Management
Birol Yildiz • ilert
10:30 Coffee break
11:00 Staying in your lane: incident response for leaders
Kat Gaines • PagerDuty
11:30 Platform Engineering for Private Cloud
Michael Cote • VMware / Pivotal
12:00 DevOps XXL: How to Scale-up in Large Setups with SRE?
Dr. Michael Hüttermann • huettermann.net
12:30 Automating SRE Operations with Multi-Agent AI: InfraAssistant Approach
Alina Astapovich • Electrolux Group
13:00 Lunch & networking
14:00 From Code to Cluster: Orchestrating 10,000+ Kubernetes deployments with 1 pipeline
Andrada Raducanu • ING Hubs Romania
14:30 How to tame chaos effectively?
Agnieszka Welian • Pegasystems
15:00 Cloud infrastructure with AWS Cloud Development Kit
Bohdan Pohorilets • Freelance SRE
15:30 Things Fall Apart: Navigating Managed Databases for Over a Decade as a Non-DBA
Renato Losio • Funambol
16:00 From Firefighting to Zero-Touch: How AIOps Empowers SREs to Build Resilient, Business-Driven Platforms
Michele Dodic • Accenture
16:30 From Text to Visuals: Observe your GitHub repos efficiently
Syed Usman Ahmad • Grafana Labs
17:00 Wrap up

Speakers

Agnieszka Welian
Pegasystems
Alina Astapovich
Electrolux Group
Andrada Raducanu
ING Hubs Romania
Birol Yildiz
ilert
Bohdan Pohorilets
Freelance SRE
Dr. Michael Hüttermann
huettermann.net
Kat Gaines
PagerDuty
Michael Cote
VMware / Pivotal
Michele Dodic
Accenture
Renato Losio
Funambol
Syed Usman Ahmad
Grafana Labs

Venue

The offices of Ilert.com

Bayenstraße 65,
50678 Köln, Germany

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one