SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.
As modern software systems become increasingly complex, Site Reliability Engineering (SRE) teams face mounting operational challenges. Traditional methods—manual triage, rule-based alerts, fragmented dashboards—are no longer sufficient to manage the scale and dynamism of today's distributed architectures. In response, the "AI-driven SRE" paradigm is emerging: agentic systems that emulate expert diagnostic reasoning, integrate service knowledge, and maintain a continuous improvement loop across the incident management lifecycle. Unlike legacy automation or general-purpose AI solutions, these SRE agents are tailored for the investigative, causal-analysis workflows unique to SRE—incorporating telemetry, change history, service topology, and post-incident knowledge. In their most advanced form, they transition from advising to semi-autonomous or fully autonomous operation, triaging incidents, hypothesizing root causes, executing remediation actions, and learning from outcomes in production.
In this talk we will walkthrough recent research and industry initiatives that embody this shift, highlight key architectural patterns (multi-agent coordination, model-context protocols, observability pipelines), and discuss avenues for safe adoption—governance, auditability, and gradual autonomy. In doing so, we argue that AI-driven SRE is poised not merely to augment reliability engineers but to transform how reliability is engineered and operated at scale.... Read more
Management feels messy, but it’s just another complex system - full of incidents, dependencies, and feedback loops. In this talk, we will discuss how to apply engineering principles to leadership: observability, reliability, and iterative improvement for people instead of servers.... Read more
The next evolution of incident response isn’t faster alerts, it’s autonomous resolution. Join ilert CEO Birol Yildiz as he shows how AI SRE agents now diagnose and remediate outages without waking anyone up. Learn how these systems combine observability data, deployment context, and code intelligence to restore services in minutes and hand over clean incident reports instead of 3 a.m. pages.... Read more
At Mollie, as teams grow and new projects are launched, we experienced firsthand how quickly observability costs can start to grow unsustainably. As a response to the limited support from our vendor, we developed a, Vector-powered, telemetry data ingestion solution to help manage ingestion at scale.... Read more
Many teams stop improving OpenTelemetry once basic data is collected, but low-quality instrumentation limits observability. This talk shares practical strategies to enhance signal consistency, completeness, and correlation, plus methods for validating and maintaining quality over time. Learn how to move beyond “done” to build reliable, insightful OpenTelemetry instrumentation.... Read more
As organizations scale their Kubernetes workloads, they often encounter a painful truth: without a solid FinOps strategy, cloud costs can spiral out of control. In this session, we'll explore real-world lessons from scaling Kubernetes environments, and how applying FinOps principles early can avoid firefighting later. You'll learn practical strategies for improving cost visibility, allocating expenses, optimizing resources, and creating a culture of financial accountability. Whether you're just starting with Kubernetes or managing it at scale, this talk will equip you with the tools to align technical growth with business outcomes—without blowing your cloud budget.... Read more
Over the past two years, I’ve been part of building and maintaining a platform where Infrastructure as Code isn’t just a tool — it is the product. Along this journey, we’ve faced a unique set of challenges that go far beyond writing Terraform modules or defining cloud resources.
This talk explores what happens when the abstractions created for scalability and flexibility begin to clash with user experience. From bridging the knowledge gap between the underlying technology and end users, to managing the complexity that comes with more configuration options, I’ll share what worked, what didn’t, and the lessons learned.
We’ll also dive into one of the hardest aspects of maintaining large IaC systems: testing. How do you ensure reliability when every configuration combination can generate a different infrastructure outcome? And how can we build meaningful observability that focuses not only on infrastructure health, but also on giving users actionable insights to optimize performance and control costs?
If you’re building or scaling a product powered by IaC, this talk provides a candid look at the trade-offs, challenges, and strategies to make your platform both powerful and user-friendly.... Read more
AI Agents are non-deterministic, tool-using systems, so typical unit tests and quality monitoring frameworks are not up to the task. In this talk, I will share a practical, end-to-end framework for evaluating our AI agents at PagerDuty, both offline and online.... Read more
**Actionable runbooks** close the gap between *“what to do”* and *“doing it.”*
This talk shows how to design and ship a runbook platform where steps can be clicked and executed safely during incidents.
**What the system does**
- Create and manage runbooks with tags and Markdown.
- Blocks include: instruction, command, API call, conditional, and timer.
- Execute a full runbook or a single block with outputs captured in history.
- Use RBAC, encrypted credential store, versioning, and containerized environments to keep execution safe and repeatable.
- Core entities and API surface: `Runbook`, `RunbookVersion`, `Block`, `ExecutionJob`, `Credentials`, plus endpoints for runbooks, versions, execution, and credentials.
**Architecture at a glance**
- React SPA communicates with a FastAPI backend and MongoDB.
- An execution worker runs jobs and streams results.
**Demo flow**
1. Create a runbook with tags and Markdown instructions.
2. Add a command block and an API call block that uses a stored credential.
3. Assign a custom Docker execution environment to the runbook.
4. Run a single block, then run the entire runbook and watch outputs land in history.
**What you will learn**
- Design principles for truly actionable runbooks and how they differ from static docs.
- How to implement safe execution with RBAC, audit, and container isolation.
- Patterns for versioning and rollbacks so teams can iterate without fear.
- How this approach complements existing incident tooling and industry guidance on making runbooks actionable
([Incident][1], [resources.rundeck.com][2]).
**Who should attend**
- SRE, platform, security, and backend engineers who own on-call and incident response.
- Engineering managers who want safer self-service for ops tasks.
... Read more
Terraform modules are the backbone of modern infrastructures, but, unit testing even in small DevOps teams is a challenge. In this talk I would like to present multiple approaches to address this challenge by deploying ephemeral environments using CI/CD.... Read more
SRE leaders must move beyond automation to lead augmented teams. They must leverage human-AI collaboration and champion explainable AI along with upskilling teams to master this new symbiosis, ensuring AI becomes a reliable partner. The future is human-machine collaboration.... Read more
As an SRE, you rarely get to choose the SLO. But you can influence how the error budget is spent by making smaller, more frequent releases, testing in creative ways and learning to effectively nudge leadership to have a higher change tolerance, you can extend the value of your error budget.... Read more
17:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
Allo | Alcântara Lisbon Offices, Av. da Índia 10,
1300-299 Lisboa, Portugal
Sponsors & Partners
Want to become a sponsor? Get in touch!
Diogo Cebola & Andre Bento
Mollie
Increase observability, not costs: controlling telemetry data ingestion at scale
Abstract
At Mollie, as teams grow and new projects are launched, we experienced firsthand how quickly observability costs can start to grow unsustainably. As a response to the limited support from our vendor, we developed a, Vector-powered, telemetry data ingestion solution to help manage ingestion at scale.
Bio
Hello, my name is Diogo Cebola. I have a background in Computer Engineering, specialising in distributed systems and secure multiparty computation.
My passion for automation and building robust, and simple to deploy applications, led me to the field of platform engineering. I have around 2 years of experience as a platform engineer working in the fields of Reliability and Observability.
I'm currently part of the Reliability team at Mollie. Here I enable teams to build reliable systems and, alongside my colleagues, maintain our observability platform.
My co-speaker André Bento is a Platform Engineer at Mollie, a leading FinTech company, and an invited Professor at the University of Coimbra, where he teaches Distributed Systems, Introduction to Programming, and Systems Integration.
He earned his PhD in Informatics Engineering from the University of Coimbra, with a thesis on Optimizing Availability and Resource Utilization of Cloud Services. André also holds a BSc in Informatics Engineering from the Coimbra Institute of Engineering and an MSc from the University of Coimbra, where his research centered on Observing and Controlling Performance in Microservices. With deep expertise in distributed systems, cloud computing, microservices, observability, and resource optimization, he is passionate about advancing cloud-based solutions and continuously exploring innovative technologies and methodologies.
Yuri Oliveira Sa
OllyGarden
Beyond 'Done': Strategies for Improving OpenTelemetry Instrumentation Quality
Abstract
Many teams stop improving OpenTelemetry once basic data is collected, but low-quality instrumentation limits observability. This talk shares practical strategies to enhance signal consistency, completeness, and correlation, plus methods for validating and maintaining quality over time. Learn how to move beyond “done” to build reliable, insightful OpenTelemetry instrumentation.
Bio
In his professional path, Yuri Sa has always been involved in helping companies achieve the next level of automation in their infrastructure environment. Throughout his 15 years of experience, he worked in critical environments as SysAdmin, SRE, and DevOps Engineer. One of his central beliefs is that all barriers between Development and Operations teams should be removed; for that reason, he decided to contribute to open-source projects focused on observability in the past few years.
Ricardo Miguel Magalhaes
xgeeks
FinOps at Scale: Lessons from Scaling Kubernetes Workloads
Abstract
As organizations scale their Kubernetes workloads, they often encounter a painful truth: without a solid FinOps strategy, cloud costs can spiral out of control. In this session, we'll explore real-world lessons from scaling Kubernetes environments, and how applying FinOps principles early can avoid firefighting later. You'll learn practical strategies for improving cost visibility, allocating expenses, optimizing resources, and creating a culture of financial accountability. Whether you're just starting with Kubernetes or managing it at scale, this talk will equip you with the tools to align technical growth with business outcomes—without blowing your cloud budget.
Bio
Ricardo Magalhães is a Senior DevOps Engineer at xgeeks. He drives the infrastructure for one of the world’s leading automotive companies. Always striving for bigger and better, Ricardo brings an unmatched determination to elevate projects and, of course, build a great story along the way—because everything is better with a good story.
Luis Serra
synvert, a GlogalLogic company
When Infrastructure as Code Becomes a Product: Lessons from the Trenches
Abstract
Over the past two years, I’ve been part of building and maintaining a platform where Infrastructure as Code isn’t just a tool — it is the product. Along this journey, we’ve faced a unique set of challenges that go far beyond writing Terraform modules or defining cloud resources.
This talk explores what happens when the abstractions created for scalability and flexibility begin to clash with user experience. From bridging the knowledge gap between the underlying technology and end users, to managing the complexity that comes with more configuration options, I’ll share what worked, what didn’t, and the lessons learned.
We’ll also dive into one of the hardest aspects of maintaining large IaC systems: testing. How do you ensure reliability when every configuration combination can generate a different infrastructure outcome? And how can we build meaningful observability that focuses not only on infrastructure health, but also on giving users actionable insights to optimize performance and control costs?
If you’re building or scaling a product powered by IaC, this talk provides a candid look at the trade-offs, challenges, and strategies to make your platform both powerful and user-friendly.
Bio
Luis is passionate about the Cloud-Native ecosystem, with Kubernetes playing a central role in all the projects he has worked on. He continuously hones his skills in this area, exploring the CNCF landscape for innovative technologies that solve real-world challenges.
Outside of work, Luis is committed to knowledge sharing. He has published several articles and is an active participant in multiple discussion groups, contributing to the wider DevOps community.
Murilo Venturin & Ricardo Moreira
PagerDuty
Building End-to-End Observability for AI Agents
Abstract
AI Agents are non-deterministic, tool-using systems, so typical unit tests and quality monitoring frameworks are not up to the task. In this talk, I will share a practical, end-to-end framework for evaluating our AI agents at PagerDuty, both offline and online.
Bio
Murilo Venturin is a Machine Learning Engineer specializing in AI agents, LLMs, and generative AI systems. At PagerDuty, he designs and deploys multi-agent AI architectures, integrates reasoning workflows, and builds evaluation frameworks to measure performance of generative AI in production.
Ricardo Moreira is an AI Engineer and Full-Stack Developer specializing in LLMs, deep learning, and scalable AI systems. He is a Senior Applied Scientist at PagerDuty, where he built the company’s unified AI agent testing and evaluation framework.
Rajat Gupta
Paymenttools
Building an Actionable Runbook Platform
Abstract
Actionable runbooks close the gap between “what to do” and “doing it.”
This talk shows how to design and ship a runbook platform where steps can be clicked and executed safely during incidents.
What the system does
- Create and manage runbooks with tags and Markdown.
- Blocks include: instruction, command, API call, conditional, and timer.
- Execute a full runbook or a single block with outputs captured in history.
- Use RBAC, encrypted credential store, versioning, and containerized environments to keep execution safe and repeatable.
- Core entities and API surface: Runbook, RunbookVersion, Block, ExecutionJob, Credentials, plus endpoints for runbooks, versions, execution, and credentials.
Architecture at a glance
- React SPA communicates with a FastAPI backend and MongoDB.
- An execution worker runs jobs and streams results.
Demo flow
1. Create a runbook with tags and Markdown instructions.
2. Add a command block and an API call block that uses a stored credential.
3. Assign a custom Docker execution environment to the runbook.
4. Run a single block, then run the entire runbook and watch outputs land in history.
What you will learn
- Design principles for truly actionable runbooks and how they differ from static docs.
- How to implement safe execution with RBAC, audit, and container isolation.
- Patterns for versioning and rollbacks so teams can iterate without fear.
- How this approach complements existing incident tooling and industry guidance on making runbooks actionable
([Incident][1], [resources.rundeck.com][2]).
Who should attend
- SRE, platform, security, and backend engineers who own on-call and incident response.
- Engineering managers who want safer self-service for ops tasks.
Bio
I’m a Senior Engineering Manager at Paymenttools in Berlin, leading platform teams across SRE and Security. I focus on reliability, observability, Kubernetes, and policy as code, and I drive green-field work from idea to production. Lately I’ve been applying GenAI into the SRE and platform domains. I like clear processes, data-backed decisions, and practical solutions, and I write for peers to share what works and what doesn’t.
Hakoub Esfahani
Veritas Technologies
Modular Terraform, Unit Testing IaC, and CI/CD
Abstract
Terraform modules are the backbone of modern infrastructures, but, unit testing even in small DevOps teams is a challenge. In this talk I would like to present multiple approaches to address this challenge by deploying ephemeral environments using CI/CD.
Bio
Hakoub Esfahani is the DevOps Manager and Product Security lead for Arctera Insight Capture, where he leads CI/CD, Cloud infrastructure, Cloud Security and Product Security efforts.
Over the course of 8 years, Hakoub has closely worked with multiple engineering, security, and product teams on different projects, allowing him to develop a wide array of skills.
He started his career in 2017 by working with clients on technical problems in complex enterprise environments, giving him the opportunity to firsthand understand both the functional and security requirements of the field.
In 2019, he started to shift his focus to DevOps, architecture, and cloud technologies, eventually creating the first DevOps team at the company.
Oscar Manzano
Roche Diagnostics
The Augmented SRE Leader: Steering Human-Machine Synergy in the Age of AIOps
Abstract
SRE leaders must move beyond automation to lead augmented teams. They must leverage human-AI collaboration and champion explainable AI along with upskilling teams to master this new symbiosis, ensuring AI becomes a reliable partner. The future is human-machine collaboration.
Bio
As People Lead in the Infrastructure Reliability Engineering chapter at Roche, I empower my team to deliver key capabilities for innovative digital health solutions. My leadership philosophy centers on mentorship and fostering professional growth.
Beyond people leadership, I have extensive experience in optimizing engineering operations. DevOps is a mindset that I still use and practice. In my previous software engineer experience I developed impactful applications, from mobile apps to assist visually impaired individuals with indoor navigation, decision-support tools for reducing global data center energy consumption, or a full platform designed to accelerate recovery for knee-injured patients.
Along the years in the academia, I authored scientific articles and worked as a software engineer at University College Cork, Ireland, on research projects related to Supply Chain Optimization, Energy Efficiency, Cloud Computing, and e-Health, designing solutions for complex challenges. Finally, I created and directed a Master's program in DevOps, bridging the gap between academia and industry at University of Barcelona, Spain.
Jake Page
MetalBear
Get More Bang for Your Error Budget's Buck: 3 Ways to Extend Reliability Without Changing Your SLOs
Abstract
As an SRE, you rarely get to choose the SLO. But you can influence how the error budget is spent by making smaller, more frequent releases, testing in creative ways and learning to effectively nudge leadership to have a higher change tolerance, you can extend the value of your error budget.
Bio
I'm Jake a DevOps engineer turned DevRel.
Over the last 5 years I have been heavily focused on the world of Cloud Native Dev tooling, from cloud FinOps, Packaging and Software Delivery, not exclusively but many times in the context of Kubernetes clusters.
Having transitioned from a previous career as a high school teacher, any chance I get to speak in front of a crowd on topics that I'm passionate about I try to take. I'm a Lisbon resident and love to frequent the local meetup scene. So if you see me around, don't be a stranger and let's chat.
Joao Freitas
PagerDuty
KeynoteThe Rise of AI-driven SRE
Abstract
As modern software systems become increasingly complex, Site Reliability Engineering (SRE) teams face mounting operational challenges. Traditional methods—manual triage, rule-based alerts, fragmented dashboards—are no longer sufficient to manage the scale and dynamism of today's distributed architectures. In response, the "AI-driven SRE" paradigm is emerging: agentic systems that emulate expert diagnostic reasoning, integrate service knowledge, and maintain a continuous improvement loop across the incident management lifecycle. Unlike legacy automation or general-purpose AI solutions, these SRE agents are tailored for the investigative, causal-analysis workflows unique to SRE—incorporating telemetry, change history, service topology, and post-incident knowledge. In their most advanced form, they transition from advising to semi-autonomous or fully autonomous operation, triaging incidents, hypothesizing root causes, executing remediation actions, and learning from outcomes in production.
In this talk we will walkthrough recent research and industry initiatives that embody this shift, highlight key architectural patterns (multi-agent coordination, model-context protocols, observability pipelines), and discuss avenues for safe adoption—governance, auditability, and gradual autonomy. In doing so, we argue that AI-driven SRE is poised not merely to augment reliability engineers but to transform how reliability is engineered and operated at scale.
Bio
João Freitas is General Manager and Engineering Lead for AI at Pager Duty. João leads PagerDuty AI initiatives and he is the main representative of the Lisbon office, being responsible for its growth, expansion, and culture. With about 20 years of experience in Software Development, Machine Learning, and as a People Manager, he was previously CTO at a startup in the area of Artificial Intelligence and has taken several roles at Microsoft in the areas of Speech Technologies and Data Engineering. João also holds a PhD in the areas of speech technology and human-computer interaction, filed several patents, and published over 40 articles in peer-reviewed international conferences and journals. He is also a regular speaker at AI conferences and the author and co-author of book chapters and one book.
William Mendes
Coralogix
KeynoteManagement is a Hard Job. That’s Why You Should Do It Like an Engineer
Abstract
Management feels messy, but it’s just another complex system - full of incidents, dependencies, and feedback loops. In this talk, we will discuss how to apply engineering principles to leadership: observability, reliability, and iterative improvement for people instead of servers.
Bio
William Mendes is an Engineering Leader with over 15 years of experience designing, scaling, and leading high-performance systems and teams.
Birol Yildiz
ilert
KeynoteWhen Incidents Fix Themselves: AI SRE in action
Abstract
The next evolution of incident response isn’t faster alerts, it’s autonomous resolution. Join ilert CEO Birol Yildiz as he shows how AI SRE agents now diagnose and remediate outages without waking anyone up. Learn how these systems combine observability data, deployment context, and code intelligence to restore services in minutes and hand over clean incident reports instead of 3 a.m. pages.
Bio
Birol Yildiz is the Co-founder and CEO of ilert, adeptly steering the company with a rare combination of technical and product expertise. His prior experience includes a significant role as Chief Product Owner for Big Data products at REWE Digital. With a strong foundation in computer science, Birol bridges the gap between developer and product strategist, constantly striving to innovate and provide customer-centric solutions at ilert.