SREday

Site Reliability, DevOps and Cloud

November 27, 2025 PagerDuty, Lisbon, Portugal

1
Day
10+
Speakers
1
Track
50+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

Coralogix, ilert, MetalBear, Mollie, OllyGarden, PagerDuty, Paymenttools, Roche Diagnostics, synvert, a GlogalLogic company, Veritas Technologies, xgeeks

Topics so far:

This is a past event, what's next?

Schedule

November 27, 2025 single track 9:30AM - 5:30PM Lisbon, in-person
view as table
main room • Track 1

09:30

Joao Freitas

KeynoteThe Rise of AI-driven SRE

PagerDuty
As modern software systems become increasingly complex, Site Reliability Engineering (SRE) teams face mounting operational challenges. Traditional methods—manual triage, rule-based alerts, fragmented dashboards—are no longer sufficient to manage the scale and dynamism of today's distributed architectures. In response, the "AI-driven SRE" paradigm is emerging: agentic systems that emulate expert diagnostic reasoning, integrate service knowledge, and maintain a continuous improvement loop across the incident management lifecycle. Unlike legacy automation or general-purpose AI solutions, these SRE agents are tailored for the investigative, causal-analysis workflows unique to SRE—incorporating telemetry, change history, service topology, and post-incident knowledge. In their most advanced form, they transition from advising to semi-autonomous or fully autonomous operation, triaging incidents, hypothesizing root causes, executing remediation actions, and learning from outcomes in production. In this talk we will walkthrough recent research and industry initiatives that embody this shift, highlight key architectural patterns (multi-agent coordination, model-context protocols, observability pipelines), and discuss avenues for safe adoption—governance, auditability, and gradual autonomy. In doing so, we argue that AI-driven SRE is poised not merely to augment reliability engineers but to transform how reliability is engineered and operated at scale.... Read more

10:00

William Mendes

KeynoteManagement is a Hard Job. That’s Why You Should Do It Like an Engineer

Coralogix
Management feels messy, but it’s just another complex system - full of incidents, dependencies, and feedback loops. In this talk, we will discuss how to apply engineering principles to leadership: observability, reliability, and iterative improvement for people instead of servers.... Read more

10:30

Birol Yildiz

KeynoteWhen Incidents Fix Themselves: AI SRE in action

ilert
The next evolution of incident response isn’t faster alerts, it’s autonomous resolution. Join ilert CEO Birol Yildiz as he shows how AI SRE agents now diagnose and remediate outages without waking anyone up. Learn how these systems combine observability data, deployment context, and code intelligence to restore services in minutes and hand over clean incident reports instead of 3 a.m. pages.... Read more

11:00

Coffee break

Main lobby

11:30

Diogo Cebola & Andre Bento

Increase observability, not costs: controlling telemetry data ingestion at scale

Mollie
At Mollie, as teams grow and new projects are launched, we experienced firsthand how quickly observability costs can start to grow unsustainably. As a response to the limited support from our vendor, we developed a, Vector-powered, telemetry data ingestion solution to help manage ingestion at scale.... Read more

12:00

Yuri Oliveira Sa

Beyond 'Done': Strategies for Improving OpenTelemetry Instrumentation Quality

OllyGarden
Many teams stop improving OpenTelemetry once basic data is collected, but low-quality instrumentation limits observability. This talk shares practical strategies to enhance signal consistency, completeness, and correlation, plus methods for validating and maintaining quality over time. Learn how to move beyond “done” to build reliable, insightful OpenTelemetry instrumentation.... Read more

12:30

Ricardo Miguel Magalhaes

FinOps at Scale: Lessons from Scaling Kubernetes Workloads

xgeeks
As organizations scale their Kubernetes workloads, they often encounter a painful truth: without a solid FinOps strategy, cloud costs can spiral out of control. In this session, we'll explore real-world lessons from scaling Kubernetes environments, and how applying FinOps principles early can avoid firefighting later. You'll learn practical strategies for improving cost visibility, allocating expenses, optimizing resources, and creating a culture of financial accountability. Whether you're just starting with Kubernetes or managing it at scale, this talk will equip you with the tools to align technical growth with business outcomes—without blowing your cloud budget.... Read more

13:00

Luis Serra

When Infrastructure as Code Becomes a Product: Lessons from the Trenches

synvert, a GlogalLogic company
Over the past two years, I’ve been part of building and maintaining a platform where Infrastructure as Code isn’t just a tool — it is the product. Along this journey, we’ve faced a unique set of challenges that go far beyond writing Terraform modules or defining cloud resources. This talk explores what happens when the abstractions created for scalability and flexibility begin to clash with user experience. From bridging the knowledge gap between the underlying technology and end users, to managing the complexity that comes with more configuration options, I’ll share what worked, what didn’t, and the lessons learned. We’ll also dive into one of the hardest aspects of maintaining large IaC systems: testing. How do you ensure reliability when every configuration combination can generate a different infrastructure outcome? And how can we build meaningful observability that focuses not only on infrastructure health, but also on giving users actionable insights to optimize performance and control costs? If you’re building or scaling a product powered by IaC, this talk provides a candid look at the trade-offs, challenges, and strategies to make your platform both powerful and user-friendly.... Read more

13:30

Lunch & networking

Main lobby

14:30

Murilo Venturin & Ricardo Moreira

Building End-to-End Observability for AI Agents

PagerDuty
AI Agents are non-deterministic, tool-using systems, so typical unit tests and quality monitoring frameworks are not up to the task. In this talk, I will share a practical, end-to-end framework for evaluating our AI agents at PagerDuty, both offline and online.... Read more

15:00

Rajat Gupta

Building an Actionable Runbook Platform

Paymenttools
**Actionable runbooks** close the gap between *“what to do”* and *“doing it.”* This talk shows how to design and ship a runbook platform where steps can be clicked and executed safely during incidents. **What the system does** - Create and manage runbooks with tags and Markdown. - Blocks include: instruction, command, API call, conditional, and timer. - Execute a full runbook or a single block with outputs captured in history. - Use RBAC, encrypted credential store, versioning, and containerized environments to keep execution safe and repeatable. - Core entities and API surface: `Runbook`, `RunbookVersion`, `Block`, `ExecutionJob`, `Credentials`, plus endpoints for runbooks, versions, execution, and credentials. **Architecture at a glance** - React SPA communicates with a FastAPI backend and MongoDB. - An execution worker runs jobs and streams results. **Demo flow** 1. Create a runbook with tags and Markdown instructions. 2. Add a command block and an API call block that uses a stored credential. 3. Assign a custom Docker execution environment to the runbook. 4. Run a single block, then run the entire runbook and watch outputs land in history. **What you will learn** - Design principles for truly actionable runbooks and how they differ from static docs. - How to implement safe execution with RBAC, audit, and container isolation. - Patterns for versioning and rollbacks so teams can iterate without fear. - How this approach complements existing incident tooling and industry guidance on making runbooks actionable ([Incident][1], [resources.rundeck.com][2]). **Who should attend** - SRE, platform, security, and backend engineers who own on-call and incident response. - Engineering managers who want safer self-service for ops tasks. ... Read more

15:30

Hakoub Esfahani

Modular Terraform, Unit Testing IaC, and CI/CD

Veritas Technologies
Terraform modules are the backbone of modern infrastructures, but, unit testing even in small DevOps teams is a challenge. In this talk I would like to present multiple approaches to address this challenge by deploying ephemeral environments using CI/CD.... Read more

16:00

Networking & sponsor crawl

Main lobby

16:30

Oscar Manzano

The Augmented SRE Leader: Steering Human-Machine Synergy in the Age of AIOps

Roche Diagnostics
SRE leaders must move beyond automation to lead augmented teams. They must leverage human-AI collaboration and champion explainable AI along with upskilling teams to master this new symbiosis, ensuring AI becomes a reliable partner. The future is human-machine collaboration.... Read more

17:00

Jake Page

Get More Bang for Your Error Budget's Buck: 3 Ways to Extend Reliability Without Changing Your SLOs

MetalBear
As an SRE, you rarely get to choose the SLO. But you can influence how the error budget is spent by making smaller, more frequent releases, testing in creative ways and learning to effectively nudge leadership to have a higher change tolerance, you can extend the value of your error budget.... Read more

17:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room
09:30 Keynote: The Rise of AI-driven SRE
Joao Freitas • PagerDuty
10:00 Keynote: Management is a Hard Job. That’s Why You Should Do It Like an Engineer
William Mendes • Coralogix
10:30 Keynote: When Incidents Fix Themselves: AI SRE in action
Birol Yildiz • ilert
11:00 Coffee break
11:30 Increase observability, not costs: controlling telemetry data ingestion at scale
Diogo Cebola & Andre Bento • Mollie
12:00 Beyond 'Done': Strategies for Improving OpenTelemetry Instrumentation Quality
Yuri Oliveira Sa • OllyGarden
12:30 FinOps at Scale: Lessons from Scaling Kubernetes Workloads
Ricardo Miguel Magalhaes • xgeeks
13:00 When Infrastructure as Code Becomes a Product: Lessons from the Trenches
Luis Serra • synvert, a GlogalLogic company
13:30 Lunch & networking
14:30 Building End-to-End Observability for AI Agents
Murilo Venturin & Ricardo Moreira • PagerDuty
15:00 Building an Actionable Runbook Platform
Rajat Gupta • Paymenttools
15:30 Modular Terraform, Unit Testing IaC, and CI/CD
Hakoub Esfahani • Veritas Technologies
16:00 Networking & sponsor crawl
16:30 The Augmented SRE Leader: Steering Human-Machine Synergy in the Age of AIOps
Oscar Manzano • Roche Diagnostics
17:00 Get More Bang for Your Error Budget's Buck: 3 Ways to Extend Reliability Without Changing Your SLOs
Jake Page • MetalBear
17:30 Wrap up

Speakers

Birol Yildiz
ilert
Diogo Cebola
& Andre Bento
Mollie
Hakoub Esfahani
Veritas Technologies
Jake Page
MetalBear
Joao Freitas
PagerDuty
Luis Serra
synvert, a GlogalLogic company
Murilo Venturin
& Ricardo Moreira
PagerDuty
Oscar Manzano
Roche Diagnostics
Rajat Gupta
Paymenttools
Ricardo Miguel Magalhaes
xgeeks
William Mendes
Coralogix
Yuri Oliveira Sa
OllyGarden

Venue

PagerDuty

Allo | Alcântara Lisbon Offices, Av. da Índia 10,
1300-299 Lisboa, Portugal

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one