SREday

Site Reliability, DevOps and Cloud

June 25, 2026 PagerDuty, Lisbon, Portugal

1
Day
10+
Speakers
1
Track
50+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

Elastic, Five9, Gigapipe, Miniclip, Miro, PagerDuty, Sateliot, Spectro Cloud, Tripadvisor, UNREAL Performance

Topics so far:

This is a past event, what's next?

Schedule

June 25, 2026 • single track • 9:30AM - 5PM • Lisbon, in-person
view as table
main room • Track 1

09:30

Joao Freitas

10:00

Ralph Bird

AI Coded It, AI Shipped It, Who Owns It? The Future of SREs in Agentic Operations

PagerDuty
As AI agents increasingly generate and deploy code into production, the traditional DevOps mantra "Code it. Ship it. Own it." is being disrupted. In this talk, I will explore how agentic AI is changing the software development lifecycle and what happens when AI “owns” large portions of the pipeline. Code development, operational knowledge and even incident response are increasingly being delegated to AI systems, but that doesn't guarantee infallibility! What happens when it breaks and nobody has seen the code before? I will examine how AI is already changing both software engineering and SRE practices, where the advances are effective and where human expertise remains critical. I will challenge you to rethink how we do service ownership, what skills are required for SREs in the agentic age, and how to collaborate with AI to ensure resilient, trustworthy systems. Is the DevOps revolution over, or are we simply at its next frontier?... Read more

10:30

Coffee break

Main lobby

11:00

Tiago Costa

Reliability Patterns for AI-Powered Apps in Azure AI Foundry

Azure Cloud & AI Architect and Advisor
AI applications promise transformative capabilities but introduce unique failure modes, such as hallucinations, latency spikes, data drift, and cost overruns, that traditional SRE practices don't fully address. This session explores proven reliability patterns, retries with exponential backoff, circuit breakers, graceful degradation, and human-in-the-loop guardrails tailored for Microsoft Foundry workloads, where observability covers agent traces, quality metrics, and safety evaluations. You'll see live demos of implementing these in Foundry projects using zone-redundant storage, multi-region failover, and built-in content filters to achieve production-grade resilience while maintaining trust and compliance.... Read more

11:30

David Jacovkis

DevOps in Space: Lessons from Low Earth Orbit

SateliotWatch
Remember when you knew all your servers by name and deployed changes with a single SSH command? Now imagine doing that when your link comes in 10-minute windows every few hours, and bandwidth is a precious commodity measured in kbps. Welcome to DevOps in Space. At Sateliot, our Infrastructure & Software Engineering team works at the intersection of aerospace, telecommunications and software. We build modern distributed systems and then send them to an environment where normal expectations -low latency, immediate feedback, up-to-date systems- simply don’t apply. In this talk I’ll give a high-level tour of the challenges that make space and telco different: intermittent and low-bandwidth connectivity, long feedback loops, constrained devices, and the cultural gaps between aerospace, telco and software teams. I'll also share some adaptations that allow us to apply some of the DevOps processes and tools that have become industry standards.... Read more

12:00

Ricardo Amaro

From DevOps Adoption to Reliable Delivery

Elastic
DevOps is widely adopted in name, but successful adoption in practice remains uneven across organizations. In this talk, I will present the main outcomes of my PhD research on how IT organizations can improve DevOps adoption. I will focus on the capabilities, metrics, and lifecycle practices that have the strongest impact on delivery performance and operational reliability. I will share a practical framework that helps teams identify common adoption barriers, connect engineering practices to measurable outcomes, and design improvement strategies that support both faster delivery and more resilient operations. The session is intended for leaders and practitioners who want a clearer, evidence-based view of what makes DevOps work beyond tooling and slogans.... Read more

12:30

Rita Lopes

404: Observability Without Trade-offs Not Found - Until We Built Gigapipe

GigapipeWatch
Every SRE team knows the drill. Prometheus for metrics. Loki for logs. Jaeger or Tempo for traces. Pyroscope bolted on last for profiling. Four services, four storage backends, four things to scale, debug, and pay for. And when an incident hits, you're jumping between all of them trying to correlate signals that were never designed to live together.The trade-off is always the same: observe everything and pay an astronomical bill, or cut costs and fly blind precisely when it matters most. We got tired of it. So we built Gigapipe. Gigapipe is an open-source polyglot observability stack that unifies logs, metrics, traces, and profiles on a single backend, using ClickHouse as the database. It natively implements the industry standards like Prometheus, Loki, Jaeger, Tempo, Pyroscope, and OpenTelemetry APIs, just to name a few. Your existing agents, exporters, and Grafana dashboards work on day one, unchanged. In this talk, we'll cover three things: 1. Why we built it: the real pain that made four services feel unsustainable, 2. How it works: the architecture decisions behind native multi-protocol compatibility on a single storage engine, 3. How you can get involved or run it yourself today: Gigapipe is open-source, under AGPLv3 license, runs as a single binary, and is live in production environments. This is a talk for engineers who've felt the weight of their observability stack and wondered if there's a better way. Spoiler: There is! And you can run it tonight.... Read more

13:00

Lunch & networking

Main lobby

14:00

Mauro Morales

From Boot to Rollback: How Image-Based Operating Systems Change Kubernetes Operations

Spectro CloudWatch
You replace pods instead of patching them. So why is the OS underneath your nodes still the last hand-patched snowflake in an otherwise declarative stack? Image-based operating systems turn nodes into versioned artifacts that can be upgraded and rolled back predictably, reducing drift and making failure recovery part of the platform design rather than an incident-time improvisation. As a Kairos maintainer, I will show how this model works in practice: what remains immutable, what persists, how atomic upgrades and rollback work, and what happens when a node fails to boot. We will then use Kubernetes to drive an upgrade and rollback of an immutable node. Because application reliability is difficult enough without discovering during an incident that no one knows exactly what the node underneath it has become.... Read more

14:30

Helia Barroso & Iris Dyrmishi

Observability Through Users' Eyes and How to Fix the Knowledge Disconnect

Five9 & MiroWatch
Logs, metrics and traces nowadays are common words in every engineer's vocabulary, as Observability is everywhere with the rise of OpenTelemetry and improved standards. Instrumentation? No problem. Getting data from any system? Done. But then soon after, things start to fall apart, telemetry piles up unused, cloud bills soar, dashboards rot and alerts confuse even seasoned engineers. Clearly there is something missing. In this talk, Iris and Hélia share some experiences from building Observability platforms, patterns they have discovered of how users interact with telemetry data. From common questions across different teams and companies, pitfalls and also, strategies that empower teams to move from confusion to confidence. So get practical insights to bridge the gap between Observability data and User Knowledge, whether you're building platforms, supporting teams, or trying to make sense of your own telemetry chaos.... Read more

15:00

Pooria Ghaedi

AI-Assisted Cloud Troubleshooting

MiniclipWatch
Cloud infrastructure incidents are often hard to troubleshoot because the problem is not always visible from the application layer. A service may look healthy, the instances may be running, and the dashboards may not clearly show the issue, but connectivity can still fail because of routing, security rules, DNS, or cloud configuration problems. In this talk, I will show how AI-assisted workflows can support SREs and cloud engineers during infrastructure troubleshooting. The focus is not on using AI to replace engineering judgment, but on using it to investigate faster, ask better questions, review configuration, interpret CLI output, and reduce the time spent jumping between documentation, consoles, and logs. As a practical demo, I will use AWS MCP servers to investigate a real-world style cloud connectivity issue and work toward the root cause. We will look at what context is useful to give the AI assistant, how MCP can help connect the assistant to AWS documentation and environment information, and where the engineer still needs to validate the output. The goal is to give attendees a realistic view of what AI can and cannot do during real infrastructure incidents. The session will focus on practical troubleshooting patterns, safe usage of AI during production-like investigations, and the importance of human validation when dealing with cloud reliability and infrastructure security.... Read more

15:30

Networking & sponsor crawl

Main lobby

16:00

Miguel Borges

AI Agents in Production: What Actually Breaks

UNREAL Performance
Running AI agents in production looks nothing like the demos. Over the past 18 months I have shipped four production AI systems as a solo engineer - multi-tenant customer service on WhatsApp, invoice extraction pipelines, HR automation, and an agent fleet management tool. Each one broke in ways I did not expect. This talk is a practical post-mortem on the failure modes that matter in production AI: LLM calls that silently return wrong answers with high confidence, cascading cost spikes that signal upstream failures before users notice, and observability gaps that make debugging a model’s decision nearly impossible. I will walk through the reliability patterns that actually helped - confidence cascades that cut cost per document by 85%, human-in-the-loop as a circuit breaker rather than a fallback, and treating AI coding agents as infrastructure with explicit boundaries. No benchmarks, no theory. Just what breaks, why it breaks, and what I changed.... Read more

16:30

Tiago Rodrigues

How Easter sparked the hunt for villains

Tripadvisor
Following Easter 2024, our messaging platform underwent a significant transformation. This presentation explores how an engineering team transitioned from developing a simple messaging product to constructing a phishing detection platform. More importantly, how we adopted a mindset akin to fraudsters, always brainstorming potential attack vectors. Join us at this engineering storytelling session that delves into how Tripadvisor leveraged straightforward tools such as text analysis, domain detection and asynchronous processing to implement a phishing detection capability while maintaining a seamless user experience.... Read more

17:00

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room
09:30 Introduction, meet & greet
Joao Freitas • PagerDuty
10:00 AI Coded It, AI Shipped It, Who Owns It? The Future of SREs in Agentic Operations
Ralph Bird • PagerDuty
10:30 Coffee break
11:00 Reliability Patterns for AI-Powered Apps in Azure AI Foundry
Tiago Costa • Azure Cloud & AI Architect and Advisor
11:30 DevOps in Space: Lessons from Low Earth Orbit
David Jacovkis • Sateliot
12:00 From DevOps Adoption to Reliable Delivery
Ricardo Amaro • Elastic
12:30 404: Observability Without Trade-offs Not Found - Until We Built Gigapipe
Rita Lopes • Gigapipe
13:00 Lunch & networking
14:00 From Boot to Rollback: How Image-Based Operating Systems Change Kubernetes Operations
Mauro Morales • Spectro Cloud
14:30 Observability Through Users' Eyes and How to Fix the Knowledge Disconnect
Helia Barroso & Iris Dyrmishi • Five9 & Miro
15:00 AI-Assisted Cloud Troubleshooting
Pooria Ghaedi • Miniclip
15:30 Networking & sponsor crawl
16:00 AI Agents in Production: What Actually Breaks
Miguel Borges • UNREAL Performance
16:30 How Easter sparked the hunt for villains
Tiago Rodrigues • Tripadvisor
17:00 Wrap up

Speakers

David Jacovkis
Sateliot
Helia Barroso
& Iris Dyrmishi
Five9 & Miro
Joao Freitas
PagerDuty
Mauro Morales
Spectro Cloud
Miguel Borges
UNREAL Performance
Pooria Ghaedi
Miniclip
Ralph Bird
PagerDuty
Ricardo Amaro
Elastic
Rita Lopes
Gigapipe
Tiago Costa
Azure Cloud & AI Architect and Advisor
Tiago Rodrigues
Tripadvisor

Venue

PagerDuty

Allo | Alcântara Lisbon Offices, Av. da Índia 10,
1300-299 Lisboa, Portugal

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one