SREday

Site Reliability, DevOps and Cloud

June 27, 2025 Catawiki, Amsterdam

1
Day
15+
Speakers
1
Track
100+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

Absolute Value, AuthZed, Catawiki, Coralogix, DR. Insight, ITQ, Jumbo, MeloMar IT, OSInet, Payble Technologies, Swisscom, Twingate

Topics so far:

This is a past event, what's next?

Schedule

June 27, 2025 • single track • 9AM - 5PM • Amsterdam, in-person
view as table
main room • Track track 1

09:00

Pablo Martin Guijo

From Local to Cloud-Based Development: Our Migration Story

Catawiki
What does it take to move an entire engineering organization from traditional local development setups to developing directly in the cloud? At Catawiki, this shift was more than just swapping tools—it was a cultural and technical transformation aimed at improving reliability, feedback loops, and developer velocity. This talk walks through the highlights of our journey: from the early signs that local development was holding us back, to the decisions, challenges, and breakthroughs that shaped our move to a cloud-based development workflow.... Read more

09:30

Coffee break

Main lobby

10:00

Sohan Maheshwar

How Google built a Consistent, Global Authorization System (and you can too!)

AuthZed
Google Zanzibar is the singular authorization service that powers permissions and sharing across all Google properties, including Docs, YouTube, and Cloud IAM. Creating a consistent, global-scale authorization system that can process "more than 10 million client queries per second” is not a trivial task. The talk will cover how the paper lays out an engineer-friendly blueprint for building a highly scalable distributed system with flexible consistency guarantees. This talk will start with foundational knowledge of Relationship Based Access Control (ReBAC) and then cover the technical implementations behind Zanzibar - How Google solved for correctness, scale and speed. The presentation will cover the different APIs for interacting with the system and also a deep-dive into how the “New Enemy” problem was solved. The talk will conclude with how you can use open source tools to build authZ into your application.K... Read more

10:30

Wouter Lagerweij & Krijn Mossel

Governing the Cloud: A Guide to Data Sovereignty

Independent & DR. Insight
With changing political relations and deteriorating trust in the large cloud providers, Data Sovereignty is on everyone’s mind. In this session, we will show you that the hard work you’ve already done on the road to DevOps and SRE will serve you just as well to make your systems sovereign by design. Join us to learn how to: - Apply architectural patterns like hybrid cloud and edge computing to enforce data locality and control. - Use IaaS software engineering principles to enhance the portability of your infrastructure. - Challenge the hyperscaler default by "right-sizing" infrastructure, - Make compliance and governance the default by automating reporting and transparency. You’ll come away from this talk with a set of tweaks and changes to your current approach to your infrastructure that will give you the flexibility to deal with the challenges of data sovereignty in a world of shifting alliances.... Read more

11:00

Frederic G. Marand

TLA+: Your Secret Weapon Against Concurrency Hell

OSInet
Let's be honest: writing correct concurrent and distributed code is hard. Race conditions? Deadlocks? They're the bane of modern software development, especially with tools like Kafka, Kubernetes, and microservices. Testing helps, but it rarely finds all the subtle, timing-dependent bugs. Imagine having a way to mathematically verify your design's logic before you code it. That's TLA+. Developed by Turing Award winner Leslie Lamport, it's a language for specifying and checking algorithms, especially concurrent ones. The catch? Most intros are dense and theoretical. This session is different. It's TLA+ for the rest of us – the software engineers in the trenches. We'll skip the deep math and focus on practical application. You'll learn what TLA+ is, how to model simple concurrent interactions, and how it can save you from deployment disasters by catching design flaws early. Walk away ready to explore how formal methods can become your secret weapon for building bulletproof systems. Key Takeaways: - Recognize the inherent difficulties in validating concurrent designs. - Get a practical, hands-on feel for TLA+ concepts and syntax. - Understand how TLA+ helps prevent race conditions and deadlocks at the design stage. - Be inspired to apply formal verification techniques to real-world problems.... Read more

11:30

Martin McLarnon

Using OpenTelemetry to Improve Service Level Objectives

Coralogix
As systems grow more complex, engineering teams must move beyond guesswork when resolving production issues and instead focus on delivering consistent, measurable user value. This talk explores how OpenTelemetry empowers teams to define, measure and improve Service Level Objectives with precision and confidence. Drawing on real-world experiences of a software team navigating the challenges of legacy code, cloud migration and production instability, this session demonstrates how OpenTelemetry helps teams gain the insights they need to reduce Mean Time To Resolution, understand root causes and proactively improve system reliability. We’ll dive into hands-on examples using the OpenTelemetry Demo Application and show how to help turn technical metrics into business-aligned SLOs that actually mean something. Whether you're an SRE, developer, or engineering leader, this talk will give you the clarity—and tools—to move from reactive firefighting to proactive service reliability.... Read more

12:00

Daniel Bojczuk

DevOps: Why do we keep messing up?

Jumbo
Summary: DevOps isn't a new concept, but a lot of companies still struggle to get it right. Even when they think they've nailed it, they often run into some not-so-fun side effects, like unhappy developers or unexpected costs that sneak up on them. In this presentation, we'll see why this can happen and how to dodge these pitfalls.... Read more

12:30

Mike Kotsur

Platform as a Product: Drive Adoption for Your Internal Tools

Absolute Value
Are your platform team's brilliant ideas struggling for adoption and meeting resistance? Why do some internal platform tools thrive while others flop? Many technically sound SRE/DevOps initiatives fall short because they aren't treated as products with real users. This session cuts through the noise, revealing why a product mindset is crucial for platform success. The best SREs walk a mile in developers’ shoes! I’ll share with you a potent blend of Lean, User Journey, and Story Mapping principles to uncover genuine user needs (your fellow engineers!) and business goals. Learn how to define impactful MVPs for your internal tools that actually reduce cognitive load and boost productivity. I'll share a practical framework, refined across diverse projects, to ensure your platform work delivers real value, gets adopted quickly, and fosters collaboration. Leave with actionable strategies and a board template to start small, validate assumptions, and build platforms your developers will love.... Read more

13:00

Lunch & networking

Main lobby

14:00

Devrim Demiroz

Lightning Thoughts: Charting the AlphaEvolve → OTel Journey

Swisscom
In this session, I invite you to join a thought experiment I conducted just after DeepMind AlphaEvolve reached the event horizon. I’ll explore how evolutionary AI usage architectures might weave into the OpenTelemetry‑powered observability landscape of the CNCF, turning runtime traces and static code into a self‑optimizing loop. No deep dives—just sparks of ideas to inspire the next steps in agent‑driven software evolution. I hope this quick sketch encourages questions and collaboration.... Read more

14:30

Roosevelt Elias

Reliability at the Edge: Building SRE Culture in Resource-Constrained Environments

Payble Technologies
How do you build reliable systems where downtime is not just an inconvenience, but a threat to trust, income, and even safety? In this talk, Roosevelt Elias, founder of Payble, explores how to establish SRE principles in markets where infrastructure is unreliable, cloud access is intermittent, and talent pipelines are still emerging. Drawing from real-world experience building financial infrastructure in Africa, we’ll discuss culturally aware incident management, lightweight observability stacks, distributed troubleshooting with limited tooling, and what it means to bake resilience into the DNA of both product and process even when the odds are stacked against you. This talk is ideal for SREs, platform engineers, and product leaders building for emerging markets or aiming to design more resilient systems globally.... Read more

15:00

Marcel Koert

Why Your DevOps Isn’t Reliable (And It’s Not the Tools)

MeloMar IT
We blame tools. Kubernetes crashed. Terraform misfired. Alerts didn’t trigger. But what if the real reason your DevOps isn’t reliable… is Dave? In this engaging and eye-opening session, we flip the script on traditional DevOps blame. I’ll show you why the root cause of most outages isn’t technical—it’s human. From missed Slack messages to rushed hotfixes and undocumented changes, it’s behavior—not infrastructure—that most often undermines reliability. This talk dives deep into the eight behavioral pillars of reliability. You’ll hear real-world examples of how poor communication, lack of ownership, and overreliance on tools have taken down multi-million-euro systems. We’ll explore how strong team habits, collaborative culture, and psychological safety do more for uptime than your latest CI/CD upgrade ever could. Expect thought-provoking insights on: • Why communication is your first defense—not a soft skill. • How processes protect more than they restrict. • Why blameless postmortems aren’t about being nice—they’re essential for truth. • What it really means to build a learning culture in SRE and DevOps. You’ll walk away with practical strategies to strengthen your team’s human infrastructure and reshape how you think about reliability. Because here’s the kicker: Tools don’t build resilient systems—people do. If you’ve ever seen an outage sparked by a well-meaning engineer, or if you’re wondering why your metrics look fine but your reliability feels fragile, this talk is for you. Join me to explore the most overlooked layer of your stack: the human one.... Read more

15:30

Michael Meelis

Building Platforms breaking code: how your org chart wrecks everything

ITQ
Do you wonder why your developers are slow to deliver? Are you stuck in an ever-growing loop of Jira tickets? Does communication break down as you scale? If your company's org chart is more complex than your product architecture, this talk is for you. Business and software development are governed by principles like Brook's Law, Murphy’s Law and Conway’s Law. In this talk I will show that we can break Murphy's law! Let's think outside of the box and inverse the communication patterns. Showing that Conway was right all along. We’ll explore Conway’s Law and Domain-Driven Design (DDD) to bridge the gap between tech and business—and I’ll prove that we are not the problem, the org chart is. TL;DR; Sit back and relax and let me explain how we can use Conway's law and DDD to optimise the developer experience.... Read more

16:00

Birol Ertekin

Reducing Blast Radius

Twingate
As site reliability / infrastructure engineers we do a lot of things for better reliability. Some of them are easy, especially if your cloud provider supports it, like adding HA to your database. And some of them require more thorough process and planning , for example reducing blast radius. You can start small with multi-zone / multi-region setup for your compute, but then you will most likely still end up with SPOFs like database and load balancer. Sharding is not new and there are many ways to accomplish it. We'll go over what Twingate did for it's own sharding strategy in this talk that eliminated all single point of failures and reduced the impact of database migrations and infrastructure changes.... Read more

16:30

Joao Rosa

Leveraging Team Topologies for Software Evolution

Independent Consultant
Have you ever encountered roadblocks in software development due to disjointed team structures or interactions? You’re not alone. Misalignments between software and the domain, siloed teams focusing on discrete tasks, or processes dictating software architecture often result in rigid software that fails to meet the customer's needs. Enter Team Topologies, a pattern language, and a set of principles and practices to ensure a smooth flow of value while honoring human-centric aspects like trust boundaries and cognitive load. This perspective raises a question: What if we allow team interactions to evolve beyond the organization chart? What would happen to the reliability of the technical systems? What would such a world look like? Join me for this talk to explore real-world use cases, understand the Team Topologies principles, and the skills required in a digital world that is constantly changing.... Read more

17:00

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room
09:00 From Local to Cloud-Based Development: Our Migration Story
Pablo Martin Guijo • Catawiki
09:30 Coffee break
10:00 How Google built a Consistent, Global Authorization System (and you can too!)
Sohan Maheshwar • AuthZed
10:30 Governing the Cloud: A Guide to Data Sovereignty
Wouter Lagerweij & Krijn Mossel • Independent & DR. Insight
11:00 TLA+: Your Secret Weapon Against Concurrency Hell
Frederic G. Marand • OSInet
11:30 Using OpenTelemetry to Improve Service Level Objectives
Martin McLarnon • Coralogix
12:00 DevOps: Why do we keep messing up?
Daniel Bojczuk • Jumbo
12:30 Platform as a Product: Drive Adoption for Your Internal Tools
Mike Kotsur • Absolute Value
13:00 Lunch & networking
14:00 Lightning Thoughts: Charting the AlphaEvolve → OTel Journey
Devrim Demiroz • Swisscom
14:30 Reliability at the Edge: Building SRE Culture in Resource-Constrained Environments
Roosevelt Elias • Payble Technologies
15:00 Why Your DevOps Isn’t Reliable (And It’s Not the Tools)
Marcel Koert • MeloMar IT
15:30 Building Platforms breaking code: how your org chart wrecks everything
Michael Meelis • ITQ
16:00 Reducing Blast Radius
Birol Ertekin • Twingate
16:30 Leveraging Team Topologies for Software Evolution
Joao Rosa • Independent Consultant
17:00 Wrap up

Speakers

Birol Ertekin
Twingate
Daniel Bojczuk
Jumbo
Devrim Demiroz
Swisscom
Frederic G. Marand
OSInet
Joao Rosa
Independent Consultant
Marcel Koert
MeloMar IT
Martin McLarnon
Coralogix
Michael Meelis
ITQ
Mike Kotsur
Absolute Value
Pablo Martin Guijo
Catawiki
Roosevelt Elias
Payble Technologies
Sohan Maheshwar
AuthZed
Wouter Lagerweij
& Krijn Mossel
Independent & DR. Insight

Venue

Catawiki

Sint Jorissteeg 2, Floor 4, 1012 XV Amsterdam, Netherlands

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one