SREday

Site Reliability, DevOps and Cloud

April 14, 2025 Redmond, WA, USA

1
Day
16+
Speakers
1
Track
150+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

Fairwinds, Google, Komodor, Microsoft, Pomerium

Topics so far:

We'll announce more talks soon, stay tuned!

This is a past event, what's next?

Schedule

April 14, 2025 • single track • 10AM - 3:30PM • Redmond, in-person
view as table
main room • Track 1

10:00

Miko Pawlikowski

KeynoteSRE’s worst nightmare

SRE Author
What's the worst that can happen? Join me for a story about reliability, secrecy and potential 200k fatalities.... Read more

10:30

Coffee break

Main lobby

11:00

Andrew Suderman

Case Study: Re-Thinking Our Infrastructure Tooling

Fairwinds
When you're managing dozens of Kubernetes clusters, across three different clouds, for dozens of individual companies in their own accounts, the challenge of (re)designing tooling is complex. Come hear how we worked through all the many possible options (centralized IaC vs templating, whether to use cluster managers like Rancher, etc.), what high-level tradeoffs were made (i.e. ease-of-use vs speed of changes vs centralized control), and the tangible outcomes of this process. The focus here will be on the process and decision making, given all the different drivers such as technology, business, people, process, etc. This talk will cover how we undertook the process, while giving you tangible next steps to take back to your desk. From this talk, you'll take away: - Some reasons why you should or should not attempt a rewrite of your Infrastructure-as-Code process. - An overview of the inputs that you should consider when designing an Infrastructure-as-Code tooling stack. - Some idea of how to be successful with a tooling rewrite.... Read more

11:30

Nick Taylor

Zero Trust: From Airports to Identity-Aware Proxies

Pomerium
Zero Trust doesn't have to be intimidating. Learn how Identity-Aware Proxies transform service access from perimeter-based to continuous verification, explained through the universal experience of airport security.... Read more

12:00

Sangeetha Rajkumar

Urgent Cloud Deployment Practices and Insights

Microsoft
In today's fast-paced world, Agile & security is paramount. Urgent Cloud Services Deployment Practices ensure seamless and secure cloud deployments under tight timelines. Gain actionable insights, maintain robust security, and enhance team coordination for efficient, reliable cloud infrastructure.... Read more

12:30

Lunch & networking

Main lobby

13:30

Gary Edgar

Hidden Signals in K8s Clusters: What Your Systems are Trying to Tell You

Komodor
Learn how to identify hidden patterns in cluster behavior that impact reliability. Discover methods to detect issues like pod delays and service dependencies, using clustering algorithms and proactive strategies for improved resilience and resource optimization.... Read more

14:00

Peter De Tender

Stress-Testing tools in Azure to optimize your workload reliability

Microsoft
For years, you've hopefully been relying on the Microsoft Well-Architected Framework (WAF) as a starting point to have your workloads set up in a highly available, redundant architecture. Following those guidelines is key to get reliably running business critical applications for your users or customers. But even after following all the guidelines, did you really know your applications would run as expected? With this session, I'm going to take the guesswork out of that question, and walk you through different tools in Azure, Azure Load Testing and Azure Chaos Studio, which help by stress-testing your critical workloads, running in App Services, Containers, CosmosDB or traditional Virtual Machines. As usual in Peter's sessions, this will be packed with live demos and only a bare minimum of slides. ... Read more

14:30

Leonid Yankulin

SLOs that Delight: Mastering Multi-Service SLO Calculations

Google
This talk is for DevOps and SRE engineers who are looking to optimize the number of SLOs in their system and find methods to manage a large number of various SLIs for their complex, multi-level service architecture. The talk covers practical methods for calculating SLOs based on multiple SLIs and dependencies, including the reliability of third-party service providers and unknowns. Attendees are expected to be familiar with basic SRE principles and terminology. The talk includes some probability math. Overwhelmed by the complexity of managing SLOs across your microservices architecture? This talk provides DevOps and SRE engineers with actionable strategies for defining, calculating, and optimizing SLOs in systems with a large number of SLIs. We’ll explore practical techniques for: - **SLO Targets:** Right-guessing and improving your SLO goals. - **SLO Optimization:** Right-sizing the number of SLOs to ensure effective monitoring without unnecessary overhead. - **SLO Composition:** Mastering aggregation and other methods to derive meaningful SLOs from diverse SLIs. - **Dependency Management:** Incorporating the reliability of third-party services and unknowns into your SLO calculations. The talk will include some probability math to illustrate key concepts. ... Read more

15:00

More Pizza Time!

Main lobby

15:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room
10:00 Keynote: SRE’s worst nightmare
Miko Pawlikowski • SRE Author
10:30 Coffee break
11:00 Case Study: Re-Thinking Our Infrastructure Tooling
Andrew Suderman • Fairwinds
11:30 Zero Trust: From Airports to Identity-Aware Proxies
Nick Taylor • Pomerium
12:00 Urgent Cloud Deployment Practices and Insights
Sangeetha Rajkumar • Microsoft
12:30 Lunch & networking
13:30 Hidden Signals in K8s Clusters: What Your Systems are Trying to Tell You
Gary Edgar • Komodor
14:00 Stress-Testing tools in Azure to optimize your workload reliability
Peter De Tender • Microsoft
14:30 SLOs that Delight: Mastering Multi-Service SLO Calculations
Leonid Yankulin • Google
15:00 More Pizza Time!
15:30 Wrap up

Speakers

Andrew Suderman
Fairwinds
Gary Edgar
Komodor
Leonid Yankulin
Google
Miko Pawlikowski
SRE Author
Nick Taylor
Pomerium
Peter De Tender
Microsoft
Sangeetha Rajkumar
Microsoft

Venue

Microsoft Reactor Redmond

3709 157th Ave NE, Redmond,
WA 98052, United States

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one