SREday

Site Reliability, DevOps and Cloud

April 11, 2025 San Francisco, CA, USA

1
Day
16+
Speakers
2
Tracks
100+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

Altinity, Arista Networks, AWS, CAST AI, Fairwinds, Harness, Ironridge, Kyndryl, Mezmo, Neubird AI, Pomerium, Postman, Randoli, Rapydo, reboot.dev, Revyl.ai, RSI Security, Salesforce, Sawmills, Tigris Data, Twilio

Topics so far:
Incident Management

This is a past event, what's next?

Schedule

April 11, 2025 • 2 parallel tracks • 10AM - 5:30PM • San Francisco, in-person
view as table
main room • Track track 1

10:00

Puneet Saraswat

KeynoteFrom Growing Pains to Enterprise Scale: How Harness Transformed Its Infrastructure

Harness
Scaling infrastructure is never just about adding more machines—it’s about evolving architecture, managing complexity, and maintaining reliability while growing rapidly. In this session, we’ll take you through the journey of how Harness scaled its infrastructure to support growing customer demands, improve performance, and enhance operational efficiency. From database optimizations to rethinking our deployment strategies, learn key takeaways and practical lessons that can help any engineering team build for scale.... Read more

10:30

Coffee break

Main lobby

11:00

Francois Martel

AI Teammates for SREs: How will they impact SRE Teams?

Neubird AI
Explore how AI teammates are revolutionizing SRE work by handling routine investigations, providing context-aware analysis, and enabling teams to focus on engineering. Learn about real implementation challenges and how to prepare your team for this transformation.... Read more

11:30

Gunnar Grosch

A proactive approach to resilience in modern applications

AWS
Discover how adopting a proactive resilience strategy can help modern applications withstand failures and maintain performance. This talk explores key techniques for identifying vulnerabilities, implementing fault tolerance, and ensuring continuous availability in dynamic environments.... Read more

12:00

John Shin

Awareness Security in the age of A.I.

RSI Security
In an increasingly digital world, where technology and data have become integral parts of our daily lives, the importance of cybersecurity cannot be overstated. This topic on **"Awareness Security"** is an exciting opportunity to enlighten your mind about the significance of this subject, the threats it addresses, and the measures individuals and organizations can take to protect themselves in the digital realm. ... Read more

12:30

Scott Davis

Visibility, Insight, and Action with Cast AI, Prometheus, and Grafana

CAST AI
Achieving full-stack observability requires seamless integration of monitoring, analysis, and automation. In this session, we’ll explore how Cast AI, Prometheus, and Grafana work together to provide real-time visibility, actionable insights, and automated optimizations for cloud-native environments. Learn how to leverage these powerful tools to monitor performance, optimize costs, and take proactive actions to enhance system reliability and efficiency.... Read more

13:00

Lunch & networking

Main lobby

14:00

Tucker Callaway

Streamlining Telemetry Data: Building a Telemetry Pipeline to Handle High-Cardinality Metrics

Mezmo
Handling high-cardinality telemetry data efficiently is crucial for modern observability systems. In this session, we will explore strategies for designing a scalable telemetry pipeline that can process large volumes of diverse metrics without performance bottlenecks. We’ll cover best practices for data ingestion, storage optimization, and query performance, along with real-world techniques to mitigate challenges like high dimensionality and resource constraints. Attendees will gain insights into building a robust telemetry pipeline that balances accuracy, efficiency, and cost-effectiveness.... Read more

14:30

Matan Nataf

Managing Databases in the Cloud in Broken

Rapydo
In this talk, Matan Nataf, co-founder, and CEO of Rapydo, addresses the complexities of managing relational databases in the cloud era. Focusing on the challenges brought by microservices and multi-tenant architectures, he underscores the limitations of current tools in achieving scalability and performance. Nataf highlights the key pillars for modern database management: observability, resiliency, and cost-efficiency. He introduces Rapido's solutions that offer consolidated visibility and automated query management to mitigate database issues. Finally, a live demo showcases Rapido's capabilities in optimizing database performance and stability.... Read more

15:00

Nick Taylor

Zero Trust: From Airports to Identity-Aware Proxies

Pomerium
Zero Trust doesn't have to be intimidating. Learn how Identity-Aware Proxies transform service access from perimeter-based to continuous verification, explained through the universal experience of airport security.... Read more

15:30

Robert Hodges

Fast, Cheap, DIY Observability with Open Source Analytics and Visualization

Altinity
Commercial observability tools are expensive and complex - but you can build a fast, cost-effective solution yourself! This talk shows how to use ClickHouse, OpenTelemetry, and Grafana to create scalable, DIY monitoring with open source tools.... Read more

16:00

David Argent

A Tale of Two Outages

Salesforce
Having been at ground zero for two outages that you can still look up on Google, this talk is designed to give an oral history of the causes, reactions, solutions, and aftermath to the Danger Sidekick outage and another major outage whose company I cannot mention by name. The first is a story about how a single individual decision cascaded into a massive failure that took six weeks to clean up and was deemed an impossible recovery by both Oracle and the SAN Vendor. It was arguably one of the largest cases of putting Humpty Dumpty back together again, as a 15TB Oracle DB was reassembled, one 4K page at a time, to provide 99.9% data recovery on what was deemed “impossible”. We called it the Lazarus Project for a reason. We had PC desktop tower units littering the aisles of our otherwise Unix datacenter, as they were the recovery fleet. The third-party data recovery team that helped achieve the impossible built their recovery tools on Windows. The second is a story about how several choices and design flaws can all come together in a perfect storm. We had poor designs involving infinite retries on a non-caching interface. We had no testing of content posted minutes after the start of a major sale which led to back-end loads from one internal service hitting the main NoSQL DB being larger than the entire service was designed to handle. There were design constraints in the NoSQL platform making it unable to shed “dead” transactions in queue, lengthening the time to recover. We had an overall rendering engine requiring hundreds to thousands of calls to all succeed in order to render a single retail page, and a need to retry all calls in the event of a single failure. The human decision to start the major sale globally at the same time, rather than staggering it according to time zones was the capstone in a Taj Mahal of an outage. While you couldn’t see it from the outside, over 98% of all transactions were successful against the NoSQL DB even at the worst part of the outage, though the design decisions elsewhere led to a much lower observed availability than that.... Read more

16:30

Himank Chaudhary

Building a distributed persistent queue on FoundationDB

Tigris Data
Tigris is a globally distributed object storage system where objects are stored all over the world. The system needs to have asynchronous tasks to keep objects cached, redundantly stored, and moved around in response to changes in access patterns. To solve this, we built an asynchronous task system on top of FoundationDB, just like our metadata. This simplifies our architecture by letting us enqueue tasks in the same transaction as we insert data, providing us atomicity with all or nothing semantics. This wouldn’t be as easy if we used an external system because we’d need distributed coordination between our database and our message queue. We also wanted to take advantage of our existing FoundationDB clusters and expertise to avoid having to make things too complicated for our SREs. Today I’ll cover how FoundationDB makes this easy, the other advantages it’s given us, and some lessons we learned in the process.... Read more

17:00

Rod Anami & Gregory Pruett

The duality of adopting AI: Can SREs become AI/ML engineers?

Kyndryl
AI/ML is not new in the business world. It has been used for some time, but Generative AI (GenAI) initiated a new disruptive force in recent years. Many businesses and technical processes are taking advantage of embedded GenAI capabilities. However, companies must put more effort into leveraging AI patterns and pre-formatted AI/ML solutions, which require AI-specialized engineering capacity, a skill that is scarce in today's market. Also, AI-powered apps have intricate complexities, such as setting up indexing pipelines to continually update RAG data, continuous processes for data preparation, model maintainability policies, and guardrails for generated AI outputs. Site Reliability Engineers (SREs) have been working with AI for over a decade, using AIOps tools to make sense of large amounts of observability data, making them acquainted with AI/ML. On the other hand, the AI/ML model and MLOps pipeline observability have many challenges and hurdles to overcome. Can SREs become future AI/ML or LLMOps engineers? Can they apply site reliability engineering principles and practices to AI/ML models? This session sheds some light on possible paths for this question since SREs can both use and support AI deployments, therefore, a duality in embracing AI.... Read more

17:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
mini room • Track track 2

10:30

Coffee break

Main lobby

11:00

Andrew Suderman

Case Study: Re-Thinking Our Infrastructure Tooling

Fairwinds
When you're managing dozens of Kubernetes clusters, across three different clouds, for dozens of individual companies in their own accounts, the challenge of (re)designing tooling is complex. Come hear how we worked through all the many possible options (centralized IaC vs templating, whether to use cluster managers like Rancher, etc.), what high-level tradeoffs were made (i.e. ease-of-use vs speed of changes vs centralized control), and the tangible outcomes of this process. The focus here will be on the process and decision making, given all the different drivers such as technology, business, people, process, etc. This talk will cover how we undertook the process, while giving you tangible next steps to take back to your desk. From this talk, you'll take away: - Some reasons why you should or should not attempt a rewrite of your Infrastructure-as-Code process. - An overview of the inputs that you should consider when designing an Infrastructure-as-Code tooling stack. - Some idea of how to be successful with a tooling rewrite.... Read more

11:30

Shubhanshu Surana

CAP Theorem Reloaded- AKA How to Optimize Distribution of Telemetry Data

Sawmills
The backbone of SRE - like all engineering, is monitoring and observability. When designing and implementing platforms, understanding how monitoring and observability telemetry data impacts your systems is critical to scale.... Read more

12:00

Anca Ghenade

Cloud Integration Testing Made Easy for Your AWS Cloud Apps

Postman
Integration testing for cloud-native AWS applications is complex due to service dependencies, deployment intricacies, and high costs. To make integration testing faster and easier, this presentation shows you how to emulate AWS environments locally using Testcontainers and LocalStack. By simulating real-world scenarios, we can test cloud apps with AWS services without relying on mocks or remote AWS accounts. This approach improves test coverage, reduces costs, and enables quick, isolated testing, bringing the simplicity of unit tests to integration tests for cloud applications.... Read more

12:30

Riley Scheid

Reliable Serverless Needs Distributed Transactions

reboot.dev
Cloud native and serverless application platforms give teams encapsulation, flexibility, and reduced deployment dependencies. But the movement onto the cloud has so far trained us to accept that decomposing your application into multiple loosely coupled functions or services requires eventual consistency (due to event sourcing/buses, queues, durable execution, etc.) Instead, we’d like to propose that transactions are a perfect fit for serverless, and how SREs can enable application developers to not (always) settle for eventual consistency. The closer your applications are to being developed like a monolith, but deployed as decoupled serverless components, the easier they are to develop, deploy, and scale! To that end, we'll talk about how reliable, consistent distributed transactions enable your teams to develop serverless applications like a monolith, and why the CAP theorem has faded in relevance for modern applications.... Read more

13:00

Lunch & networking

Main lobby

14:00

Ren Lee

Fearless SREs

Arista Networks
In this talk, we'll explore both the technical but human and very emotional side of what makes mythical SREs perform at their best in an incident – but also how we not only find such engineers but help others grow into one. Together, we will explore the following: - **Fearless ≠ Heroism/Invincibility** - **Unique Qualities that Seem Mythical** - **Fear is a Risk** - **Natural Inclination vs. Trainable Skills – Do They Overlap?** - **What Does Your Incident Response Training Look Like?** - **Your Role as Lead/Manager Within SRE** - **Celebrating Your Fearless SRE Team** ... Read more

14:30

Mohit Menghnani

Unlocking Observability with React & Node.js

Twilio
This talk shares the secrets of observability with React and Node.js! We will discuss practical strategies for debugging, monitoring performance, and delivering seamless user experiences. The session will end with tips for reforming your dev workflow and ensuring your apps run flawlessly.... Read more

15:00

Anam Hira

DragonCrawl: Revolutionizing Mobile Testing with Generative AI

Revyl.ai
The challenges of traditional mobile testing at scale (3,000+ simultaneous experiments) Architecture and implementation of DragonCrawl using MPNet and embedding techniques Real-world examples of DragonCrawl's adaptive behavior and problem-solving capabilities Practical strategies for handling LLM challenges like hallucinations and adversarial cases Results and metrics from production deployment Live demonstration of DragonCrawl in action We'll explore the technical details of model selection, embedding evaluation, and the specific guardrails implemented to ensure reliable testing. Attendees will see how we achieved 99%+ stability in production while eliminating maintenance overhead.... Read more

15:30

Robin Sarkar

Driving Solar Energy Efficiency with Predictive Analytics and AI: The Role of Site Reliability in Renewable Energy

Ironridge
The solar energy sector is at the forefront of innovation, leveraging predictive analytics and artificial intelligence (AI) to tackle the operational complexities of large-scale solar farms. As global solar capacity surpasses 760 GW, these advanced technologies are transforming site reliability, enabling operators to predict equipment failures, optimize panel performance, and ensure grid stability. With applications such as AI-powered performance optimization increasing energy yield by 2.8% and predictive maintenance reducing downtime by 18%, the role of SRE practices in solar energy systems has become critical. This session will delve into the intersection of AI, IoT, and SRE principles in renewable energy, highlighting technical challenges, solutions, and the growing market for AI-driven tools in solar energy. Attendees will gain data-driven insights into deploying predictive analytics for enhanced reliability, reduced operational costs, and improved energy production. Discover how SREs can lead the way in accelerating AI adoption for a sustainable energy future.... Read more

16:00

Rajith Muditha Attapattu

Mastering Kubernetes Cost Optimization for Sustainable Cloud Operations

Randoli
Cloud-native platforms like Kubernetes offer unparalleled flexibility and scalability, but they often come with a hidden price tag. Without intentional cost management, organizations risk overspending due to inefficient resource utilization, over-provisioning, and lack of visibility into cloud expenses. As cloud bills rise, optimizing Kubernetes costs is no longer optional—it's a necessity for maintaining financial sustainability while delivering value. This session will provide a practical guide to minimizing costs in Kubernetes environments while maximizing resource efficiency. We will explore strategies for gaining visibility, rightsizing workloads, leveraging cost-aware scheduling, intelligent auto-scaling, implementing resource quotas and promoting accountability to avoid unnecessary expenses. The discussion will include leveraging tools like OpenCost, Prometheus, and Kubernetes-native features to monitor and control spending in real time to support above mentioned strategies. By the end of this session, participants will walk away with actionable insights and tools to identify and eliminate cost drains in their Kubernetes clusters, ensuring that every dollar spent contributes to business outcomes. This talk is ideal for SREs, DevOps engineers, and decision-makers looking to optimize cloud expenditures without sacrificing operational needs.... Read more

16:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room mini room
10:00 KeynoteFrom Growing Pains to Enterprise Scale: How Harness Transformed Its Infrastructure
Puneet Saraswat • Harness
10:30 Coffee break
11:00 AI Teammates for SREs: How will they impact SRE Teams?
Francois Martel • Neubird AI
Case Study: Re-Thinking Our Infrastructure Tooling
Andrew Suderman • Fairwinds
11:30 A proactive approach to resilience in modern applications
Gunnar Grosch • AWS
CAP Theorem Reloaded- AKA How to Optimize Distribution of Telemetry Data
Shubhanshu Surana • Sawmills
12:00 Awareness Security in the age of A.I.
John Shin • RSI Security
Cloud Integration Testing Made Easy for Your AWS Cloud Apps
Anca Ghenade • Postman
12:30 Visibility, Insight, and Action with Cast AI, Prometheus, and Grafana
Scott Davis • CAST AI
Reliable Serverless Needs Distributed Transactions
Riley Scheid • reboot.dev
13:00 Lunch & networking
14:00 Streamlining Telemetry Data: Building a Telemetry Pipeline to Handle High-Cardinality Metrics
Tucker Callaway • Mezmo
Fearless SREs
Ren Lee • Arista Networks
14:30 Managing Databases in the Cloud in Broken
Matan Nataf • Rapydo
Unlocking Observability with React & Node.js
Mohit Menghnani • Twilio
15:00 Zero Trust: From Airports to Identity-Aware Proxies
Nick Taylor • Pomerium
DragonCrawl: Revolutionizing Mobile Testing with Generative AI
Anam Hira • Revyl.ai
15:30 Fast, Cheap, DIY Observability with Open Source Analytics and Visualization
Robert Hodges • Altinity
Driving Solar Energy Efficiency with Predictive Analytics and AI: The Role of Site Reliability in Renewable Energy
Robin Sarkar • Ironridge
16:00 A Tale of Two Outages
David Argent • Salesforce
Mastering Kubernetes Cost Optimization for Sustainable Cloud Operations
Rajith Muditha Attapattu • Randoli
16:30 Building a distributed persistent queue on FoundationDB
Himank Chaudhary • Tigris Data
Wrap up
17:00 The duality of adopting AI: Can SREs become AI/ML engineers?
Rod Anami & Gregory Pruett • Kyndryl
17:30 Wrap up

Speakers

Anam Hira
Revyl.ai
Anca Ghenade
Postman
Andrew Suderman
Fairwinds
David Argent
Salesforce
Francois Martel
Neubird AI
Gunnar Grosch
AWS
Himank Chaudhary
Tigris Data
John Shin
RSI Security
Matan Nataf
Rapydo
Mohit Menghnani
Twilio
Nick Taylor
Pomerium
Puneet Saraswat
Harness
Rajith Muditha Attapattu
Randoli
Ren Lee
Arista Networks
Riley Scheid
reboot.dev
Robert Hodges
Altinity
Robin Sarkar
Ironridge
Rod Anami
& Gregory Pruett
Kyndryl
Scott Davis
CAST AI
Shubhanshu Surana
Sawmills
Tucker Callaway
Mezmo

Venue

Harness.io HQ

55 Stockton St, San Francisco,
CA 94108, United States

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one