SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.
This panel will examine the challenges, lessons learned, and best practices around deploying, testing, troubleshooting, observability, and developer experience for large-scale IoT systems. We’ll demystify IoT + SRE/DevOps and showcase real-world perspectives.... Read more
Customer impacting incidents happen. No team knows where the bodies are buried better than DevOps and for that reason they are often called on to play a key role. When incident response is fragmented however, teams struggle to respond quickly and effectively. Once the immediate issue is resolved, gaps in visibility, accountability, and responsibility remain, leaving leaders and teams unsure of how to prevent future incidents or improve outcomes. This keynote will explore how these shortcomings emerge during and after an incident, and provide Dev leaders with practical guidance on bridging the gaps between DevOps, ITOps, ITSM, and ITIM. Learn how to break down silos, streamline the incident management process, and create a unified approach that drives efficiency, accountability, and better long-term results.... Read more
SLOs allow teams to prioritize the more impactful or important aspects of their services. Metrics that center the user experience gives teams focus and goals. Highlighting the metrics that matter most gives teams space to disable alerts that contribute nothing to overall customer happiness.
Automation gives us more hands to deal with more issues without taking time away from more interesting work. Machine Learning tools group alerts together and help add context when things are noisy and distracting.
These tools combined help teams tackle a common incident response problem - alert fatigue. Full Service Ownership teams looking to improve the quality of their services can experiment with their SLIs and SLOs to find the work that will benefit their systems the most while also preserving their own sanity. Prevent the stress and anxiety that can arise from unexpected system failures by setting clear expectations and allowing for planned responses to potential issues based on pre-agreed norms.... Read more
How should AI based automated test generation make us rethink the software testing pyramid?
We compare two automated test generation approaches (LLM vs SBST), discuss open source tools & industry use cases, and provide developers with recommendations for integrating AI into their testing workflows.... Read more
At Cast AI, we’ve developed Container Live Migration to automatically consolidate these workloads, ensuring continuous uptime, reducing resource fragmentation, and cutting costs. Join us to see how we’re making Kubernetes work for stateful applications in a practical demo.... Read more
All services emit telemetry data, but ensuring it is useful can be challenging. Too much, and you have a lot of noise; too little, and you can’t properly identify and troubleshoot issues. In 2023, it became clear to us at Mezmo that we had complex high cardinality metrics, needed to improve our observability practices, and consolidate our multiple Prometheus instances. We wanted that elusive “single pane of glass” experience.
This session will detail our transformation from a Sysdig-centric approach to a more flexible, centralized observability strategy. We'll explore how we:
* Migrated from Sysdig to OpenTelemetry Collector
* Implemented data transformation techniques to enhance metric context
* Consolidated Prometheus instances
* Developed targeted data identification and processing methods... Read more
We SREs and infrastructure engineers love GitOps. After all, what’s not to love about declarative infrastructure-as-code, a single source-of-truth, change tracking, and automated CI/CD pipelines that make the magic happen?
However, we must ask whether GitOps is the silver bullet that solves all problems -- including climate change and world peace. Or are there limitations and situations where GitOps might not be the best approach?
This talk will explore why and how Bloomberg's cloud-native infrastructure engineering team migrated from GitOps to what we are calling “OperatorOps" -- a mix of declarative specifications and Kubernetes operators to handle Day 2 operations for our fully managed etcd-as-a-service system.... Read more
Fancy a peek into how AWS FIS bring Chaos Engineering into AWS Lambda? As usage of serverless technology grows, chaos engineering for serverless becomes even more crucial for ensuring reliable and available applications. Join us as we demo new capabilities for testing AWS Lambda, and unpack how these new faults have been built and run under the hood. Finally, learn valuable lessons gleaned from our customers' experiences with modern serverless applications.... Read more
From a CEO's perspective, integrating DevOps with machine learning pipelines is key to strategic advantage, driving innovation, market agility, and operational efficiency. This presentation underscores real-world successes and views DevOps as crucial for future business growth.... Read more
Race cars are built for speed, but a challenging track can hold them back. Is the network slowing down your AI/ML jobs? Join me to learn how to boost speed and resilience by tackling link flapping and network congestion. Turbocharge your AI/ML workloads, just like a race car at top speed.... Read more
In this talk, Matan Nataf, co-founder, and CEO of Rapydo, addresses the complexities of managing relational databases in the cloud era. Focusing on the challenges brought by microservices and multi-tenant architectures, he underscores the limitations of current tools in achieving scalability and performance. Nataf highlights the key pillars for modern database management: observability, resiliency, and cost-efficiency. He introduces Rapido's solutions that offer consolidated visibility and automated query management to mitigate database issues. Finally, a live demo showcases Rapido's capabilities in optimizing database performance and stability.... Read more
Join for an interesting talk on 'Secret Management'. We'll dive into the challenges of protecting sensitive data in dynamic, distributed environments. Learn about tools that offer secure, scalable, and reliable solutions. Don't miss this opportunity to enhance your DevOps practices!... Read more
17:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
Stateful Workloads Made Easy: A Practical Demo of Live Migration
Abstract
At Cast AI, we’ve developed Container Live Migration to automatically consolidate these workloads, ensuring continuous uptime, reducing resource fragmentation, and cutting costs. Join us to see how we’re making Kubernetes work for stateful applications in a practical demo.
Bio
As Field CTO, Phil is responsible for working with customers to educate and encourage kubernetes best practices that lead to optimal cloud costs. With more than 15 years of experience in a wide range of positions he is able to balance resiliency, performance and cost to help customers achieve their goals.
Previously, Phil was a Director of Engineering for Security Products at Oracle cloud. This experience helped shape his understanding of cloud scale technology and best practices.
Jon Duarte
Mezmo
Streamlining Telemetry Data: Building a Telemetry Pipeline to Handle High-Cardinality Metrics
Abstract
All services emit telemetry data, but ensuring it is useful can be challenging. Too much, and you have a lot of noise; too little, and you can’t properly identify and troubleshoot issues. In 2023, it became clear to us at Mezmo that we had complex high cardinality metrics, needed to improve our observability practices, and consolidate our multiple Prometheus instances. We wanted that elusive “single pane of glass” experience.
This session will detail our transformation from a Sysdig-centric approach to a more flexible, centralized observability strategy. We'll explore how we:
Migrated from Sysdig to OpenTelemetry Collector
Implemented data transformation techniques to enhance metric context
Consolidated Prometheus instances
Developed targeted data identification and processing methods
Bio
Jon Duarte is a Site Reliability Engineer II at Mezmo with expertise in Terraform, Kubernetes, and GitHub. His career journey includes roles as a Linux application support engineer and Microsoft SQL DevOps Database Administrator at iHeartMedia. Jon holds a BBA in Infrastructure Assurance with a minor in Information Systems from The University of Texas at San Antonio. Known for his curiosity and problem-solving skills, Jon is dedicated to continuous learning and improvement.
Sachin Kamboj
Bloomberg
The Evolution of GitOps to OperatorOps
Abstract
We SREs and infrastructure engineers love GitOps. After all, what’s not to love about declarative infrastructure-as-code, a single source-of-truth, change tracking, and automated CI/CD pipelines that make the magic happen?
However, we must ask whether GitOps is the silver bullet that solves all problems -- including climate change and world peace. Or are there limitations and situations where GitOps might not be the best approach?
This talk will explore why and how Bloomberg's cloud-native infrastructure engineering team migrated from GitOps to what we are calling “OperatorOps" -- a mix of declarative specifications and Kubernetes operators to handle Day 2 operations for our fully managed etcd-as-a-service system.
Bio
Sachin Kamboj is one of the principal software engineers working on Bloomberg's internal Kubernetes-as-a-Service (KaaS) platform. He has been involved in implementing various open source cloud-native projects at Bloomberg since 2016.
Sachin loves to play with new technologies and to learn how things really work behind the scenes. To this end, he contributed to some of the open source projects for the Kubernetes community that were started at and published by Bloomberg, including PowerfulSeal and Goldpinger.
Before joining Bloomberg, Sachin was active in the multi-agent systems and distributed systems communities, where he presented his work at various academic conferences and was nominated multiple times for "best paper" awards. He has also presented at various internal Bloomberg conferences, and has mentored and taught the company's new software engineers.
Iris Sheu & Saurabh Kumar
AWS
Serverless Chaos Engineering: AWS FIS Lambda Actions Under the Hood
Abstract
Fancy a peek into how AWS FIS bring Chaos Engineering into AWS Lambda? As usage of serverless technology grows, chaos engineering for serverless becomes even more crucial for ensuring reliable and available applications. Join us as we demo new capabilities for testing AWS Lambda, and unpack how these new faults have been built and run under the hood. Finally, learn valuable lessons gleaned from our customers' experiences with modern serverless applications.
Bio
Iris Sheu is a Sr. Technical Product Manager with AWS Reliability Services, and joined AWS in 2022. As a PM of AWS Fault Injection Service, she enjoys being able to help customers realize the value of resilience testing. She is focused on delivering products and tools that help customers build more confidence in their operational resilience. Iris is based in Washington, DC.
Saurabh Kumar is a Senior Solutions Architect based in North Carolina, USA, with a strong focus on resilience and chaos engineering. He is passionate about helping customers solve their business challenges and technical problems from migration to modernization & optimization leveraging his 20+ years of experience in tech industry. He is also an active contributor to the tech community, having authored several publications on chaos engineering, AWS Fault Injection Service, and observability strategies.
Ben Savage
Veritas Automata
Mastering the DevOps Machine Learning Pipeline for Unrivaled Innovation - A CEO's Perspective on Cool DevOps
Abstract
From a CEO's perspective, integrating DevOps with machine learning pipelines is key to strategic advantage, driving innovation, market agility, and operational efficiency. This presentation underscores real-world successes and views DevOps as crucial for future business growth.
How can DevOps, when seamlessly integrated with machine learning pipelines, become a powerhouse for innovation and competitive advantage? This exploration, from a CEO's perspective, unveils DevOps not just as a collection of practices and tools but as a pivotal asset in strategic business positioning and market leadership. Learn how melding DevOps with machine learning amplifies its strategic importance, driving product innovation, enhancing customer experiences, and facilitating agile responses to market dynamics. We will highlight concrete examples where the synergy of DevOps and machine learning has led to notable business successes, including market expansion, the swift introduction of innovative products or features, and achieving operational efficiencies. The presentation concludes with a visionary outlook on DevOps, enriched with machine learning, as a transformative force in business growth and evolution.
Bio
Ben Savage, serving as the visionary CEO of Veritas Automata, brings to the forefront an extensive background in revolutionizing autonomous transaction processing through the strategic integration of blockchain, smart contracts, and machine learning. His tenure at Veritas Automata is marked by an unwavering commitment to innovation, operational efficiency, and leveraging technology for transformative business solutions. Prior to leading Veritas Automata, Ben was instrumental in shaping the future of IoT solutions as the CTO and Chief Innovation Officer at Apex Supply Chain Technologies, where he pioneered the development of the groundbreaking Pizza Portal for Little Caesars among other complex IoT applications.
With a career spanning over two decades, Ben's expertise encompasses product development, embedded systems, cloud computing, and supply chain management, honed through significant roles at Apex, Maersk, and Computer Science Corporation. His pioneering work has been recognized with numerous domestic and international patents, underscoring his contribution to advancing technology and business practices. Furthermore, Ben's thought leadership and strategic insights have made him a valued member of multiple advisory boards, where he continues to influence the direction of technological innovation and industry standards.
At the helm of Veritas Automata, Ben Savage is not just leading a company; he is steering an industry towards a future where technology serves as a cornerstone for efficiency, security, and growth, embodying the ethos of innovation, improvement, and inspiration that Veritas Automata champions.
Lerna Ekmekcioglu
Clockwork Systems
Turbocharging AI/ML workloads: Revving Up Speed and Resilience
Abstract
Race cars are built for speed, but a challenging track can hold them back. Is the network slowing down your AI/ML jobs? Join me to learn how to boost speed and resilience by tackling link flapping and network congestion. Turbocharge your AI/ML workloads, just like a race car at top speed.
Bio
Lerna is a Senior Solutions Engineer at Clockwork Systems where she helps customers meet their performance goals with software solutions built on Clockwork.io’s foundational research. Prior to this, she was a Senior Solutions Architect serving Global Financial Services customers at AWS for 3 years. Before that, Lerna spent 17 years as an infrastructure engineer in large financial services companies working on authentication systems, distributed caching, and multi region deployments using IaC and CI/CD to name a few. In her spare time, she enjoys hiking, sightseeing and backyard astronomy.
Matan Nataf
Rapydo
Managing Databases in the Cloud in Broken
Abstract
In this talk, Matan Nataf, co-founder, and CEO of Rapydo, addresses the complexities of managing relational databases in the cloud era. Focusing on the challenges brought by microservices and multi-tenant architectures, he underscores the limitations of current tools in achieving scalability and performance. Nataf highlights the key pillars for modern database management: observability, resiliency, and cost-efficiency. He introduces Rapido's solutions that offer consolidated visibility and automated query management to mitigate database issues. Finally, a live demo showcases Rapido's capabilities in optimizing database performance and stability.
Bio
Matan Nataf is the Co-Founder & CEO of Rapydo, leading cloud and database management innovations. Previously, he founded CloudPort and held leadership roles at Spot.io, ControlUp, and HPE, specializing in cloud solutions, virtualization, and enterprise IT. With deep expertise in AWS, DBMS, and team management, he brings years of experience in scaling tech-driven businesses.
Michel Schildmeijer
SSC-ICT
Which vault? Don’t tell me your secrets!
Abstract
Join for an interesting talk on 'Secret Management'. We'll dive into the challenges of protecting sensitive data in dynamic, distributed environments. Learn about tools that offer secure, scalable, and reliable solutions. Don't miss this opportunity to enhance your DevOps practices!
Secret management is a crucial aspect of DevOps, as it protects sensitive data used by applications and services. Secrets can include API keys, credentials, tokens, certificates, and passwords that grant access to various resources and systems. If these secrets are compromised, attackers can exploit them to cause damage, steal information, or disrupt operations. The challenges of secret management is how to securely store, distribute, and rotate secrets in a dynamic and distributed environment. Traditional methods of hard-coding secrets in configuration files or environment variables are not secure, scalable, or reliable. Moreover, secrets need to be updated frequently to comply with security policies and regulations to prevent unauthorized access. To address these challenges, several tools and frameworks have been developed to provide secret management solutions for DevOps.These tools can help DevOps teams to implement best practices for secret management.
Bio
Michel started his career as a medical officer in the Royal Dutch Airforce, with a focus on pharma. After the air force, he continued in pharma, followed by time working in clinical pharmacology. While there, he transitioned to IT by learning UNIX and MUMPS, and developed a system for managing patients’ medical records. As his career developed, his responsibility shifted from a deep technical perspective to a more visionary role.
At the end of 2011, Michel authored a book on WebLogic Administration for beginners. He joined Qualogy in April 2012 where he expanded his repertoire significantly, serving a wide range of customers with his knowledge about Java Application Servers, Middleware and Application Integration. He also increased his multiple-industry knowledge in his role as Solutions or IT architect by working for customers in a range of sectors, including financials, telecom, public transportation and government organizations.
In 2012, he received the IT Industry-recognized title of Oracle ACE for bein
Ale Paredes, Jessica Garson, Brian Annis & Vinny Ruia
Viam, Elastic, Place Exchange & FIreFly Automatix
KeynotePanel Discussion: Smart Deployments: The Intersection of IoT & SRE
Abstract
This panel will examine the challenges, lessons learned, and best practices around deploying, testing, troubleshooting, observability, and developer experience for large-scale IoT systems. We’ll demystify IoT + SRE/DevOps and showcase real-world perspectives.
Bio
Ale Paredes is the Director of Engineering at Viam. With a strong background in tech leadership, she has driven impactful projects at companies like Stripe and Code Climate, advancing product scalability, reliability, and team performance. As the co-founder of Latinas in Tech NYC, Ale advocates for diversity and inclusion in the tech industry. She is passionate about fostering a collaborative environment where everyone can grow their skills, deliver impactful solutions, and build resilient, scalable systems.
Jessica Garson is a Python programmer, educator, and artist. She currently works at Elastic as a Senior Developer Advocate. She previously worked in developer relations at Twitter for four years. She has spoken at conferences worldwide, from PyCon to Write the Docs. She uses code and modular synthesizers to make music and audio-reactive video art in her spare time.
Brian Annis is the Director of Site Reliability Engineering at Place Exchange. Previously he served as the Lead SRE at Intersection, a consortium member of LinkNYC. There he built a novel configuration management platform to manage thousands of disparate IoT devices across the globe. He also worked collaboratively with multiple transit providers including SETPA, CTA, and LA Metro to build innovative digital out of home networks at scale. A passionate believer in knowledge sharing, Brian has spoken at many conferences, including Ansiblefest, to share insights and challenges with the broader community. He currently lives and works in NYC.
Vinny Ruia is a robotics engineer at Firefly Automatix. In the past four years, he has helped develop significant portions of the company’s autonomous mowing program from the ground up. Vinny has also participated in agricultural robotics community events and developed autonomous boats for collegiate competitions. Vinny is excited to see the different ways that robotics and technology will be used to make lives easier.
Phil Christianson
Xurrent
KeynoteIncident Mgmt: IT Ops and DevOps Collide
Abstract
Customer impacting incidents happen. No team knows where the bodies are buried better than DevOps and for that reason they are often called on to play a key role. When incident response is fragmented however, teams struggle to respond quickly and effectively. Once the immediate issue is resolved, gaps in visibility, accountability, and responsibility remain, leaving leaders and teams unsure of how to prevent future incidents or improve outcomes. This keynote will explore how these shortcomings emerge during and after an incident, and provide Dev leaders with practical guidance on bridging the gaps between DevOps, ITOps, ITSM, and ITIM. Learn how to break down silos, streamline the incident management process, and create a unified approach that drives efficiency, accountability, and better long-term results.
Bio
Phil Christianson is the Chief Product Officer at Xurrent, where he drives the strategic direction and execution of the platform roadmap, focusing on metrics-driven leadership to secure a winning market position. With extensive experience in product management, he previously led large teams and multimillion-dollar R&D initiatives at Wayfair, overseeing pricing systems and competitive intelligence. A University of Iowa graduate with a BBA in Management of Information Systems, Phil brings over two decades of expertise in building high-performing teams and delivering innovative solutions across industries.
Mandi Walls
PagerDuty
KeynoteTackling Alert Fatigue with SLOs, Automation, and Machine Learning
Abstract
SLOs allow teams to prioritize the more impactful or important aspects of their services. Metrics that center the user experience gives teams focus and goals. Highlighting the metrics that matter most gives teams space to disable alerts that contribute nothing to overall customer happiness.
Automation gives us more hands to deal with more issues without taking time away from more interesting work. Machine Learning tools group alerts together and help add context when things are noisy and distracting.
These tools combined help teams tackle a common incident response problem - alert fatigue. Full Service Ownership teams looking to improve the quality of their services can experiment with their SLIs and SLOs to find the work that will benefit their systems the most while also preserving their own sanity. Prevent the stress and anxiety that can arise from unexpected system failures by setting clear expectations and allowing for planned responses to potential issues based on pre-agreed norms.
Surbhi Madan
Google
KeynoteRedefining The Software Testing Pyramid with AI-powered test generation
Abstract
How should AI based automated test generation make us rethink the software testing pyramid?
We compare two automated test generation approaches (LLM vs SBST), discuss open source tools & industry use cases, and provide developers with recommendations for integrating AI into their testing workflows.
Cohn’s Test Pyramid has been a foundational aspect of agile development, guiding developers in the delicate balance between unit, integration, and e2e testing. The recent surge in the adoption of AI based code generation necessitates a need to rethink the canonical form of the testing pyramid as test driven development takes on a “shift-left” approach, pushing complex integration testing to earlier in the workflow. In this talk, we start by discussing the two main automated test generation approaches in use today (LLM-based vs SBST-based) and tradeoffs between them, and then we dive into some open source generation software and current industry use-cases using AI based test generation. Finally, we address some common myths surrounding future speculation about advancements in this area, and close with informed recommendations for developers on when to consider AI based solutions for your testing workflows and which approach to use.
Key takeaways for developers:
A practical understanding of the tradeoffs between LLM and SBST-based test generation approaches.
Available test generation tools that can be used for your use-cases. - Recommendations for when and how to integrate automated test generation in your workflow
Bio
Hi, I’m Surbhi, currently working as a Senior Software Engineer at Google in NYC. My background includes 7 years of experience working in Android native client development, platform engineering, performance optimization, backend architectures, and server-driven UIs all of which have involved test-driven development. My work has primarily been in Java and C++. I graduated from Brown University with a degree in Computer Science where I contributed to the department’s undergraduate teaching in data structures algorithms, focusing on making courses inclusive and collaborative and introducing unit testing and the testing pyramid paradigm to the course curriculum, which has since become a mainstay fixture. Since then, I have at Google, contributed extensively towards intern hiring, mentoring early career professionals, and helping foster inclusive team cultures. I enjoy following pro tennis, running, biking, cooking, board games, walking, eating good food, and exploring the city. Reach me at [email protected] .