SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.
The New York Times Building
5620 8th Avenue, 45th Floor,
New York, NY 10018, United States
Sponsors & Partners
Want to become a sponsor? Get in touch!
Tyler Schade
GEICO
Observing the un-observable: instrumentation for your legacy workloads
Abstract
Every enterprise has a brownfield morass of legacy software that isn't exciting anymore, but still runs in production. These workloads are often business-critical, but receive little investment outside of what it takes to keep the lights on. As SREs, these workloads can be particularly annoying - legacy platforms and code without instrumentation often mean our only source of observability data comes from logs. In this talk, I'll show you how you can transparently use cloud-native tricks to instrument an application without touching the application's code, whether it runs on a VM, a container, or elsewhere. Then, I'll show you how the data collected can be used strategically to generate the four golden signals of observability that we as SREs leverage to keep production online. If you're managing the morass of legacy software in production, this talk is for you.
Bio
Tyler Schade is a Distinguished Engineer at GEICO, where he leads software engineering and architecture for GEICO's infrastructure. An avid proponent of open source software, Tyler currently serves on the Istio steering committee and as a maintainer for a number of CNCF-hosted networking and security projects. Prior to GEICO, Tyler held roles at Solo.io, Ironnet, and Comcast, building deep networking, security and distributed systems experience. He lives in Philadelphia with his wife and their newborn.
Dor Amir
Nadir
Cache Is King: The Hidden Cost of LLM Optimization
Abstract
A cheaper model does not always mean a cheaper run. Routing between models, trimming context, splitting work across agents, or adding tools and skills mid-run can look like smart optimizations. But when those changes disrupt prompt-cache reuse, they can turn discounted, reusable context into input you pay to process again—undermining the savings you expected. OpenAI Developers
This talk explores why prompt caching deserves to be an architectural consideration, not an afterthought, in production LLM systems. We’ll examine how caching interacts with model routing, context management, subagents, and dynamic tool loading. Through concrete examples, we’ll compare preserving a warm cache with starting fresh, explore when breaking the cache is worth the trade-off, and show how to make these decisions based on the economics of an entire run rather than the price of an individual request.
Attendees will leave with practical patterns for cache-aware workflows and an observability checklist covering cache reuse, cache-read and write costs, latency, and cost per successfully completed task. The takeaway: don’t optimize away your biggest savings. Optimize the whole run—not just the token count.
Bio
Dor Amir is the founder of Nadir, an LLM infrastructure company focused on making AI systems more efficient. Previously, he worked on machine learning and personalization at Dropbox, Amazon, Guesty, and Fiverr. His work focuses on production AI systems, LLM routing, agents, and reducing the cost and latency of running AI at scale.
Romaric Philogene
Qovery
SRE in the Era of AI: Lessons Learned from Hundreds of Teams I've met Over 3 Years
Abstract
Over the past three years, I've collaborated with hundreds of SRE teams as they've integrated AI into their workflows. What has truly evolved - and what remains the same? This talk will provide practical insights and observations from the forefront of AI adoption in infrastructure. I'll discuss:
The evolving skill set: What SREs need to learn today and which skills stand the test of time
AI in incident management: Identifying where AI agents add value, where they fall short, and how to distinguish between the two
AI SRE: What's real, and what's just hype
Bio
Romaric Philogene is CEO & Co-founder of Qovery, the Agentic infrastructure platform powering hundreds of companies like Alan, Talkspace, Zoom..With 20+ years in infrastructure engineering, he started his career when cloud didn't exist - managing physical servers in HFT, getting 2 AM pages, and building runbooks from scratch. Prior to Qovery (founded 2020), he was CEO of two other companies self funded and sold to major brands in ERP for automotive.
Yossi Eliaz
Incredibuild
Reliability Lessons from Building Sandboxes for AI Agents
Abstract
AI agents now run thousands of ephemeral coding tasks a day, and each one needs a clean, fast, trustworthy environment to execute in. This talk shares hands-on lessons from building secure execution infrastructure for AI workloads at Incredibuild and islo.dev, covering CI runners, sandboxes, snapshots, caching, traces, and reproducible environments. It walks through a real sandbox provider for ephemeral task isolation, presented at an Apache Airflow Monthly Town Hall, and a meta-harness loop that took a coding agent from 0 of 5 to 5 of 5 passing tasks in about two seconds across four proposer steps. Drawing on earlier container security work at Twistlock, the talk covers what breaks when isolation, caching, and observability have to work together at CI speed.
Bio
Yossi Eliaz is a Principal Engineer, AI Systems at Incredibuild, building secure execution infrastructure for AI workloads: CI runners, sandboxes, snapshots, caching, and reproducible environments. He previously worked on container security at Twistlock. He holds a PhD in Physics and an MSc in Computer Science.
Elliot Gardiner
Mirror Panel Inc.
Production Has Observability. Your Coding Agent Doesn’t
Abstract
Claude Code, Codex, and other coding agents are becoming long-running, tool-using systems that read repositories, execute commands, modify code, call external services, and make decisions with limited human supervision. Yet most teams observe them primarily through token counts and final pull requests. This talk applies familiar SRE practices to agentic development. We’ll instrument a coding-agent workflow with OpenTelemetry, examine traces of tool calls, failures, retries, tests and code changes, and look at what those traces reveal about reliability that the final output hides. We’ll cover practical failure patterns, useful signals, and where traditional observability concepts translate well - and where they don’t.
Bio
Elliot Gardiner is a software engineer and former principal SRE with 13 years of experience building and operating large-scale systems. He is the founder of Code Telemetry, an observability platform for AI coding agents, and co-founder of Mirror Panel, a synthetic research platform focused on measuring and improving the reliability of AI-generated audience insight. His current work sits at the intersection of AI systems, observability, developer infrastructure, and reliability.
Aaron Hunter
AWS
The tokenomics of Model Context Protocol at scale
Abstract
Everyone's connecting MCP servers to their AI agents right now. What almost nobody's doing is checking the bill. Every tool you register and every server you connect quietly pads your token count... often before the model does anything useful. This talk digs into the real tokenomics of MCP: why tool schemas are a fixed tax on every request, how that overhead compounds into serious cost and latency at scale, and what you can actually do about it. You'll see how to measure the token cost your MCP setup adds, using live examples, plus practical ways to trim it without losing the tools you rely on. If you're running AI agents in production, or about to be, this is the cost conversation worth having before your invoice has it for you.
Bio
Aaron Hunter is a Principal Developer Advocate at AWS based in Frisco, TX. With over 15 years of experience spanning system administration, networking, and training, he brings more than a decade of cloud expertise to help Engineers, Developers, Builders, and tech enthusiasts master AWS technologies. His philosophy is simple: always be learning something new, and have fun while doing it! Aaron shares his knowledge through workshops, online courses, mentoring, and live streaming – making complex cloud concepts accessible and enjoyable. When he's not building in the cloud, you'll find him exploring craft beer scenes in new cities or teaming up with friends in Marvel Rivals on his PlayStation, because even superheroes need good teammates!
Denys Zhak
Imubit
What AWS Lambda Was Hiding: Runtime Assumptions Exposed by a Migration
Abstract
Earlier this year, I wrapped up a six-month migration of seven Python data-ingestion Lambdas into services. Most of the code moved unchanged into an environment where concurrent deliveries shared a filesystem, memory space, and process. I’ll use five cases to show what changed in practice. Duplicate deliveries collided in /tmp. Six SQLAlchemy engines created per request eventually filled a 4 GiB process. Lambda’s async invocation retries disappeared when the same work moved behind RabbitMQ. A blocking database call caused liveness probes to fail. A DLQ replay kept taking its pod down. In each case, application code depended on behavior supplied by the runtime. Filesystem isolation, process lifetime, delivery semantics, network topology, and resource boundaries formed the runtime contract around our code. The services had a runtime contract of their own. We had to identify which guarantees carried over and which responsibilities now belonged in the application. I’ll finish with a checklist engineers can use to find those dependencies before moving a workload between managed runtimes.
Bio
Denys Zhak is a software engineer with nine years of experience in backend engineering and distributed systems, focused on production reliability. At Imubit, he builds cloud and on-prem software for industrial machine-learning deployments. He led the six-month migration of seven Python ingestion Lambdas into services described in this talk. Outside work, he contributes to open source. Much of that work has been on Astral’s Ruff and ty. He writes about production engineering at denyszhak.com.
Itamar Knafo
Dalton AI
You Build It, You Run It Is Dead: Reliability When the Author Is an Agent
Abstract
You build it, you run it. Fifteen years of SRE gospel, and almost nobody questions it. The rule was never about writing code. It was about a loop. Whoever wrote it understands it, so whoever wrote it carries the pager. Understanding was the point. The pager was just enforcement. Agents cut that loop. The code still ships. The author is gone. I take the assumption apart, one joint at a time. No human author, so who holds the pager. Nobody who can explain it, so the author cannot answer the 3am question. Review cannot keep up with the output, so your gate is a rubber stamp. Then the fixes that do not work, and why each one folds. Human in the loop. Make the agent better. Keep agent code out of prod. Then the reframe. If writing it no longer means understanding it, reliability has to come from somewhere else. Operability stops being something you hope the author cared about, and becomes something you check for. One real incident carries the argument. A single S3 copy call, a 5 GiB limit nobody checked, a migration that never sealed, and a cluster crash-looping while the log line said migrating objects: <nil>. None of it was hard. Nobody could see it. The principle is not dead. The assumption under it is.
Bio
Itamar Knafo is the co-founder and CEO of Dalton AI. Previously, has spent five years building and leading software, DevOps, observability, and operations teams. He is a hands-on engineer with a deep background of cloud infrastructure, and open source, and now focuses on applying AI to production engineering.
Giovanni Martinez
Tiger Data
Agents Can Write Migrations. They Shouldn’t Ship Them.
Abstract
AI agents can write migrations, but production changes need guardrails. This live demonstration walks through an agent-assisted PostgreSQL workflow using Tiger Cloud and the Tiger CLI: provisioning an isolated development service, generating versioned migrations, and testing correctness and performance before human-approved production deployment. We’ll cover scoped access, automated checks, lock and rewrite detection, deployment monitoring, and rollback criteria. Attendees will leave with a practical blueprint for accelerating database delivery without giving agents unchecked authority over production.
Bio
Giovanni Martinez is a Database Support Engineer at Tiger Data and an AWS Community Builder with more than 10 years of PostgreSQL experience. He specializes in database performance, production reliability, cloud-native PostgreSQL, TimescaleDB, and agent-assisted database operations. Through The Postgres Guy and his open-source iqtoolkit project, Giovanni creates practical tools and educational content that help engineers troubleshoot, optimize, and operate PostgreSQL systems safely.
Anand Mattah Chenna Kesavalu
Grainger
Navigate RUM Without Limits: Cutting 80% of Browser Telemetry Noise While Improving Observability
Abstract
Modern observability platforms make it easy to collect browser telemetry. The real challenge is deciding what not to collect.
At Grainger, our digital platforms process approximately 1.5 billion requests per day across multiple brands, countries, and customer-facing applications. As our Real User Monitoring (RUM) adoption expanded, we discovered that a significant portion of our telemetry was generated by bots, synthetic monitoring solutions, automated testing platforms, internal tooling, and duplicate browser events. While these events increased observability costs, they provided little value when understanding real customer experiences.
This session shares our journey modernizing browser observability during a Datadog RUM migration and transforming RUM from a "collect everything" model into a scalable, business-focused observability platform.
Attendees will learn how we:
• Identified and filtered bot, synthetic, and automated traffic before it reached our observability platform
• Reduced browser event ingestion by nearly 80% (7.81 million to 1.67 million events) while maintaining visibility into real user experiences
• Improved session accuracy by addressing split-session behavior, cookie inconsistencies, and telemetry fragmentation
• Implemented intelligent filtering across edge, application, and browser layers to reduce noise and improve signal quality
• Balanced engineering observability, business analytics, compliance requirements, and platform costs
• Established a repeatable framework for determining which browser telemetry delivers operational value and which should be excluded
Beyond the technical implementation, this talk explores the broader challenge facing modern SRE and observability teams: collecting the right data rather than simply collecting more data.
The result was a cleaner signal, faster incident investigations, reduced observability spend, and greater confidence that our dashboards, alerts, and analytics reflect real customer behavior—not automated traffic.
Whether you use Datadog, OpenTelemetry, New Relic, Dynatrace, Elastic, or another observability platform, you'll leave with practical techniques and architectural patterns to navigate RUM without limits.
Bio
Anand is a Staff Software Engineer and Platform SRE at Grainger, where he serves as the enterprise subject matter expert for CDN, edge security, bot management, observability, and web performance. He is responsible for reliability and performance initiatives supporting more than 1,000 services and APIs across Grainger's digital platforms.
Anand has over 20 years of experience in software engineering, site reliability engineering, and cloud infrastructure. He has led large-scale implementations involving Datadog, Akamai, DataDome, AWS, and enterprise observability platforms. His work includes browser observability optimization, bot mitigation, edge security modernization, and platform reliability engineering. Anand is also the author of a several guest blogs at leading vendors including Datadog RUM [https://www.datadoghq.com/blog/grainger-rum-costs-from-bot-traffic/]
He holds numerous industry certifications including AWS Machine Learning Speciality, AWS Solutions Architect Professional, AWS Certified Generative AI Developer - Professional, Datadog Certified Professional, Github Copilot, Terraform Associate, Splunk certifications, Akamai Professional, Gremlin Certified Chaos Engineer and Google GenAI credentials etc.,
A frequent conference speaker, Anand has presented at multiple Grainger technology conferences on observability, bot management, and security engineering.