SREday

Site Reliability, DevOps and Cloud

June 6, 2026 Datadog, New York, US

1
Day
10+
Speakers
2
Tracks
80+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

Altinity, Antimetal, AWS, Bank Of JuliusBaer, Cleric, Datadog, DocuSOR, Extraterrestrial Incorporated, FUSSMOBILE, Game Plan Tech, Gatling, GitGuardian, LoopStudio, Palo Alto Networks, Providence, Seismic, Teladoc Health

Topics so far:
Resilience & Chaos Engineering
Production Engineering

This is a past event, what's next?

Schedule

June 6, 2026 • 2 parallel tracks • 9AM - 4:30PM • NYC, in-person
view as table
main room • Track 1

09:00

Intro

KeynoteIntro by Mark Pawlikowski

Walking everyone through the agenda, warming up for a great event !... Read more

09:30

Ajuna Kyaruzi

KeynoteObservability and SRE in the AI Era

DatadogWatch
Agentic tools and AI-assisted code review have made shipping faster than ever. The reliability practices that keep those systems running are under pressure to match. The systems themselves are changing too. Most teams are already running multiple models in production. When something breaks, the cause is often a rate limit, a prompt update, or a model that changed upstream rather than anything that would show up in a deploy log. Ajuna Kyaruzi, Manager of SRE and Platform Advocacy at Datadog, will share how we can keep reliability in step with development velocity. The observability signals that help give us the complete picture, how incident response changes when systems can drift without a deployment, and what teams operating at scale are learning.... Read more

10:00

Coffee break

Main lobby

10:30

Heather Thacker

Choose Your Weapon: The Performance Testing Arsenal

GatlingWatch
Your app works great on your laptop, in the dev environment. Then production hits 10x expected traffic during a marketing campaign and everything falls apart. Or maybe not, instead six months of data accumulates, causing the response times to be painfully slow. Load testing, stress testing, soak testing, and spike testing, they all sound similar, but address completely different problems, and most teams only do one, if at all. This talk breaks down the essential types of performance testing every app and web developer should understand. Learn when to use each approach, what problems they uncover, and how to integrate them into your development workflow without drowning in complexity. We'll cover real-world scenarios where each testing type saves production systems, helping you choose the right weapon for your performance battles. What you'll learn: - The differences between load, stress, soak, and spike testing, and when each matters - Which performance testing types reveal which production problems, before they happen - How to integrate performance testing into CI/CD without slowing down development and deployments - Practical criteria for deciding which tests your application actually needs... Read more

11:00

Shreyas Iyer

Building a World Model for Production

AntimetalWatch
Antimetal builds production engineering agents that operate across hundreds of customer software systems. To be effective, these agents require the same context an experienced engineer has: what depends on what, what changed recently, how failures propagate, what each piece means to the team that runs it. Most of that information is not explicitly documented anywhere. To solve this problem, we’ve built a world model: a unified, machine-legible representation of any customer’s software systems. This talk presents the engineering and process behind constructing it, including a provider-agnostic ontology, linking runtime to code, streaming updates in, causal reasoning and inference, among much more. We close with where the model breaks down today and the open problems we are still working through.... Read more

11:30

Ian Miller

Shifting Left - Evolution of the Software Development Life Cycle

SeismicWatch
Over the past decade, four trends in enterprise software have reduced the need for large, specialized teams and placed operational capabilities directly in the hands of individual engineers: the shift from monoliths to micro-services, Kubernetes, Site Reliability Engineering, and AI-assisted development. Each trend follows the same pattern. What once required a dedicated specialist or team became a tool or practice that any engineer could own. AI-assisted development accelerates this pattern most dramatically, compressing capabilities previously owned by specialized teams directly into the development workflow. Think about QE and E2E test generation and maintenance, security analysis, error budget analysis, and performance testing - these capabilities historically required large amounts of human and technical capital to execute properly. As AI agents begin to handle increasingly critical aspects of engineering, the trajectory is clear: the agentic software development lifecycle is one where specialized expertise shifts left, resulting in each individual engineer becoming the focal point of shipping products.... Read more

12:00

Lunch & networking

Main lobby

13:00

Josh Lee

OpenTelemetry: Playtime Is Over

AltinityWatch
So you finally got your organization to invest in OpenTelemetry. You carefully evaluated observability backends and picked the perfect one. Everything is awesome. Then twelve months later, your costs have skyrocketed and you can’t explain why. What happened? This talk examines how to emit meaningful telemetry while keeping costs under control, by exploring the following: - What to actually instrument - Which metrics to focus on - Pipeline efficiency with OTel Arrow Applying sampling, filtering, and intentional instrumentation to cut down on noise Schema management and validation with tools like Weaver We’ll review the ingredients of a mature observability implementation with OpenTelemetry: one that grows with you instead of overwhelming you. You’ll learn how to apply cost-effective techniques to achieve meaningful observability. Speaker Notes (visible to organizers only) As more organizations embrace OpenTelemetry and mature their Observability practice, we find ourselves coming out of that Observability and OpenTelemetry "honeymoon period". We've gone from, “We’re using OpenTelemetry, therefore we have Observability” to "How do we actually make this work for us?" This talk will equip organizations to build a sustainable and long-lasting observability practice built on OpenTelemetry.... Read more

13:30

Dwayne McDaniel

From Pets To Cattle To Agents: Evolving Identity And Security For Workloads

GitGuardianWatch
“Treat servers like cattle, not pets” captured one of the biggest shifts in how we run infrastructure. Going from servers we named with masking tape on the case to infrastructure-as-code to deploy thousands of processes changed the mental model around authentication and authorization. Now, AI agents are forcing another mental shift, as we anthropomorphize these nondeterministic autonomous processes. Many people in the industry are now contemplating how we should handle delegation, security, and identity for these agentic systems. We will look to answer why human and non-human identity management diverged and what that means to various parts of our organizations. This talk will look at the state of authentication and authorization across services today and how we got here. The audience will walk away with a much better sense of where we are headed with machine-to-machine authentication, covering: - SPIFFE/SPIRE - WIMSE and Workload Identity Tokens (WIT) - Security Token Services (STS) - Federated Identities on Cloud Providers... Read more

14:00

Greg Spektor

AI for Product Management: Working Faster, Smarter, and More Strategically

Extraterrestrial IncorporatedWatch
Product Managers are expected to balance strategy, customer needs, stakeholder communication, prioritization, analytics, documentation, and delivery, often while operating under constant time pressure. As AI capabilities mature, they offer a powerful opportunity to reduce administrative overhead and accelerate many of the information-heavy tasks that consume a PM’s day. In this session, Greg Spektor explores how AI can augment the product management lifecycle across customer discovery, product strategy, prioritization, documentation, analytics, and delivery coordination. Through practical examples, attendees will learn how AI can synthesize research, analyze customer feedback, generate product documentation, automate reporting, identify risks, and support better decision-making. The talk also examines the limitations of AI, including hallucinations, bias, governance concerns, and the areas where human judgment remains irreplaceable. Attendees will leave with a practical framework for adopting AI within product organizations and a roadmap for moving from quick wins to more strategic AI-enabled workflows.... Read more

14:30

Networking & sponsor crawl

Main lobby

15:00

Willem Pienaar

Perfect Reasoning, Wrong Answer

Cleric
RLHF rewards confident, complete answers. Production incident investigation requires the opposite: withholding judgment, maintaining competing hypotheses, and seeking disconfirming evidence. The training incentive points directly away from the task. At Cleric, we built an AI agent that investigates production incidents. When a cascading failure produces dozens of correlated anomalies, the agent latches onto the loudest signal and builds a coherent narrative around it. The reasoning is sound. The answer is wrong. We taxonomized our agent's production failures and found that premature convergence (the agent finding a plausible answer and stopping) is one of the most common failure modes. Red herrings in production aren't random noise. They're symptoms that look like causes: a downstream service timing out because an upstream database is slow. The timeout is the loudest signal. The database is the root cause. This talk covers the architectural patterns we use to counteract it: forcing multiple competing hypotheses before commitment, requiring evidence that distinguishes between hypotheses rather than confirms the leading one, and a devil's advocate layer that argues against the top hypothesis. We also cover causal inference techniques that make the problem tractable, cascading failures propagate with measurable latency through dependency graphs, and dozens of anomalies collapse into a short causal chain when you follow timestamps instead of severity.... Read more

15:30

Edward Rodriguez

Kubernetes at the Edge: SRE Lessons from Disconnected Clusters

FUSSMOBILE
Most Kubernetes reliability discussions assume stable cloud networking, elastic infrastructure, and always-available control planes. But the real operational challenge starts when Kubernetes has to run closer to the edge: in constrained, bandwidth-limited, intermittently connected, or partially air-gapped environments where standard cloud assumptions no longer hold. This talk shares practical SRE lessons from operating managed Kubernetes environments across edge and hybrid infrastructure using technologies such as GKE Enterprise, EKS Anywhere, Bottlerocket OS, GitOps workflows, observability tooling, and production-grade operational runbooks. I will cover what changes when clusters are no longer “just in the cloud”: upgrade planning, image and artifact distribution, node OS lifecycle, observability under constrained bandwidth, incident response, storage behavior, and the tradeoffs between automation and safe human control. The session is intended to be candid and technical, focused on lessons learned rather than theory. Attendees will walk away with a practical mental model for designing and operating Kubernetes platforms in environments where connectivity is imperfect, upgrades require choreography, observability must be selective, and reliability depends as much on operational discipline as it does on tooling.... Read more

16:00

Paul Caplan

Agent Harnesses: From Slot Machines to Safety Nets

Teladoc Health
An agent is made up of just two parts: the LLM and the harness. But what exactly is a harness? Is it the core “ReAct loop,” tool interface, and context management? Or the instructions, memory layer, skills, and context that you provide? Spoiler alert: the answer is “yes.” This talk is for anyone who wants more out of their agents, whether you're building agentic applications, using coding agents like Claude Code or Codex, or all of the above. More predictability and reliability. Higher quality output. How to be the engineer in the loop, not the agent babysitter. As many in the industry move toward “thin” harnesses, this talk argues for a strong “outer” harness: a coordination and control layer outside the core agent itself. Let the model reason and judge. Let the harness make the work repeatable, observable, and safe.... Read more

16:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
mini room • Track 2

10:00

Coffee break

Main lobby

10:30

Anisha Manoharan

When Milliseconds Cost Millions: Designing Low-Latency AI Systems in Fintech

Bank Of JuliusBaerWatch
In fintech systems, milliseconds don’t just impact performance — they directly translate into financial risk. AI-driven decisions such as fraud detection, credit scoring, and transaction authorization operate under strict latency budgets, often under 100 milliseconds. When these systems slow down, the consequences go beyond user experience: delayed decisions can increase exposure, impact compliance, and lead to measurable financial loss. This talk dives into the engineering challenges of running low-latency AI systems reliably in production. We’ll explore how to define meaningful SLOs for AI workloads, manage tail latency (p95/p99), and design resilient inference pipelines on Kubernetes. Through real-world scenarios, we’ll unpack trade-offs between model accuracy and response time, how bottlenecks emerge under scale, and what typically fails first in high-pressure environments. We’ll also cover observability patterns for AI systems — from detecting latency regressions early to maintaining consistent performance under fluctuating load. If you build or operate systems where delays are unacceptable, this session will provide practical insights to design AI platforms that are not only intelligent, but reliably fast under real-world conditions.... Read more

11:00–12:00

60 min
Aaron Hunter & Saurabh Rob Dahal

1h Hands-On with Kiro: Use AI Agents to Build Your Website and Optimize Your Resume

####Free Kiro Access for Attendees#### Kiro is a new agentic IDE from AWS that turns your ideas into working software through specs, not just prompts. In this hands-on workshop, you'll use Kiro to build a website from scratch using spec-driven development - where requirements, design, and tasks are generated before a single line of code is written. Then you'll use Kiro's agent skills to generate a polished, tailored resume in minutes. No prior Kiro experience needed. Every attendee gets a free Kiro Pro coupon... Read more

12:00

Lunch & networking

Main lobby

13:00

Marcos Novelli Harispe

Beyond Prevention: Mastering Incident Response and Post-Mortems with Ephemeral Environments

LoopStudioWatch
We know Ephemeral Environments (EEs) reduce unexpected errors in production, but what happens when an issue inevitably slips through? In this follow-up to “Making Life Easier for SREs with EEs”, we shift our focus from prevention to the high-pressure world of incident response and post-mortems. This session explores how EEs can help us become faster firefighters and more effective investigators. We will walk through a live Kubernetes microservice demo to experience how EEs serve as a dual-purpose tool when facing real production failures: first, as a firefighting asset to quickly reproduce and mitigate production failures eliminating the “it only breaks in production” problem; and second, as a time machine to rebuild the exact configuration and code state that triggered the failure for deep debugging and root cause analysis leading to permanent fixes. We will also touch upon common concerns, including how to manage the cost and persistence of these temporary environments in a real-world SRE workflow.... Read more

13:30

Zachary Gruenberg

Who Watches the AI Agents? SRE for Non-Human Identity at Scale

Palo Alto NetworksWatch
AI agents (MCP servers, LangChain tools, autonomous coding assistants) are becoming production workloads that SREs own, but nobody’s treating their identities, secrets, or privilege escalation paths with the same rigor as human operators. Walk through real patterns: Conjur-brokered credentials for AI agents, JIT privilege grants scoped to agent tasks, and what an incident looks like when an agent’s session token gets over-permissioned.... Read more

14:00

Christopher Tineo

Free Software Isn't Gratis: why companies should treat Open Source as part of their Infrastructure

Game Plan TechWatch
Maintaining an open source project is hard. It requires managing a group of people who are largely working for free to build something that other people profit off of, usually distributed across the globe, with limited resources. The whole time you’re doing this, you’re receiving demands from users and businesses alike for features or bug fixes on a timeline that works for them, not you and your (possibly very limited) group of contributors that you can’t exactly order around, since they aren’t being paid. It’s stressful, and it can be overwhelming. When one of these projects is the victim of an attack that takes advantage of the fact that there are only one or two maintainers, or eventually has to shut down due to rising technical debt and falling contributor numbers, the public blame falls on us, not on the businesses that didn’t offer contributors in time.... Read more

14:30

Networking & sponsor crawl

Main lobby

15:00

Omari Gaskins Jr

Your Outdated Docs are Costly: Why You should be writing tests for your docs

DocuSORWatch
Internal Developer Platforms aim to create golden paths, reduce cognitive load, and standardize how teams build and operate software. But there’s a blind spot: documentation. READMEs, onboarding guides, and operational runbooks are often treated as static artifacts. Over time they drift from reality, even as the platform itself evolves. The result is slower onboarding, brittle self-service workflows, and increased support burden on platform teams. In this talk, I’ll explore documentation as an unverified surface within Internal Developer Platforms, and introduce a practical approach to making it executable. By treating markdown instructions and runbooks as workflows that can be executed and validated inside CI or ephemeral environments, platform teams can detect drift automatically. Setup steps, service bootstrapping commands, and health checks become verifiable contracts rather than informal guidance. We’ll cover: • Why documentation drift increases cognitive load and support tickets • How to integrate executable documentation into platform pipelines • How this approach strengthens golden paths and reduces onboarding time • What changes in platform team workflows when documentation becomes testable Document drift is a universal problem, and a costly one. It's time we implmented an actual fix.... Read more

15:30

Shahram Anver

Agentic Learning Patterns for SRE

ClericWatch
Most AI agents in production process each incident from scratch. No memory of what worked last time, no feedback loops, no adaptation. Stateless tools pretending to be smart. At Cleric, we build an autonomous AI SRE. Getting the agent to diagnose incidents was the easy part. Getting it to retain and apply what it learned from previous ones is where the real engineering problems lie. When one engineer figures out that an OOM spike is always the Redis sidecar, the agent should know that too. And it should know it across teams, across services, across time. We built a three-layer operational memory architecture — semantic, episodic, and procedural — that enables the agent to retain context across investigations and to improve over time. Semantic memory captures what the agent knows about infrastructure and relationships. Episodic memory records specific investigations and their outcomes. Procedural memory encodes the patterns that worked and when to apply them. But captured knowledge decays. The runbook from six months ago references a service that has since been decomposed into three microservices. The fix that worked in Q3 causes a different failure in Q1 because traffic patterns shifted. This talk covers how we detect and handle staleness, the feedback loops that update agent behavior based on resolution outcomes, and what we've learned about building agents that actually get better at their job over time.... Read more

16:00

Vivek Shah

99.9% Fun: A Game-Based Guide to SLOs and SLAs

Providence
This presentation provides a comprehensive and engaging overview of Service Level Objectives (SLOs) and Service Level Agreements (SLAs), using a scroller game built in with HTML Canvas and Vanilla JS to illustrate the concepts. The three sections of the scroller game cover availability, latency, and error rate. For each metric, the accompanied presentation explains the math behind the metrics in an accessible way. It also reasons why certain percentiles or thresholds may be set based on situation. The talk ends with the exploration of case studies that illustrate how SLOs/SLAs have helped support the core values and product values of companies (ex, through supporting customer-first development and the delivery of high-quality results).... Read more

16:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room mini room
09:00 KeynoteIntro by Mark Pawlikowski
Intro
09:30 KeynoteObservability and SRE in the AI Era
Ajuna Kyaruzi • Datadog
10:00 Coffee break
10:30 Choose Your Weapon: The Performance Testing Arsenal
Heather Thacker • Gatling
When Milliseconds Cost Millions: Designing Low-Latency AI Systems in Fintech
Anisha Manoharan • Bank Of JuliusBaer
11:00 Building a World Model for Production
Shreyas Iyer • Antimetal
1h Hands-On with Kiro: Use AI Agents to Build Your Website and Optimize Your Resume
Aaron Hunter & Saurabh Rob Dahal • AWS
60 min
11:30 Shifting Left - Evolution of the Software Development Life Cycle
Ian Miller • Seismic
12:00 Lunch & networking
13:00 OpenTelemetry: Playtime Is Over
Josh Lee • Altinity
Beyond Prevention: Mastering Incident Response and Post-Mortems with Ephemeral Environments
Marcos Novelli Harispe • LoopStudio
13:30 From Pets To Cattle To Agents: Evolving Identity And Security For Workloads
Dwayne McDaniel • GitGuardian
Who Watches the AI Agents? SRE for Non-Human Identity at Scale
Zachary Gruenberg • Palo Alto Networks
14:00 AI for Product Management: Working Faster, Smarter, and More Strategically
Greg Spektor • Extraterrestrial Incorporated
Free Software Isn't Gratis: why companies should treat Open Source as part of their Infrastructure
Christopher Tineo • Game Plan Tech
14:30 Networking & sponsor crawl
15:00 Perfect Reasoning, Wrong Answer
Willem Pienaar • Cleric
Your Outdated Docs are Costly: Why You should be writing tests for your docs
Omari Gaskins Jr • DocuSOR
15:30 Kubernetes at the Edge: SRE Lessons from Disconnected Clusters
Edward Rodriguez • FUSSMOBILE
Agentic Learning Patterns for SRE
Shahram Anver • Cleric
16:00 Agent Harnesses: From Slot Machines to Safety Nets
Paul Caplan • Teladoc Health
99.9% Fun: A Game-Based Guide to SLOs and SLAs
Vivek Shah • Providence
16:30 Wrap up

Speakers

Aaron Hunter
& Saurabh Rob Dahal
AWS
Ajuna Kyaruzi
Datadog
Anisha Manoharan
Bank Of JuliusBaer
Christopher Tineo
Game Plan Tech
Dwayne McDaniel
GitGuardian
Edward Rodriguez
FUSSMOBILE
Greg Spektor
Extraterrestrial Incorporated
Heather Thacker
Gatling
Ian Miller
Seismic
Intro
Josh Lee
Altinity
Marcos Novelli Harispe
LoopStudio
Omari Gaskins Jr
DocuSOR
Paul Caplan
Teladoc Health
Shahram Anver
Cleric
Shreyas Iyer
Antimetal
Vivek Shah
Providence
Willem Pienaar
Cleric
Zachary Gruenberg
Palo Alto Networks

Venue

Datadog

The New York Times Building
5620 8th Avenue, 45th Floor,
New York, NY 10018, United States

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one