SREday

Site Reliability, DevOps and Cloud

October 2, 2026 San Francisco, CA, USA

1
Day
16+
Speakers
2
Tracks
100+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

AppMaster, Autolake, Cisco, Fidian, FloQast, Google, Grafana Labs, Harness, Hippo Harvest, Hotdata, LangChain, Microsoft, Nexsys Systems LLC, NOFire AI, Red Hat, Reliaburger, Serra Labs, Signadot, Socket, SpaceXAI, StackGen, TestSprite, TruffleRoot, ViralCircle, Wack Incorporated, Walmart Global Tech, WSO2

Topics so far:
Performance & Scalability

Event Starts In:

Tickets

Schedule

October 2, 2026 • 2 parallel tracks • 9AM - 6:30PM • San Francisco, in-person
view as table
main room • Track 1

09:00

Matt Schillerstrom

KeynoteYour Customers Are Already Building Your Roadmap

Harness
What if the best feature request isn't a ticket but something your customer has already built? AI agents are making that possible. Customers can now create specialized workflows on a platform using their own context, tools, and operational knowledge. At Harness, we are seeing this firsthand with Worker Agents. In this talk and demo, we’ll show how customers are extending the platform, what those agents reveal through usage and telemetry, and how that gives product teams a new way to discover what customers need next. The strongest signal is not just what customers ask for. It's what they actually build and continue to use. The feature request of the future may be something the customer has already built.... Read more

09:30

Arjun Iyer

Keynote10x the PRs, Same Error Budget: Validating Agent-Written Code Before Prod

Signadot
Coding agents changed the math on software delivery. Teams that shipped ten PRs a day now see a hundred, and most of that code was never run by a human before it hit CI. The error budget didn't grow with it. Change is still the leading cause of incidents. When change volume goes up an order of magnitude and validation stays flat, one of two things happens. The pipeline backs up, or teams cut corners to clear it. Either way, the risk lands in production and on the people who run it. The usual fixes (more CI runners, more mocks, test in prod) don't hold at this scale. In a microservices system, the failures that reach production are rarely unit-level. Integration failures across services, queues, and data stores are the ones that slip into production, and the only way to catch them before merge is to run the change in an environment against real dependencies. Teams build them two ways today: shared or fully duplicated stacks. I'll show why neither can absorb 10x the change, then walk through a third approach that gives every PR its own isolated environment without duplicating the stack. You'll leave with a framework for what to verify before merge, what to leave to production, and how to keep the error budget intact when agents write most of the code.... Read more

10:00

John Jamie & Dharani Vijayakumar

KeynoteThe Reliability Factory: What 109,000 Incidents Say Comes Next for SRE

StackGen
We analyzed 109,000 incidents from the public status pages of 392 companies, along with hundreds of published post-mortems. We started with resolution times and then went under the hood, classifying each incident by failure mode, root cause and the remediation the team applied. The result is one of the most detailed public pictures of how production systems fail and how teams recover them. It also shows that a team's own response practice explains three times more of its resolution time than its industry does. The data also shows how AI is changing incidents. AI-related incidents grew from under 2% of the total in 2023 to more than 10% in 2026. Incidents caused by AI agents acting on production systems are rising sharply this year. With AI coding pushing change volume higher every quarter, SRE teams will face much more change than current practice was designed to handle. In the second half of the talk, we describe where we believe SRE goes next. Faster incident response on its own won't keep pace. The focus has to shift from mitigation to prevention, and the mitigation that remains has to become autonomous. Reliability becomes part of an operations factory, in which agentic systems check each change against known failure modes before it ships and mitigate what still breaks within the team's SLOs. Agents can only do that with a world model of the systems they run and a safe place to test a fix before applying it. Today's service and context graphs map what is deployed and how it connects. A world model also needs intended state, change history, SLOs, upstream providers and the failure patterns from our research. The SRE becomes a reliability architect, who designs this system and decides what agents are allowed to act on. We'll also propose an autonomy index to track how much operational work agents carry and how well they do it. Attendees will leave with insights from the incident research and a roadmap toward autonomous SRE.... Read more

10:30

Coffee break

Main lobby

11:00

Smita Pasumarthi & Chandan Chilumula

From Observability to Causality: Teaching Machines How Your System Fails

TruffleRoot
Modern observability is very good at telling us what is unhealthy. It is much worse at telling us why. A database can be on fire because it caused an incident—or because something several hops away made it the place where the damage accumulated. An LLM can read all of the telemetry and still make the same mistake. The problem is not always a lack of data or intelligence. Sometimes what is missing is an understanding of how this particular system behaves when things go wrong. While building TruffleRoot, we’ve been exploring a different approach: give machines an explicit, evolving model of an environment—how it is connected, what changes over time, how failures can propagate, and what evidence supports or contradicts an explanation. This talk is about that shift from observability to causality: why generic assumptions about failure break down in real systems, what it means for a system to learn how your system fails, and where deterministic reasoning and LLMs each fit into that picture. The goal isn’t an AI that tells a better story about an outage. It’s a system that learns how your system fails and can show its work.... Read more

11:30

Aron Eidelman

Don't Let Your AI's Patch Break Prod: SRE Patterns for Autonomous Remediation

Google
Software and security teams have never had to patch so much code so quickly. AI agents can generate vulnerability fixes in seconds, but pushing unverified, machine-authored code to production has led to trading a CVE for a production outage. To solve this problem, SRE principles are now gaining new traction in application security and development teams that realize mitigating one risk shouldn't create another. To benefit at all from agents in SSDLC, we need operational cultures that recognize reliability as a combined social and technical measure, rather than just another tooling challenge. In this talk, we’ll walk through the practical mechanics of safely remediating vulnerabilities with agents. We’ll cover how TDD provides the baseline contract to keep patches from breaking existing features, why symptom-based test coverage catches regressions traditional tests miss, and how sandboxed runtimes isolate untrusted agent execution. Finally, we’ll look at closing the loop with automated canary analysis and SLO-driven rollbacks so fixes can deploy without risking production. We'll stick to real scenarios that helped us patch hundreds of repositories without breaking any relying applications (at least as far as users could tell).... Read more

12:00

Artemis Leonardou

The Junior Engineer Is Gone and Your On-Call Has Not Noticed

NOFire AI
The career ladder most teams still hire into does not really describe the people joining them now, and on-call is where that gap shows up first. I started shipping production code at fourteen. I never got the version of this job where someone walks you through syntax and then slowly widens your scope, because by the time I started, that was not how anyone around me was learning. We went from tutorials straight to orchestrating systems with agents. That produces engineers who move unusually fast and who have real gaps in unusual places, and most runbooks and escalation paths assume neither of those things. In fifteen minutes I want to be specific about what that looks like from the inside. Where we are genuinely faster, where we are genuinely worse, and the two things I would change about code review and on-call rotation if you are hiring people like me. You would probably rather hear it from someone currently living it than from a report about us.... Read more

12:30

Lunch & networking

Main lobby

13:30

Wei-Chin Call

From Thumbs-Down to RCA: Closing the Loop on AI Agents

Grafana Labs
Your on-call rotation doesn't page you when an AI agent gives a customer bad advice, but it should. As LLM-powered agents move from chatbot novelty into production systems that make decisions, recommend actions, and interact directly with users, SRE practices built around traditional services need to evolve to account for non-deterministic, natural-language behavior. In this talk, I’ll walk through a live AI agent, an intentionally “broken” education assistant, to demonstrate what agent observability looks like in practice. We’ll follow distributed traces across the LLM call and its downstream tool calls, see an automated evaluator catch a harmful response before it reaches a real audience, and trace a deliberate service outage end-to-end through the agent’s error-handling path. I’ll then close the loop the way an SRE would: a simple thumbs-up/down control on every agent response feeds directly into a conversation-rating signal. A dissatisfied user is no longer just a qualitative complaint. It becomes a queryable, traceable signal that can be correlated with the trace that produced it. Finally, we’ll explore what familiar SRE concepts such as SLOs, error budgets, and incident response begin to look like when a language model becomes part of the production service. Attendees will leave with a practical framework for instrumenting and operating AI agents as production dependencies, and a healthy skepticism toward the idea that AI agents somehow don't need SREs.... Read more

14:00

Geoff White

Open Source SRE Agents for your internal Labs

Nexsys Systems LLC
At 2 a.m. during the Kubernetes buildout for my home agent lab, my SRE agent Gandalf helped me chase down a routing bug quietly breaking federation between the gateway cluster and five agent homeservers behind a VPN tunnel. A few weeks later, Gandalf moved off my laptop and onto a proper server, the same week the swarm outgrew my desk. That's the origin story of Agent-Matrix: a fleet of sovereign AI agents, each with its own identity on the Matrix protocol, letting people message agents from Element or FluffyChat on their phone exactly like messaging a colleague, without routing a single message through Telegram, Signal, or anyone else's chat cloud. This talk is about what it takes for one SRE, not a platform team, to stand up that fleet on a spare server with open protocols and open source tools. You'll meet Gandalf, the SRE agent watching for Grey Rhinos and helping stabilize Black Jellyfish style cascades, and Galadriel, the research agent drawing on OpenBrain, the fleet's shared long-term memory. I'll give an honest accounting of what's solid today (federation, per-agent identity, shared memory) and what's still being built (full end-to-end encryption via a Rust rewrite of the Matrix MCP server), so you can judge what's worth building at your own org, on your own budget.... Read more

14:30

Edward Borukhov

When the On-Call Engineer Is an Agent: Incident Response for AI-Driven Systems

Walmart Global Tech
How AI agents fail differently than traditional software—silently, non-deterministically, and in ways that drift as models update—and practical patterns for incident classification, observability, and postmortems built for probabilistic systems.... Read more

15:00

Robbie McKinstry

Overcoming the Trust Deficit with SRE Agents

Wack Incorporated
Most SRE agents only do half of the job: they wait for an incident to occur, then hand off diagnostics. While that’s better than nothing, it stops short of resolving the incident. The flipside is that no one actually wants an agent to autonomously resolve incidents if it can only deliver 90% of the time. That would mean for the other 10%, the agent has made the wrong call on business-impacting systems, exacerbating conditions and making a postmortem more difficult to conduct. While *coding* agents are allowed to act autonomously since bad output can be caught in review, incident response doesn’t have that luxury, leaving SRE agents to watch instead of act. This talk covers the limitations of bleeding-edge incident response agents, and provides a roadmap toward trustworthy, verified, and independent agent playbooks. Playbooks require agents to show their work to an independent oracle that judges whether the case they’ve made is plausible or flimsy. Only once a check passes does the agent move to the next step in the playbook, with each step gating the tools it's allowed to touch. What you end up with is your existing runbook turned into a state machine, where every transition locks down what the agent can do next. I'll walk through the architecture, what actually counts as evidence to a judge, as well as the weaknesses of this approach.... Read more

15:30

Networking & sponsor crawl

Main lobby

16:00

Yunhao Jiao

Verification After the Deploy: Catching What AI-Written Code Breaks Before Your Users Do

TestSprite
Most of the quality tooling we have added for AI-written code looks at the diff. Code review, static analysis, and pre-merge checks all ask whether the change looks correct. None of them answer the question a reliability engineer actually cares about, which is whether the running application still works. When a person wrote the code, that mattered less. They knew what they had changed and what else it might touch. When an agent writes the code, nobody knows. The pull request looks clean, the checks pass, and the first honest signal arrives after the deploy, often when a user hits it. This talk makes the case for moving verification to the other side of the deploy event, and shows what that looks like in practice. A testing agent opens the live build in a real browser, exercises it the way a user would, and hands the failure back to the coding agent with enough context to act on it: the failing step, a screenshot, a DOM snapshot, and a root-cause hypothesis. The repair happens before a human is involved. Yunhao will close with what TestSprite measured across CoderCup, a public benchmark where frontier coding agents built the same application under identical conditions, including a pattern the team calls feature retention: working features quietly disappearing as later edits land in a codebase whose author does not remember what it built last week.... Read more

16:30

Hilliary Lipsig

Burning the Ships: Migrating Our Core Logging Stack with No Way Back

Red Hat
What happens when your core logging framework has a critical exploit, rollback is a security risk, and you can’t test the new stack until it hits production? You burn the ships. In this talk, I share the high-stakes engineering story of a complete, one-way logging migration across thousands of live clusters. Forced by an active vulnerability to bypass traditional safety nets we had to rely on napkin-math resource estimations, synthetic traffic testing, table top exercises of what could go wrong, and a terrifyingly literal "test in production" rollout. I will share the exact strategies we used to pull off this zero-rollback cutover without plunging our systems into telemetry blindness. You’ll leave with concrete patterns for calculating production risk under extreme constraints, validating data pipelines on the fly, and failing forward.... Read more

17:00

Jonathon Klobucar

New Incident, Who Dis? Incident Response Beyond Engineering

Socket
SRE already solved incident response: severities, a commander, a working channel, a timeline, a blameless postmortem. Then we kept it inside engineering. Security has its own process, but it's usually less battle-tested; IT runs incidents out of a ticket queue; marketing runs them out of a group text. And the worst incidents are the ones that hit all three at once. That process is the most valuable thing an SRE team owns, and it's also the easiest way to stop being the team that says no and start being the team everyone calls. This talk covers what I learned bringing the SRE incident model to security, IT, and a PR crisis: where it worked without changes, where it fell apart, and how to show up as a partner instead of a process cop. We'll also cover where AI genuinely helps and how it hands teams with no incident experience a polished process they can't actually run. You'll leave with a severity model every department can read, a handoff checklist, and a DR/IR tabletop format the whole company practices together and, for once, actually enjoys.... Read more

17:30

Oleg Sotnikov

One Container, One Database per Tenant: Running Hundreds of Thousands of AI-Generated Apps

AppMaster
When your platform generates your customers' applications, every incident in their production is your incident. This talk is a field report from running a no-code platform's runtime plane at a peak of 250,000 concurrently hosted containers: one container and one PostgreSQL database per tenant app across multiple regions, with a control plane and a build plane feeding it. The applications themselves are produced by an AI-driven generator from customers' blueprints, which adds a failure class most fleets never see — code that nobody reviewed, regenerated on demand. I'll cover the reliability model we settled on and why: blast-radius isolation vs. density, per-tenant observability, fleet-wide regeneration and redeploy when the generator itself changes, noisy-neighbor handling on shared hosts, and the failure modes that only exist when the code in production was compiled from a customer's blueprint 10 minutes ago. Practical patterns for anyone operating a fleet of workloads you don't control.... Read more

18:00

Romain Priour

Operating Reliable Software Across Customer Clouds

LangChain
LangSmith’s Bring Your Own Cloud (BYOC) model lets customers run LangSmith in their AWS accounts across many regions. Our team is responsible for provisioning the infrastructure, shipping upgrades, monitoring deployments, and proactively resolving issues. As the fleet grows, we need to keep doing that reliably without a proportional increase in manual work. In this talk, I’ll show how we operate this fleet efficiently while keeping operational overhead low. I'll cover how we provision and continuously reconcile infrastructure, how we manage private customer environments through a shared control plane, and how we build internal tooling to keep day-to-day operations manageable as the fleet grows. We’ll follow a release through our deployment pipeline: validating changes, rolling them out across customer environments, designing upgrades to avoid downtime, and handling failures along the way. Finally, we’ll look at how we leverage telemetry to understand what’s happening across the fleet, surface failures that require attention, and use internal AI tooling to help troubleshoot incidents.... Read more

18:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
meeting room • Track 2

10:30

Coffee break

Main lobby

11:00

Vanshika Gupta

How to Turn LinkedIn Into Your #1 B2B GTM Channel

ViralCircle
Most B2B companies treat LinkedIn as a content channel. This talk breaks down how to turn it into a GTM engine, including the content formats that work (and the ones that don’t), how to build the right audience, and how to turn attention into conversations, pipeline, and customers.... Read more

11:30

Hubert Chen

Beyond Unit Tests: Bringing Intelligent Agent Evals into CI/CD Pipelines

Fidian
As software shifts from deterministic code to non-deterministic LLMs and autonomous agents, traditional CI/CD practices fall short. Treating probabilistic agentic workflows like rigid software forces tough choices around test coverage, pass rates, and build times. This talk explores how to modernize DevOps pipelines for the AI era. We'll examine how balancing mode l selection, evaluation costs, and pipeline speed affects both developer velocity and system reliability. We will dive into improvement strategies like intelligent task selection, lightwei ght judging, test tiering and eval prediction can reduce evaluation time and costs.... Read more

12:00

Ganesh Nathan

Your Bill Is an SLO: Reliability for a Data Platform With No Servers

Autolake
We run a multi-tenant data lakehouse for higher education and public sector organizations on AWS serverless services only: S3, Glue, Athena, Lambda, Step Functions, and Iceberg tables. No cluster, no Kubernetes, and no on-call rotation. That changes what reliability means, because the signals SREs usually rely on, CPU, memory, pod health, do not exist. Our answer is that cost is the most honest signal you have. A query that returns the right answer at ten times the expected cost is a degradation, and we treat it like one. This talk walks through the SLOs we built around cost per query, file count per partition, and snapshot age, and how those SLOs drive automated remediation: compaction, snapshot expiry, catalog reconciliation, and retention enforcement that run themselves and report back. We will also cover the guardrails that keep petabyte-scale storage healthy without a human in the loop: hidden partitioning so users cannot write expensive queries by accident, and per-tenant cost limits that stop a runaway workload before it becomes a bill. You will leave with a working model for SRE on serverless data infrastructure and a checklist you can apply to any Iceberg on S3 deployment.... Read more

12:30

Lunch & networking

Main lobby

13:30

Mark Pawlikowski

Kill It and Watch It Come Back: Hands-on with Reliaburger

Reliaburger
I have never run a Kubernetes cluster. Yes, I said it. Are you still with me? Cool! Over the past last five years I talked to thousands of SREs and platform engineers at the conferences I run with my team: [SREday](https://sreday.com/), [LLMday](https://llmday.com/), [PLATFORMday](https://platformday.com/) and [Conf42](https://www.conf42.com/). The same story keeps coming back. Before the first container hits production you step into the world of Kubernetes, and then you need to build a whole team just to keep it running. A kolljoy for many smaller teams. We started talking with my brother, [Miko Pawlikowski](https://www.linkedin.com/in/mikolajpawlikowski/), how we could help the community ease some of these pains. Miko has been running platform teams on Kubernetes for 10+ years, leading him to start writing Reliaburger: one Rust binary with the scheduler, gossip, Raft, eBPF service discovery, ingress, mTLS, registry, metrics, logs, GitOps and chaos testing inside, driven by TOML. An app with health checks, ingress and autoscaling is twelve lines. He wrote the software. I am here to hand it to the community and see what you do with it. This is a hands-on session, so bring a laptop (macOS on Apple silicon or Linux; Windows is not supported yet). In 30 minutes you will: 1. Install Reliaburger with one command and start a node on your laptop. No container runtime, no root, no VMs: the built-in process runtime runs plain binaries. 2. Deploy an app from a twelve-line TOML file and watch it in the terminal dashboard and the web UI. 3. Ship a deliberately sick version and watch the health checker restart it, then roll back with one command. 4. Break it on purpose with the built-in fault injector: kill an instance, freeze another, and let `relish wtf` explain what just happened. 5. Read the manual and the source code from inside the binary, so you can keep going after the session. 6. Tell us what broke and what you liked. Reliaburger is free and open source (Apache 2.0), and it is version 0.1.0, so it will break. That is where you come in. We are looking for helpers, burger ambassadors and contributors: people who will run it, break it, file the issues, write the docs and tell their teams. There are millions of practitioners running parts of Kubernetes they do not need, with more people than they should need to run it. We want this to become a community-driven project, just as our conferences are. Everyone who makes it through the six steps and gives feedback goes into a raffle at the end. There is hardware on the table, and it is worth your 30 minutes.... Read more

14:00

Eddie Tejeda

The Agent Data Bus: Shared State for Multi-Agent Systems

Hotdata
Multi-agent AI systems are distributed systems, but many still pass data through prompts. This loses information, wastes tokens, and makes failures hard to manage. This talk presents a shared data bus where agents exchange a database ID. Using a real three-agent pipeline, I’ll cover retries, conflicting writes, access control, monitoring, and cleanup. I will also provide a guide for choosing between prompts, queues, files, and databases.... Read more

14:30

Larsen Cundric

Running Untrusted Agents in Production

SpaceXAI
Giving an AI agent the ability to execute code changes the infrastructure it needs. When I was working at Browser Use, that meant rethinking where agents ran, which credentials they could access, and how their work survived a process dying. I'll walk through the architecture we built: isolated agent runtimes, scoped access through a control plane, and state stored outside the worker. Then I'll cover what still went wrong, including a failed state restore that was mistaken for a fresh session and allowed blank state to replace valid conversation history. We'll look at the recovery changes that followed and the operational tradeoffs we encountered moving between runtimes. The talk draws on publicly documented Browser Use work. Attendees will leave with practical lessons for isolating untrusted workloads, distinguishing missing state from unavailable storage, and deciding which parts of an agent platform they need to operate themselves.... Read more

15:00

Jagan Jagannathan

Different Workloads, Different Wins

Serra Labs
For some workloads, success means a smaller cloud bill; for others, it means better performance. This talk demonstrates how to put those goals into practice. We’ll walk through analyzing workload behavior, selecting an optimization mode, and evaluating recommended configurations against cost and performance requirements. Attendees will see how to turn workload data into actionable infrastructure decisions.... Read more

15:30

Networking & sponsor crawl

Main lobby

16:00

Ankush Sharma

SRE for LLMs: Observability and Chaos Engineering for AI

FloQast
LLMs break the assumptions traditional monitoring is built on, a service can be "up" while quietly hallucinating, drifting, or burning your token budget. This talk shows how to apply SRE discipline to AI systems i.e defining SLOs for inference quality and latency, token-level metrics, prompt tracing, and structured logging across LLM pipelines. We'll then go further with chaos engineering for AI, injecting failures into model-serving infrastructure and automating resilience experiments. Attendees will leave with a practical blueprint for making LLM systems as observable and reliable as the rest of their stack, drawn from the author's book *Observability for Large Language Models* (Apress, 2026).... Read more

16:30

Sameera Jayasoma

Beyond MCP and Skills: Architecting Agentic Developer Workflows

WSO2
MCP servers, Skills, and agents are a commodity now. But composing these primitives into production-ready developer workflows is not easy. Building the abstraction layers that help agents get the context they need, without burning tokens on low-level detail and wiring it all into a cohesive system architecture takes time. Once you've figured out the patterns, they're repeatable. This talk is about those patterns, the decisions and trade-offs you face when turning these building blocks into developer workflows. I'll demonstrate one workflow live to make it concrete: an alert fires, an agent assembles context, produces a root-cause report, helps a developer troubleshoot, and proposes a fix the platform executes. From that one workflow, I'll pull out patterns you can reuse anywhere, treating the platform as the agent's context source, letting agents express intent while the platform reconciles deterministically, scoping permissions to bound blast radius, and placing the human handoffs that keep people in the loop. These are repeatable patterns I learned building OpenChoreo, a CNCF project and they apply to any platform, not just this one.... Read more

17:00

Girish Konda

A Thousand Alerts a Day, and One Real Incident Hiding In Them

Microsoft
A monitoring stack that fires a thousand alerts a day does not have a monitoring problem. It has a reading problem. Nearly all of it is benign and repeats the same twenty shapes - and somewhere in there is a real incident that looks exactly like the rest until someone digs. The obvious fix is to point an agent at the queue and let it triage. We tried that. It is confidently wrong, and the confidence is the dangerous part: it returns a fluent, well-argued verdict on every alert at the same level of certainty, whether it pulled real evidence or just pattern-matched the title. At a thousand a day, even a small confidently-wrong rate is dozens of silent misfilings - and the ones it buries are the unfamiliar ones, which is exactly where the real incidents live. I am a Principal Engineer in Microsoft Fabric, where I am the tech lead and architect for capacity management. What finally worked was refusing to let one agent do both jobs. Triage got cut back until routing is all it can do. It groups by signature and hands off. It is not allowed to conclude, to close, or to declare anything benign - that is the judgement it is worst at and the one with the worst downside. Investigation was then split across dedicated agents, one per failure domain. Each knows only its own telemetry, its own known-benign patterns, and its own escalation bar. That narrow hypothesis space is what makes "I could not find evidence for this" an answer the agent will actually give, instead of a plausible story. It also means each one can be tested on its own against past incidents, which a single generalist triager never can be. The metric that matters is recall on the rare real incident, not how much of the queue got closed. I will go through the routing design, how the investigation agents are kept honest, what is escalated to a human on purpose, and the real problems we only caught because a specialist refused to answer.... Read more

17:30

An Phan

When the Site Is Actually a Site: Rethinking Site Reliability for the Physical World

Hippo Harvest
What does Site Reliability Engineering look like when the "site" is an actual physical place? At a physical site, software is only one part of the system. Networks can disappear, machines have their own state, sensors provide incomplete evidence, and environmental conditions keep changing regardless of whether our infrastructure is healthy. A reliable cloud service does not necessarily mean a reliable site. Drawing from engineering systems that span cloud infrastructure, on-site computing, machines, sensors, and physical processes, this talk explores where the traditional reliability boundary starts to break down. It proposes a broader way to think about observability, failure, recovery, and system health when the physical site itself becomes part of the system we are responsible for keeping reliable.... Read more

18:00

Hazel Mahajan

Chaos Engineering for AI Agents: What Happens When Your Tools Fail, Time Out, or Vanish

Cisco
Chaos engineering has been standard practice in distributed systems for over a decade - kill a node, inject latency, watch what breaks, fix it before production does. Agentic AI systems now chain LLM calls to external tools (via MCP or similar), but there's no equivalent discipline yet for what happens when those tools misbehave. In this 30-minute session, I'll apply classic chaos-engineering principles - fault injection, steady-state hypothesis, blast-radius containment - to an actual agent pipeline: tool calls that time out, return malformed data, or silently disappear mid-chain. I'll walk through what broke loudly vs. what failed silently and went undetected, share the one number from testing that surprised me most, and leave the audience with a minimal "chaos toolkit for agents" checklist they can apply to their own systems.... Read more

18:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room meeting room
09:00 KeynoteYour Customers Are Already Building Your Roadmap
Matt Schillerstrom • Harness
09:30 Keynote10x the PRs, Same Error Budget: Validating Agent-Written Code Before Prod
Arjun Iyer • Signadot
10:00 KeynoteThe Reliability Factory: What 109,000 Incidents Say Comes Next for SRE
John Jamie & Dharani Vijayakumar • StackGen
10:30 Coffee break
11:00 From Observability to Causality: Teaching Machines How Your System Fails
Smita Pasumarthi & Chandan Chilumula • TruffleRoot
How to Turn LinkedIn Into Your #1 B2B GTM Channel
Vanshika Gupta • ViralCircle
11:30 Don't Let Your AI's Patch Break Prod: SRE Patterns for Autonomous Remediation
Aron Eidelman • Google
Beyond Unit Tests: Bringing Intelligent Agent Evals into CI/CD Pipelines
Hubert Chen • Fidian
12:00 The Junior Engineer Is Gone and Your On-Call Has Not Noticed
Artemis Leonardou • NOFire AI
Your Bill Is an SLO: Reliability for a Data Platform With No Servers
Ganesh Nathan • Autolake
12:30 Lunch & networking
13:30 From Thumbs-Down to RCA: Closing the Loop on AI Agents
Wei-Chin Call • Grafana Labs
Kill It and Watch It Come Back: Hands-on with Reliaburger
Mark Pawlikowski • Reliaburger
14:00 Open Source SRE Agents for your internal Labs
Geoff White • Nexsys Systems LLC
The Agent Data Bus: Shared State for Multi-Agent Systems
Eddie Tejeda • Hotdata
14:30 When the On-Call Engineer Is an Agent: Incident Response for AI-Driven Systems
Edward Borukhov • Walmart Global Tech
Running Untrusted Agents in Production
Larsen Cundric • SpaceXAI
15:00 Overcoming the Trust Deficit with SRE Agents
Robbie McKinstry • Wack Incorporated
Different Workloads, Different Wins
Jagan Jagannathan • Serra Labs
15:30 Networking & sponsor crawl
16:00 Verification After the Deploy: Catching What AI-Written Code Breaks Before Your Users Do
Yunhao Jiao • TestSprite
SRE for LLMs: Observability and Chaos Engineering for AI
Ankush Sharma • FloQast
16:30 Burning the Ships: Migrating Our Core Logging Stack with No Way Back
Hilliary Lipsig • Red Hat
Beyond MCP and Skills: Architecting Agentic Developer Workflows
Sameera Jayasoma • WSO2
17:00 New Incident, Who Dis? Incident Response Beyond Engineering
Jonathon Klobucar • Socket
A Thousand Alerts a Day, and One Real Incident Hiding In Them
Girish Konda • Microsoft
17:30 One Container, One Database per Tenant: Running Hundreds of Thousands of AI-Generated Apps
Oleg Sotnikov • AppMaster
When the Site Is Actually a Site: Rethinking Site Reliability for the Physical World
An Phan • Hippo Harvest
18:00 Operating Reliable Software Across Customer Clouds
Romain Priour • LangChain
Chaos Engineering for AI Agents: What Happens When Your Tools Fail, Time Out, or Vanish
Hazel Mahajan • Cisco
18:30 Wrap up

Speakers

An Phan
Hippo Harvest
Ankush Sharma
FloQast
Arjun Iyer
Signadot
Aron Eidelman
Google
Artemis Leonardou
NOFire AI
Eddie Tejeda
Hotdata
Edward Borukhov
Walmart Global Tech
Ganesh Nathan
Autolake
Geoff White
Nexsys Systems LLC
Girish Konda
Microsoft
Hazel Mahajan
Cisco
Hilliary Lipsig
Red Hat
Hubert Chen
Fidian
Jagan Jagannathan
Serra Labs
John Jamie
& Dharani Vijayakumar
StackGen
Jonathon Klobucar
Socket
Larsen Cundric
SpaceXAI
Mark Pawlikowski
Reliaburger
Matt Schillerstrom
Harness
Oleg Sotnikov
AppMaster
Robbie McKinstry
Wack Incorporated
Romain Priour
LangChain
Sameera Jayasoma
WSO2
Smita Pasumarthi
& Chandan Chilumula
TruffleRoot
Vanshika Gupta
ViralCircle
Wei-Chin Call
Grafana Labs
Yunhao Jiao
TestSprite

Venue

Harness.io HQ

55 Stockton St, San Francisco,
CA 94108, United States

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one