SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.
Companies presenting:
AppMaster, Autolake, Cisco, Fidian, FloQast, Google, Grafana Labs, Harness, Hippo Harvest, Hotdata, LangChain, Microsoft, Nexsys Systems LLC, NOFire AI, Red Hat, Reliaburger, Serra Labs, Signadot, Socket, SpaceXAI, StackGen, TestSprite, TruffleRoot, ViralCircle, Wack Incorporated, Walmart Global Tech, WSO2
What if the best feature request isn't a ticket but something your customer has already built?
AI agents are making that possible. Customers can now create specialized workflows on a platform using their own context, tools, and operational knowledge. At Harness, we are seeing this firsthand with Worker Agents.
In this talk and demo, we’ll show how customers are extending the platform, what those agents reveal through usage and telemetry, and how that gives product teams a new way to discover what customers need next.
The strongest signal is not just what customers ask for. It's what they actually build and continue to use. The feature request of the future may be something the customer has already built.... Read more
Coding agents changed the math on software delivery. Teams that shipped ten PRs a day now see a hundred, and most of that code was never run by a human before it hit CI. The error budget didn't grow with it.
Change is still the leading cause of incidents. When change volume goes up an order of magnitude and validation stays flat, one of two things happens. The pipeline backs up, or teams cut corners to clear it. Either way, the risk lands in production and on the people who run it.
The usual fixes (more CI runners, more mocks, test in prod) don't hold at this scale. In a microservices system, the failures that reach production are rarely unit-level. Integration failures across services, queues, and data stores are the ones that slip into production, and the only way to catch them before merge is to run the change in an environment against real dependencies.
Teams build them two ways today: shared or fully duplicated stacks. I'll show why neither can absorb 10x the change, then walk through a third approach that gives every PR its own isolated environment without duplicating the stack.
You'll leave with a framework for what to verify before merge, what to leave to production, and how to keep the error budget intact when agents write most of the code.... Read more
We analyzed 109,000 incidents from the public status pages of 392 companies, along with hundreds of published post-mortems. We started with resolution times and then went under the hood, classifying each incident by failure mode, root cause and the remediation the team applied. The result is one of the most detailed public pictures of how production systems fail and how teams recover them. It also shows that a team's own response practice explains three times more of its resolution time than its industry does.
The data also shows how AI is changing incidents. AI-related incidents grew from under 2% of the total in 2023 to more than 10% in 2026. Incidents caused by AI agents acting on production systems are rising sharply this year. With AI coding pushing change volume higher every quarter, SRE teams will face much more change than current practice was designed to handle.
In the second half of the talk, we describe where we believe SRE goes next. Faster incident response on its own won't keep pace. The focus has to shift from mitigation to prevention, and the mitigation that remains has to become autonomous. Reliability becomes part of an operations factory, in which agentic systems check each change against known failure modes before it ships and mitigate what still breaks within the team's SLOs.
Agents can only do that with a world model of the systems they run and a safe place to test a fix before applying it. Today's service and context graphs map what is deployed and how it connects. A world model also needs intended state, change history, SLOs, upstream providers and the failure patterns from our research. The SRE becomes a reliability architect, who designs this system and decides what agents are allowed to act on. We'll also propose an autonomy index to track how much operational work agents carry and how well they do it.
Attendees will leave with insights from the incident research and a roadmap toward autonomous SRE.... Read more
Modern observability is very good at telling us what is unhealthy. It is much worse at telling us why.
A database can be on fire because it caused an incident—or because something several hops away made it the place where the damage accumulated. An LLM can read all of the telemetry and still make the same mistake. The problem is not always a lack of data or intelligence. Sometimes what is missing is an understanding of how this particular system behaves when things go wrong.
While building TruffleRoot, we’ve been exploring a different approach: give machines an explicit, evolving model of an environment—how it is connected, what changes over time, how failures can propagate, and what evidence supports or contradicts an explanation.
This talk is about that shift from observability to causality: why generic assumptions about failure break down in real systems, what it means for a system to learn how your system fails, and where deterministic reasoning and LLMs each fit into that picture.
The goal isn’t an AI that tells a better story about an outage. It’s a system that learns how your system fails and can show its work.... Read more
Software and security teams have never had to patch so much code so quickly. AI agents can generate vulnerability fixes in seconds, but pushing unverified, machine-authored code to production has led to trading a CVE for a production outage.
To solve this problem, SRE principles are now gaining new traction in application security and development teams that realize mitigating one risk shouldn't create another. To benefit at all from agents in SSDLC, we need operational cultures that recognize reliability as a combined social and technical measure, rather than just another tooling challenge.
In this talk, we’ll walk through the practical mechanics of safely remediating vulnerabilities with agents. We’ll cover how TDD provides the baseline contract to keep patches from breaking existing features, why symptom-based test coverage catches regressions traditional tests miss, and how sandboxed runtimes isolate untrusted agent execution. Finally, we’ll look at closing the loop with automated canary analysis and SLO-driven rollbacks so fixes can deploy without risking production.
We'll stick to real scenarios that helped us patch hundreds of repositories without breaking any relying applications (at least as far as users could tell).... Read more
The career ladder most teams still hire into does not really describe the people joining them now, and on-call is where that gap shows up first. I started shipping production code at fourteen. I never got the version of this job where someone walks you through syntax and then slowly widens your scope, because by the time I started, that was not how anyone around me was learning. We went from tutorials straight to orchestrating systems with agents. That produces engineers who move unusually fast and who have real gaps in unusual places, and most runbooks and escalation paths assume neither of those things. In fifteen minutes I want to be specific about what that looks like from the inside. Where we are genuinely faster, where we are genuinely worse, and the two things I would change about code review and on-call rotation if you are hiring people like me. You would probably rather hear it from someone currently living it than from a report about us.... Read more
Your on-call rotation doesn't page you when an AI agent gives a customer bad advice, but it should.
As LLM-powered agents move from chatbot novelty into production systems that make decisions, recommend actions, and interact directly with users, SRE practices built around traditional services need to evolve to account for non-deterministic, natural-language behavior.
In this talk, I’ll walk through a live AI agent, an intentionally “broken” education assistant, to demonstrate what agent observability looks like in practice. We’ll follow distributed traces across the LLM call and its downstream tool calls, see an automated evaluator catch a harmful response before it reaches a real audience, and trace a deliberate service outage end-to-end through the agent’s error-handling path.
I’ll then close the loop the way an SRE would: a simple thumbs-up/down control on every agent response feeds directly into a conversation-rating signal. A dissatisfied user is no longer just a qualitative complaint. It becomes a queryable, traceable signal that can be correlated with the trace that produced it.
Finally, we’ll explore what familiar SRE concepts such as SLOs, error budgets, and incident response begin to look like when a language model becomes part of the production service.
Attendees will leave with a practical framework for instrumenting and operating AI agents as production dependencies, and a healthy skepticism toward the idea that AI agents somehow don't need SREs.... Read more
At 2 a.m. during the Kubernetes buildout for my home agent lab, my SRE agent Gandalf helped me chase down a routing bug quietly breaking federation between the gateway cluster and five agent homeservers behind a VPN tunnel. A few weeks later, Gandalf moved off my laptop and onto a proper server, the same week the swarm outgrew my desk. That's the origin story of Agent-Matrix: a fleet of sovereign AI agents, each with its own identity on the Matrix protocol, letting people message agents from Element or FluffyChat on their phone exactly like messaging a colleague, without routing a single message through Telegram, Signal, or anyone else's chat cloud.
This talk is about what it takes for one SRE, not a platform team, to stand up that fleet on a spare server with open protocols and open source tools. You'll meet Gandalf, the SRE agent watching for Grey Rhinos and helping stabilize Black Jellyfish style cascades, and Galadriel, the research agent drawing on OpenBrain, the fleet's shared long-term memory. I'll give an honest accounting of what's solid today (federation, per-agent identity, shared memory) and what's still being built (full end-to-end encryption via a Rust rewrite of the Matrix MCP server), so you can judge what's worth building at your own org, on your own budget.... Read more
How AI agents fail differently than traditional software—silently, non-deterministically, and in ways that drift as models update—and practical patterns for incident classification, observability, and postmortems built for probabilistic systems.... Read more
Most SRE agents only do half of the job: they wait for an incident to occur, then hand off diagnostics. While that’s better than nothing, it stops short of resolving the incident. The flipside is that no one actually wants an agent to autonomously resolve incidents if it can only deliver 90% of the time. That would mean for the other 10%, the agent has made the wrong call on business-impacting systems, exacerbating conditions and making a postmortem more difficult to conduct. While *coding* agents are allowed to act autonomously since bad output can be caught in review, incident response doesn’t have that luxury, leaving SRE agents to watch instead of act. This talk covers the limitations of bleeding-edge incident response agents, and provides a roadmap toward trustworthy, verified, and independent agent playbooks. Playbooks require agents to show their work to an independent oracle that judges whether the case they’ve made is plausible or flimsy. Only once a check passes does the agent move to the next step in the playbook, with each step gating the tools it's allowed to touch. What you end up with is your existing runbook turned into a state machine, where every transition locks down what the agent can do next. I'll walk through the architecture, what actually counts as evidence to a judge, as well as the weaknesses of this approach.... Read more
Most of the quality tooling we have added for AI-written code looks at the diff. Code review, static analysis, and pre-merge checks all ask whether the change looks correct. None of them answer the question a reliability engineer actually cares about, which is whether the running application still works.
When a person wrote the code, that mattered less. They knew what they had changed and what else it might touch. When an agent writes the code, nobody knows. The pull request looks clean, the checks pass, and the first honest signal arrives after the deploy, often when a user hits it.
This talk makes the case for moving verification to the other side of the deploy event, and shows what that looks like in practice. A testing agent opens the live build in a real browser, exercises it the way a user would, and hands the failure back to the coding agent with enough context to act on it: the failing step, a screenshot, a DOM snapshot, and a root-cause hypothesis. The repair happens before a human is involved.
Yunhao will close with what TestSprite measured across CoderCup, a public benchmark where frontier coding agents built the same application under identical conditions, including a pattern the team calls feature retention: working features quietly disappearing as later edits land in a codebase whose author does not remember what it built last week.... Read more
What happens when your core logging framework has a critical exploit, rollback is a security risk, and you can’t test the new stack until it hits production? You burn the ships. In this talk, I share the high-stakes engineering story of a complete, one-way logging migration across thousands of live clusters. Forced by an active vulnerability to bypass traditional safety nets we had to rely on napkin-math resource estimations, synthetic traffic testing, table top exercises of what could go wrong, and a terrifyingly literal "test in production" rollout.
I will share the exact strategies we used to pull off this zero-rollback cutover without plunging our systems into telemetry blindness. You’ll leave with concrete patterns for calculating production risk under extreme constraints, validating data pipelines on the fly, and failing forward.... Read more
SRE already solved incident response: severities, a commander, a working channel, a timeline, a blameless postmortem. Then we kept it inside engineering. Security has its own process, but it's usually less battle-tested; IT runs incidents out of a ticket queue; marketing runs them out of a group text. And the worst incidents are the ones that hit all three at once.
That process is the most valuable thing an SRE team owns, and it's also the easiest way to stop being the team that says no and start being the team everyone calls. This talk covers what I learned bringing the SRE incident model to security, IT, and a PR crisis: where it worked without changes, where it fell apart, and how to show up as a partner instead of a process cop. We'll also cover where AI genuinely helps and how it hands teams with no incident experience a polished process they can't actually run. You'll leave with a severity model every department can read, a handoff checklist, and a DR/IR tabletop format the whole company practices together and, for once, actually enjoys.... Read more
When your platform generates your customers' applications, every incident in their production is your incident. This talk is a field report from running a no-code platform's runtime plane at a peak of 250,000 concurrently hosted containers: one container and one PostgreSQL database per tenant app across multiple regions, with a control plane and a build plane feeding it. The applications themselves are produced by an AI-driven generator from customers' blueprints, which adds a failure class most fleets never see — code that nobody reviewed, regenerated on demand. I'll cover the reliability model we settled on and why: blast-radius isolation vs. density, per-tenant observability, fleet-wide regeneration and redeploy when the generator itself changes, noisy-neighbor handling on shared hosts, and the failure modes that only exist when the code in production was compiled from a customer's blueprint 10 minutes ago. Practical patterns for anyone operating a fleet of workloads you don't control.... Read more
LangSmith’s Bring Your Own Cloud (BYOC) model lets customers run LangSmith in their AWS accounts across many regions. Our team is responsible for provisioning the infrastructure, shipping upgrades, monitoring deployments, and proactively resolving issues. As the fleet grows, we need to keep doing that reliably without a proportional increase in manual work.
In this talk, I’ll show how we operate this fleet efficiently while keeping operational overhead low. I'll cover how we provision and continuously reconcile infrastructure, how we manage private customer environments through a shared control plane, and how we build internal tooling to keep day-to-day operations manageable as the fleet grows.
We’ll follow a release through our deployment pipeline: validating changes, rolling them out across customer environments, designing upgrades to avoid downtime, and handling failures along the way. Finally, we’ll look at how we leverage telemetry to understand what’s happening across the fleet, surface failures that require attention, and use internal AI tooling to help troubleshoot incidents.... Read more
18:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
Most B2B companies treat LinkedIn as a content channel. This talk breaks down how to turn it into a GTM engine, including the content formats that work (and the ones that don’t), how to build the right audience, and how to turn attention into conversations, pipeline, and customers.... Read more
As software shifts from deterministic code to non-deterministic LLMs and autonomous agents, traditional CI/CD practices fall short. Treating probabilistic agentic workflows like
rigid software forces tough choices around test coverage, pass rates, and build times. This talk explores how to modernize DevOps pipelines for the AI era. We'll examine how balancing mode
l selection, evaluation costs, and pipeline speed affects both developer velocity and system reliability. We will dive into improvement strategies like intelligent task selection, lightwei
ght judging, test tiering and eval prediction can reduce evaluation time and costs.... Read more
We run a multi-tenant data lakehouse for higher education and public sector organizations on AWS serverless services only: S3, Glue, Athena, Lambda, Step Functions, and Iceberg tables. No cluster, no Kubernetes, and no on-call rotation. That changes what reliability means, because the signals SREs usually rely on, CPU, memory, pod health, do not exist. Our answer is that cost is the most honest signal you have. A query that returns the right answer at ten times the expected cost is a degradation, and we treat it like one. This talk walks through the SLOs we built around cost per query, file count per partition, and snapshot age, and how those SLOs drive automated remediation: compaction, snapshot expiry, catalog reconciliation, and retention enforcement that run themselves and report back. We will also cover the guardrails that keep petabyte-scale storage healthy without a human in the loop: hidden partitioning so users cannot write expensive queries by accident, and per-tenant cost limits that stop a runaway workload before it becomes a bill. You will leave with a working model for SRE on serverless data infrastructure and a checklist you can apply to any Iceberg on S3 deployment.... Read more
I have never run a Kubernetes cluster. Yes, I said it.
Are you still with me? Cool!
Over the past last five years I talked to thousands of SREs and platform engineers at the conferences I run with my team: [SREday](https://sreday.com/), [LLMday](https://llmday.com/), [PLATFORMday](https://platformday.com/) and [Conf42](https://www.conf42.com/).
The same story keeps coming back. Before the first container hits production you step into the world of Kubernetes, and then you need to build a whole team just to keep it running. A kolljoy for many smaller teams. We started talking with my brother, [Miko Pawlikowski](https://www.linkedin.com/in/mikolajpawlikowski/), how we could help the community ease some of these pains.
Miko has been running platform teams on Kubernetes for 10+ years, leading him to start writing Reliaburger: one Rust binary with the scheduler, gossip, Raft, eBPF service discovery, ingress, mTLS, registry, metrics, logs, GitOps and chaos testing inside, driven by TOML. An app with health checks, ingress and autoscaling is twelve lines. He wrote the software. I am here to hand it to the community and see what you do with it.
This is a hands-on session, so bring a laptop (macOS on Apple silicon or Linux; Windows is not supported yet).
In 30 minutes you will:
1. Install Reliaburger with one command and start a node on your laptop. No container runtime, no root, no VMs: the built-in process runtime runs plain binaries.
2. Deploy an app from a twelve-line TOML file and watch it in the terminal dashboard and the web UI.
3. Ship a deliberately sick version and watch the health checker restart it, then roll back with one command.
4. Break it on purpose with the built-in fault injector: kill an instance, freeze another, and let `relish wtf` explain what just happened.
5. Read the manual and the source code from inside the binary, so you can keep going after the session.
6. Tell us what broke and what you liked.
Reliaburger is free and open source (Apache 2.0), and it is version 0.1.0, so it will break. That is where you come in. We are looking for helpers, burger ambassadors and contributors: people who will run it, break it, file the issues, write the docs and tell their teams. There are millions of practitioners running parts of Kubernetes they do not need, with more people than they should need to run it. We want this to become a community-driven project, just as our conferences are.
Everyone who makes it through the six steps and gives feedback goes into a raffle at the end. There is hardware on the table, and it is worth your 30 minutes.... Read more
Multi-agent AI systems are distributed systems, but many still pass data through prompts. This loses information, wastes tokens, and makes failures hard to manage.
This talk presents a shared data bus where agents exchange a database ID. Using a real three-agent pipeline, I’ll cover retries, conflicting writes, access control, monitoring, and cleanup. I will also provide a guide for choosing between prompts, queues, files, and databases.... Read more
Giving an AI agent the ability to execute code changes the infrastructure it needs. When I was working at Browser Use, that meant rethinking where agents ran, which credentials they could access, and how their work survived a process dying.
I'll walk through the architecture we built: isolated agent runtimes, scoped access through a control plane, and state stored outside the worker. Then I'll cover what still went wrong, including a failed state restore that was mistaken for a fresh session and allowed blank state to replace valid conversation history. We'll look at the recovery changes that followed and the operational tradeoffs we encountered moving between runtimes.
The talk draws on publicly documented Browser Use work. Attendees will leave with practical lessons for isolating untrusted workloads, distinguishing missing state from unavailable storage, and deciding which parts of an agent platform they need to operate themselves.... Read more
For some workloads, success means a smaller cloud bill; for others, it means better performance. This talk demonstrates how to put those goals into practice. We’ll walk through analyzing workload behavior, selecting an optimization mode, and evaluating recommended configurations against cost and performance requirements. Attendees will see how to turn workload data into actionable infrastructure decisions.... Read more
LLMs break the assumptions traditional monitoring is built on, a service can be "up" while quietly hallucinating, drifting, or burning your token budget. This talk shows how to apply SRE discipline to AI systems i.e defining SLOs for inference quality and latency, token-level metrics, prompt tracing, and structured logging across LLM pipelines. We'll then go further with chaos engineering for AI, injecting failures into model-serving infrastructure and automating resilience experiments. Attendees will leave with a practical blueprint for making LLM systems as observable and reliable as the rest of their stack, drawn from the author's book *Observability for Large Language Models* (Apress, 2026).... Read more
MCP servers, Skills, and agents are a commodity now. But composing these primitives into production-ready developer workflows is not easy. Building the abstraction layers that help agents get the context they need, without burning tokens on low-level detail and wiring it all into a cohesive system architecture takes time. Once you've figured out the patterns, they're repeatable.
This talk is about those patterns, the decisions and trade-offs you face when turning these building blocks into developer workflows.
I'll demonstrate one workflow live to make it concrete: an alert fires, an agent assembles context, produces a root-cause report, helps a developer troubleshoot, and proposes a fix the platform executes.
From that one workflow, I'll pull out patterns you can reuse anywhere, treating the platform as the agent's context source, letting agents express intent while the platform reconciles deterministically, scoping permissions to bound blast radius, and placing the human handoffs that keep people in the loop.
These are repeatable patterns I learned building OpenChoreo, a CNCF project and they apply to any platform, not just this one.... Read more
A monitoring stack that fires a thousand alerts a day does not have a monitoring problem. It has a reading problem. Nearly all of it is benign and repeats the same twenty shapes - and somewhere in there is a real incident that looks exactly like the rest until someone digs.
The obvious fix is to point an agent at the queue and let it triage. We tried that. It is confidently wrong, and the confidence is the dangerous part: it returns a fluent, well-argued verdict on every alert at the same level of certainty, whether it pulled real evidence or just pattern-matched the title. At a thousand a day, even a small confidently-wrong rate is dozens of silent misfilings - and the ones it buries are the unfamiliar ones, which is exactly where the real incidents live.
I am a Principal Engineer in Microsoft Fabric, where I am the tech lead and architect for capacity management. What finally worked was refusing to let one agent do both jobs.
Triage got cut back until routing is all it can do. It groups by signature and hands off. It is not allowed to conclude, to close, or to declare anything benign - that is the judgement it is worst at and the one with the worst downside.
Investigation was then split across dedicated agents, one per failure domain. Each knows only its own telemetry, its own known-benign patterns, and its own escalation bar. That narrow hypothesis space is what makes "I could not find evidence for this" an answer the agent will actually give, instead of a plausible story. It also means each one can be tested on its own against past incidents, which a single generalist triager never can be.
The metric that matters is recall on the rare real incident, not how much of the queue got closed.
I will go through the routing design, how the investigation agents are kept honest, what is escalated to a human on purpose, and the real problems we only caught because a specialist refused to answer.... Read more
What does Site Reliability Engineering look like when the "site" is an actual physical place?
At a physical site, software is only one part of the system. Networks can disappear, machines have their own state, sensors provide incomplete evidence, and environmental conditions keep changing regardless of whether our infrastructure is healthy. A reliable cloud service does not necessarily mean a reliable site.
Drawing from engineering systems that span cloud infrastructure, on-site computing, machines, sensors, and physical processes, this talk explores where the traditional reliability boundary starts to break down. It proposes a broader way to think about observability, failure, recovery, and system health when the physical site itself becomes part of the system we are responsible for keeping reliable.... Read more
Chaos engineering has been standard practice in distributed systems for over a decade - kill a node, inject latency, watch what breaks, fix it before production does. Agentic AI systems now chain LLM calls to external tools (via MCP or similar), but there's no equivalent discipline yet for what happens when those tools misbehave. In this 30-minute session, I'll apply classic chaos-engineering principles - fault injection, steady-state hypothesis, blast-radius containment - to an actual agent pipeline: tool calls that time out, return malformed data, or silently disappear mid-chain. I'll walk through what broke loudly vs. what failed silently and went undetected, share the one number from testing that surprised me most, and leave the audience with a minimal "chaos toolkit for agents" checklist they can apply to their own systems.... Read more
18:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
55 Stockton St, San Francisco,
CA 94108, United States
Sponsors & Partners
Want to become a sponsor? Get in touch!
Smita Pasumarthi & Chandan Chilumula
TruffleRoot
From Observability to Causality: Teaching Machines How Your System Fails
Abstract
Modern observability is very good at telling us what is unhealthy. It is much worse at telling us why.
A database can be on fire because it caused an incident—or because something several hops away made it the place where the damage accumulated. An LLM can read all of the telemetry and still make the same mistake. The problem is not always a lack of data or intelligence. Sometimes what is missing is an understanding of how this particular system behaves when things go wrong.
While building TruffleRoot, we’ve been exploring a different approach: give machines an explicit, evolving model of an environment—how it is connected, what changes over time, how failures can propagate, and what evidence supports or contradicts an explanation.
This talk is about that shift from observability to causality: why generic assumptions about failure break down in real systems, what it means for a system to learn how your system fails, and where deterministic reasoning and LLMs each fit into that picture.
The goal isn’t an AI that tells a better story about an outage. It’s a system that learns how your system fails and can show its work.
Bio
Smita Pasumarthi builds enterprise AI applications at truffleroot.ai, drawing on extensive experience in enterprise integration and automation. Her expertise spans generative AI, LLMs, AI infrastructure, MuleSoft, Boomi, Workato, and Python.
Chandan Chilumula is a technology professional based in the San Francisco Bay Area, with experience in software solutions and enterprise technology.
Aron Eidelman
Google
Don't Let Your AI's Patch Break Prod: SRE Patterns for Autonomous Remediation
Abstract
Software and security teams have never had to patch so much code so quickly. AI agents can generate vulnerability fixes in seconds, but pushing unverified, machine-authored code to production has led to trading a CVE for a production outage.
To solve this problem, SRE principles are now gaining new traction in application security and development teams that realize mitigating one risk shouldn't create another. To benefit at all from agents in SSDLC, we need operational cultures that recognize reliability as a combined social and technical measure, rather than just another tooling challenge.
In this talk, we’ll walk through the practical mechanics of safely remediating vulnerabilities with agents. We’ll cover how TDD provides the baseline contract to keep patches from breaking existing features, why symptom-based test coverage catches regressions traditional tests miss, and how sandboxed runtimes isolate untrusted agent execution. Finally, we’ll look at closing the loop with automated canary analysis and SLO-driven rollbacks so fixes can deploy without risking production.
We'll stick to real scenarios that helped us patch hundreds of repositories without breaking any relying applications (at least as far as users could tell).
Bio
Aron Eidelman is a Senior Developer Relations Engineer and Security Advocate at Google. He works on AI software security, leading developer enablement for Google's AI Threat Defense program, CodeMender, and Antigravity security skills. He co-authored Kaggle's "5 Days of AI Agents" curriculum, which trained over 350,000 developers globally in secure coding. Aron is a frequent speaker on secure software development, with recent talks at Google Cloud Next '26 and the Zenity AI Agent Security Summit.
Outside of Google, Aron serves as the Chairman of Azure Printed Homes. The company 3D prints modular housing from recycled plastic, repurposing 150,000 plastic bottles per unit for wildfire recovery projects in LA and a new 62-unit housing village in San Luis Obispo.
Artemis Leonardou
NOFire AI
The Junior Engineer Is Gone and Your On-Call Has Not Noticed
Abstract
The career ladder most teams still hire into does not really describe the people joining them now, and on-call is where that gap shows up first. I started shipping production code at fourteen. I never got the version of this job where someone walks you through syntax and then slowly widens your scope, because by the time I started, that was not how anyone around me was learning. We went from tutorials straight to orchestrating systems with agents. That produces engineers who move unusually fast and who have real gaps in unusual places, and most runbooks and escalation paths assume neither of those things. In fifteen minutes I want to be specific about what that looks like from the inside. Where we are genuinely faster, where we are genuinely worse, and the two things I would change about code review and on-call rotation if you are hiring people like me. You would probably rather hear it from someone currently living it than from a report about us.
Bio
Artemis Leonardou is a software engineer and DevRel at NOFire AI. She has been building since she was 14 and now ships AI-generated code to production as a matter of routine, which is where most of what she talks about comes from. She was selected among the top 2000 global CS talents for Y Combinator's AI Startup School, has won first place in three hackathons, and her technical content has crossed 2 million views as Artemis Codes. She has spoken at React Summit and TechBiz, speaks at DevConf.US in September 2026, and leads a hands-on workshop at Open Conf in November. She is 19 and still a full-time student.
Wei-Chin Call
Grafana Labs
From Thumbs-Down to RCA: Closing the Loop on AI Agents
Abstract
Your on-call rotation doesn't page you when an AI agent gives a customer bad advice, but it should.
As LLM-powered agents move from chatbot novelty into production systems that make decisions, recommend actions, and interact directly with users, SRE practices built around traditional services need to evolve to account for non-deterministic, natural-language behavior.
In this talk, I’ll walk through a live AI agent, an intentionally “broken” education assistant, to demonstrate what agent observability looks like in practice. We’ll follow distributed traces across the LLM call and its downstream tool calls, see an automated evaluator catch a harmful response before it reaches a real audience, and trace a deliberate service outage end-to-end through the agent’s error-handling path.
I’ll then close the loop the way an SRE would: a simple thumbs-up/down control on every agent response feeds directly into a conversation-rating signal. A dissatisfied user is no longer just a qualitative complaint. It becomes a queryable, traceable signal that can be correlated with the trace that produced it.
Finally, we’ll explore what familiar SRE concepts such as SLOs, error budgets, and incident response begin to look like when a language model becomes part of the production service.
Attendees will leave with a practical framework for instrumenting and operating AI agents as production dependencies, and a healthy skepticism toward the idea that AI agents somehow don't need SREs.
Bio
Wei-Chin Call is a Staff Emerging Products Manager at Grafana Labs, where she helps shape the company's next generation of agentic, full-stack observability products. She joined Grafana Labs as a Senior Solutions Engineer, partnering with customers on technical proof-of-concepts and driving adoption of the Grafana Cloud stack.
Prior to Grafana Labs, Wei-Chin managed a Solutions Engineering team at Cribl and held technical evangelism and sales engineering leadership roles at AppDynamics, where she built sales enablement curricula and led cloud-native training initiatives. She holds certifications in OpenTelemetry (OTCA), Kubernetes (CKAD, CKA), and Prometheus (PCA).
Geoff White
Nexsys Systems LLC
Open Source SRE Agents for your internal Labs
Abstract
At 2 a.m. during the Kubernetes buildout for my home agent lab, my SRE agent Gandalf helped me chase down a routing bug quietly breaking federation between the gateway cluster and five agent homeservers behind a VPN tunnel. A few weeks later, Gandalf moved off my laptop and onto a proper server, the same week the swarm outgrew my desk. That's the origin story of Agent-Matrix: a fleet of sovereign AI agents, each with its own identity on the Matrix protocol, letting people message agents from Element or FluffyChat on their phone exactly like messaging a colleague, without routing a single message through Telegram, Signal, or anyone else's chat cloud.
This talk is about what it takes for one SRE, not a platform team, to stand up that fleet on a spare server with open protocols and open source tools. You'll meet Gandalf, the SRE agent watching for Grey Rhinos and helping stabilize Black Jellyfish style cascades, and Galadriel, the research agent drawing on OpenBrain, the fleet's shared long-term memory. I'll give an honest accounting of what's solid today (federation, per-agent identity, shared memory) and what's still being built (full end-to-end encryption via a Rust rewrite of the Matrix MCP server), so you can judge what's worth building at your own org, on your own budget.
Bio
Geoff White is a Principal-level Site Reliability and AI
Infrastructure Engineering leader with decades of experience designing
and operating enterprise-scale Virtual and Kubernetes platforms. His
expertise spans NVIDIA AI Ops, Kubernetes, vSphere, Terraform,
Ansible, NSX, and observability stacks such as Prometheus, Grafana, and
ELK. Geoff has participated in and led infrastructure deployments and reliability transformations for organizations including Dell and Autodesk, blending automation, DevSecOps, and Agentic Engineering. A hands-on architect fluent in Python, C/C++, and Go, he applies performance engineering and root-cause discipline to build scalable, secure, and compliant enterprise platforms. He values measurable SLOs, cost efficiency, troubleshooting, and engineering mentorship.
Edward Borukhov
Walmart Global Tech
When the On-Call Engineer Is an Agent: Incident Response for AI-Driven Systems
Abstract
How AI agents fail differently than traditional software—silently, non-deterministically, and in ways that drift as models update—and practical patterns for incident classification, observability, and postmortems built for probabilistic systems.
Bio
Edward Borukhov is the Fractional Head of Reliability & Operations at Beyond Uptime, where he helps organizations build incident response and operational maturity for both traditional and AI-driven systems. He spent 15 years in enterprise SRE and technical operations at Walmart Labs, Symantec, Harmonic, and Apple, and now works as a fractional reliability lead for companies navigating the shift from reactive incident response to proactive, AI-enabled operations.
Robbie McKinstry
Wack Incorporated
Overcoming the Trust Deficit with SRE Agents
Abstract
Most SRE agents only do half of the job: they wait for an incident to occur, then hand off diagnostics. While that’s better than nothing, it stops short of resolving the incident. The flipside is that no one actually wants an agent to autonomously resolve incidents if it can only deliver 90% of the time. That would mean for the other 10%, the agent has made the wrong call on business-impacting systems, exacerbating conditions and making a postmortem more difficult to conduct. While coding agents are allowed to act autonomously since bad output can be caught in review, incident response doesn’t have that luxury, leaving SRE agents to watch instead of act. This talk covers the limitations of bleeding-edge incident response agents, and provides a roadmap toward trustworthy, verified, and independent agent playbooks. Playbooks require agents to show their work to an independent oracle that judges whether the case they’ve made is plausible or flimsy. Only once a check passes does the agent move to the next step in the playbook, with each step gating the tools it's allowed to touch. What you end up with is your existing runbook turned into a state machine, where every transition locks down what the agent can do next. I'll walk through the architecture, what actually counts as evidence to a judge, as well as the weaknesses of this approach.
Bio
Robbie McKinstry is the founder and CEO of Wack where he’s building MultiTool, an observability tool that provides intelligent, thresholdless alerts on top of your existing telemetry. Before MultiTool Robbie was on the OSS core team at Pulumi, and before that an early employee and research scientist in HashiCorp's office of the CTO. His career has focused on taking hard ideas off paper and turning them into usable developer tools.
Yunhao Jiao
TestSprite
Verification After the Deploy: Catching What AI-Written Code Breaks Before Your Users Do
Abstract
Most of the quality tooling we have added for AI-written code looks at the diff. Code review, static analysis, and pre-merge checks all ask whether the change looks correct. None of them answer the question a reliability engineer actually cares about, which is whether the running application still works.
When a person wrote the code, that mattered less. They knew what they had changed and what else it might touch. When an agent writes the code, nobody knows. The pull request looks clean, the checks pass, and the first honest signal arrives after the deploy, often when a user hits it.
This talk makes the case for moving verification to the other side of the deploy event, and shows what that looks like in practice. A testing agent opens the live build in a real browser, exercises it the way a user would, and hands the failure back to the coding agent with enough context to act on it: the failing step, a screenshot, a DOM snapshot, and a root-cause hypothesis. The repair happens before a human is involved.
Yunhao will close with what TestSprite measured across CoderCup, a public benchmark where frontier coding agents built the same application under identical conditions, including a pattern the team calls feature retention: working features quietly disappearing as later edits land in a codebase whose author does not remember what it built last week.
Bio
Yunhao Jiao is a Yale University graduate with a master's degree and spent nearly five years at AWS as a Senior Software Engineer, where he helped build the Contract Tests framework for AWS CloudFormation, ensuring resource types operated as expected throughout their lifecycle. He has been active in NLP research since 2015 and published his first AI paper as lead author in 2017. Yunhao founded TestSprite and participates in the Techstars Miami and YC China accelerator programs.
Hilliary Lipsig
Red Hat
Burning the Ships: Migrating Our Core Logging Stack with No Way Back
Abstract
What happens when your core logging framework has a critical exploit, rollback is a security risk, and you can’t test the new stack until it hits production? You burn the ships. In this talk, I share the high-stakes engineering story of a complete, one-way logging migration across thousands of live clusters. Forced by an active vulnerability to bypass traditional safety nets we had to rely on napkin-math resource estimations, synthetic traffic testing, table top exercises of what could go wrong, and a terrifyingly literal "test in production" rollout.
I will share the exact strategies we used to pull off this zero-rollback cutover without plunging our systems into telemetry blindness. You’ll leave with concrete patterns for calculating production risk under extreme constraints, validating data pipelines on the fly, and failing forward.
Bio
Hilliary Lipsig is an Engineering Leader, Author, and Senior Principal Site Reliability Engineer at Red Hat, where she focuses on Azure Red Hat OpenShift. Known for leading follow-the-sun SRE teams, she is also the host of the popular YouTube livestream GitOps Guide to the Galaxy. Outside of her work in reliability engineering, Hilliary serves as Secretary on the Board of Directors for the Gender Education Network and is an active voice in the open-source and Kubernetes communities.
Jonathon Klobucar
Socket
New Incident, Who Dis? Incident Response Beyond Engineering
Abstract
SRE already solved incident response: severities, a commander, a working channel, a timeline, a blameless postmortem. Then we kept it inside engineering. Security has its own process, but it's usually less battle-tested; IT runs incidents out of a ticket queue; marketing runs them out of a group text. And the worst incidents are the ones that hit all three at once.
That process is the most valuable thing an SRE team owns, and it's also the easiest way to stop being the team that says no and start being the team everyone calls. This talk covers what I learned bringing the SRE incident model to security, IT, and a PR crisis: where it worked without changes, where it fell apart, and how to show up as a partner instead of a process cop. We'll also cover where AI genuinely helps and how it hands teams with no incident experience a polished process they can't actually run. You'll leave with a severity model every department can read, a handoff checklist, and a DR/IR tabletop format the whole company practices together and, for once, actually enjoys.
Bio
Jonathon Klobucar is a tech lead at Socket. He got there the long way, building product at Rdio, then a container platform at Apcera back when that idea was still weird. He founded the SRE team at Iterable and helped launch some Apple services you've used. At Gridmatic he built and ran the real-time operations desk that connected a battery fleet to the California and Texas power grids, which is where he started noticing that reliability problems and security problems are mostly the same problem. Sublime made that official: he built their incident response, MDR, and detection programs from nothing.
Oleg Sotnikov
AppMaster
One Container, One Database per Tenant: Running Hundreds of Thousands of AI-Generated Apps
Abstract
When your platform generates your customers' applications, every incident in their production is your incident. This talk is a field report from running a no-code platform's runtime plane at a peak of 250,000 concurrently hosted containers: one container and one PostgreSQL database per tenant app across multiple regions, with a control plane and a build plane feeding it. The applications themselves are produced by an AI-driven generator from customers' blueprints, which adds a failure class most fleets never see — code that nobody reviewed, regenerated on demand. I'll cover the reliability model we settled on and why: blast-radius isolation vs. density, per-tenant observability, fleet-wide regeneration and redeploy when the generator itself changes, noisy-neighbor handling on shared hosts, and the failure modes that only exist when the code in production was compiled from a customer's blueprint 10 minutes ago. Practical patterns for anyone operating a fleet of workloads you don't control.
Bio
Oleg has 25+ years in software development and shipped his first commercial app in 2003. He holds 50+ professional certifications from Microsoft, Cisco, IBM, Dell, EC-Council and others, and has built 9 startups over the past 17 years, with 2 exits. Since 2020 he has been CTO and Co-Founder of AppMaster, an AI code-generation platform launched before ChatGPT became a thing. Since 2025 he works as a Fractional CTO with 20+ companies across 12 countries, helping them transition to AI-native software engineering and SRE.
Romain Priour
LangChain
Operating Reliable Software Across Customer Clouds
Abstract
LangSmith’s Bring Your Own Cloud (BYOC) model lets customers run LangSmith in their AWS accounts across many regions. Our team is responsible for provisioning the infrastructure, shipping upgrades, monitoring deployments, and proactively resolving issues. As the fleet grows, we need to keep doing that reliably without a proportional increase in manual work.
In this talk, I’ll show how we operate this fleet efficiently while keeping operational overhead low. I'll cover how we provision and continuously reconcile infrastructure, how we manage private customer environments through a shared control plane, and how we build internal tooling to keep day-to-day operations manageable as the fleet grows.
We’ll follow a release through our deployment pipeline: validating changes, rolling them out across customer environments, designing upgrades to avoid downtime, and handling failures along the way. Finally, we’ll look at how we leverage telemetry to understand what’s happening across the fleet, surface failures that require attention, and use internal AI tooling to help troubleshoot incidents.
Bio
Romain Priour is a software engineer on LangChain’s infrastructure team, where he leads the team behind LangSmith’s Bring Your Own Cloud (BYOC) platform. Previously, he helped build an intelligent log compression engine at Grepr and worked on SAP Business Network at SAP. Romain studied computer science at UC Berkeley. Go Bears!
Vanshika Gupta
ViralCircle
How to Turn LinkedIn Into Your #1 B2B GTM Channel
Abstract
Most B2B companies treat LinkedIn as a content channel. This talk breaks down how to turn it into a GTM engine, including the content formats that work (and the ones that don’t), how to build the right audience, and how to turn attention into conversations, pipeline, and customers.
Bio
Vanshika Gupta is the founder of ViralCircle, a marketing agency that helps AI and tech companies turn founder-led content into a growth and customer acquisition channel.
Previously, she worked in marketing at YC-backed Julius AI. With a background in computer science and technology management, she brings together technical understanding and creative marketing for the AI and startup ecosystem.
Hubert Chen
Fidian
Beyond Unit Tests: Bringing Intelligent Agent Evals into CI/CD Pipelines
Abstract
As software shifts from deterministic code to non-deterministic LLMs and autonomous agents, traditional CI/CD practices fall short. Treating probabilistic agentic workflows like
rigid software forces tough choices around test coverage, pass rates, and build times. This talk explores how to modernize DevOps pipelines for the AI era. We'll examine how balancing mode
l selection, evaluation costs, and pipeline speed affects both developer velocity and system reliability. We will dive into improvement strategies like intelligent task selection, lightwei
ght judging, test tiering and eval prediction can reduce evaluation time and costs.
Bio
Hubert Chen is a San Jose–based Infrastructure Engineering leader and architect with over 30 years of experience scaling backend systems, SRE, and DevOps operations. Currently a Member of Technical Staff at Fidian focused on infrastructure for AI testing and cloud-managed Kubernetes, he brings extensive senior leadership from Palo Alto Networks as Distinguished SRE and Director of Infrastructure, where he managed hundreds of Kubernetes clusters across AWS and GCP, multi-petabyte data lakes, and tens of millions in cloud spend. His background also includes leading infrastructure and systems scaling at Branch and Shutterfly, supported by a BS in Computer Science from the University of Michigan.
Ganesh Nathan
Autolake
Your Bill Is an SLO: Reliability for a Data Platform With No Servers
Abstract
We run a multi-tenant data lakehouse for higher education and public sector organizations on AWS serverless services only: S3, Glue, Athena, Lambda, Step Functions, and Iceberg tables. No cluster, no Kubernetes, and no on-call rotation. That changes what reliability means, because the signals SREs usually rely on, CPU, memory, pod health, do not exist. Our answer is that cost is the most honest signal you have. A query that returns the right answer at ten times the expected cost is a degradation, and we treat it like one. This talk walks through the SLOs we built around cost per query, file count per partition, and snapshot age, and how those SLOs drive automated remediation: compaction, snapshot expiry, catalog reconciliation, and retention enforcement that run themselves and report back. We will also cover the guardrails that keep petabyte-scale storage healthy without a human in the loop: hidden partitioning so users cannot write expensive queries by accident, and per-tenant cost limits that stop a runaway workload before it becomes a bill. You will leave with a working model for SRE on serverless data infrastructure and a checklist you can apply to any Iceberg on S3 deployment.
Bio
Ganesh Nathan is co-founder of Autolake, a serverless data lakehouse platform built entirely on AWS native services for higher education and public sector organizations. He has spent 30 years building data infrastructure across public sector and enterprise environments, and currently runs multi-tenant, petabyte-scale Iceberg on S3 deployments with no clusters and no on-call rotation. His work focuses on what reliability, governance, and cost control look like when there are no servers to watch.
Mark Pawlikowski
Reliaburger
Kill It and Watch It Come Back: Hands-on with Reliaburger
Abstract
I have never run a Kubernetes cluster. Yes, I said it.
Are you still with me? Cool!
Over the past last five years I talked to thousands of SREs and platform engineers at the conferences I run with my team: SREday, LLMday, PLATFORMday and Conf42.
The same story keeps coming back. Before the first container hits production you step into the world of Kubernetes, and then you need to build a whole team just to keep it running. A kolljoy for many smaller teams. We started talking with my brother, Miko Pawlikowski, how we could help the community ease some of these pains.
Miko has been running platform teams on Kubernetes for 10+ years, leading him to start writing Reliaburger: one Rust binary with the scheduler, gossip, Raft, eBPF service discovery, ingress, mTLS, registry, metrics, logs, GitOps and chaos testing inside, driven by TOML. An app with health checks, ingress and autoscaling is twelve lines. He wrote the software. I am here to hand it to the community and see what you do with it.
This is a hands-on session, so bring a laptop (macOS on Apple silicon or Linux; Windows is not supported yet).
In 30 minutes you will:
Install Reliaburger with one command and start a node on your laptop. No container runtime, no root, no VMs: the built-in process runtime runs plain binaries.
Deploy an app from a twelve-line TOML file and watch it in the terminal dashboard and the web UI.
Ship a deliberately sick version and watch the health checker restart it, then roll back with one command.
Break it on purpose with the built-in fault injector: kill an instance, freeze another, and let relish wtf explain what just happened.
Read the manual and the source code from inside the binary, so you can keep going after the session.
Tell us what broke and what you liked.
Reliaburger is free and open source (Apache 2.0), and it is version 0.1.0, so it will break. That is where you come in. We are looking for helpers, burger ambassadors and contributors: people who will run it, break it, file the issues, write the docs and tell their teams. There are millions of practitioners running parts of Kubernetes they do not need, with more people than they should need to run it. We want this to become a community-driven project, just as our conferences are.
Everyone who makes it through the six steps and gives feedback goes into a raffle at the end. There is hardware on the table, and it is worth your 30 minutes.
Bio
Mark Pawlikowski runs SREday, LLMday and PLATFORMday, in-person conferences for the people who keep production up, and Conf42 online. He is not an engineer. He has spent the last five years listening to SREs and platform engineers describe what their work is actually like, and now brings Reliaburger, his brother Miko's batteries-included container orchestrator, to that community to play with, break and build on.
Eddie Tejeda
Hotdata
The Agent Data Bus: Shared State for Multi-Agent Systems
Abstract
Multi-agent AI systems are distributed systems, but many still pass data through prompts. This loses information, wastes tokens, and makes failures hard to manage.
This talk presents a shared data bus where agents exchange a database ID. Using a real three-agent pipeline, I’ll cover retries, conflicting writes, access control, monitoring, and cleanup. I will also provide a guide for choosing between prompts, queues, files, and databases.
Bio
Eddie Tejeda is co-founder and CTO of Hotdata.dev , which helps AI agents create databases in milliseconds. He is a two-time founder and previously led engineering at Mode Analytics and ThoughtSpot.
Larsen Cundric
SpaceXAI
Running Untrusted Agents in Production
Abstract
Giving an AI agent the ability to execute code changes the infrastructure it needs. When I was working at Browser Use, that meant rethinking where agents ran, which credentials they could access, and how their work survived a process dying.
I'll walk through the architecture we built: isolated agent runtimes, scoped access through a control plane, and state stored outside the worker. Then I'll cover what still went wrong, including a failed state restore that was mistaken for a fresh session and allowed blank state to replace valid conversation history. We'll look at the recovery changes that followed and the operational tradeoffs we encountered moving between runtimes.
The talk draws on publicly documented Browser Use work. Attendees will leave with practical lessons for isolating untrusted workloads, distinguishing missing state from unavailable storage, and deciding which parts of an agent platform they need to operate themselves.
Bio
Larsen Cundric works on Grok at SpaceXAI. Previously, he was a founding engineer at Browser Use, where he worked on browser agent infrastructure, sandboxing, and control planes. He writes about the architecture and operational tradeoffs behind running agents in production.
Jagan Jagannathan
Serra Labs
Different Workloads, Different Wins
Abstract
For some workloads, success means a smaller cloud bill; for others, it means better performance. This talk demonstrates how to put those goals into practice. We’ll walk through analyzing workload behavior, selecting an optimization mode, and evaluating recommended configurations against cost and performance requirements. Attendees will see how to turn workload data into actionable infrastructure decisions.
Bio
I am a founder of Serra Labs, where we help teams optimize cloud workloads for cost, performance, and business value. My background spans Cloud optimization, AIOps, and Generative AI, with current focus on AI Clouds.
Ankush Sharma
FloQast
SRE for LLMs: Observability and Chaos Engineering for AI
Abstract
LLMs break the assumptions traditional monitoring is built on, a service can be "up" while quietly hallucinating, drifting, or burning your token budget. This talk shows how to apply SRE discipline to AI systems i.e defining SLOs for inference quality and latency, token-level metrics, prompt tracing, and structured logging across LLM pipelines. We'll then go further with chaos engineering for AI, injecting failures into model-serving infrastructure and automating resilience experiments. Attendees will leave with a practical blueprint for making LLM systems as observable and reliable as the rest of their stack, drawn from the author's book Observability for Large Language Models (Apress, 2026).
Bio
Ankush Sharma is an AI systems architect and veteran technologist with over 22 years of experience across distributed systems, cloud infrastructure, and AI platform engineering. He is the author of Observability for Large Language Models: Site Reliability and Chaos Engineering for AI at Scale (Apress, 2026) and has led engineering teams at global technology companies (Microsoft, Citrix, startups) while contributing to open-source AI infrastructure projects. He is currently Head of Infrastructure (SRE & AI Engineering) at FloQast, a late-stage fintech startup. His work has been recognized through conference talks, patents, and developer forums. He is based in the San Francisco Bay Area.
Sameera Jayasoma
WSO2
Beyond MCP and Skills: Architecting Agentic Developer Workflows
Abstract
MCP servers, Skills, and agents are a commodity now. But composing these primitives into production-ready developer workflows is not easy. Building the abstraction layers that help agents get the context they need, without burning tokens on low-level detail and wiring it all into a cohesive system architecture takes time. Once you've figured out the patterns, they're repeatable.
This talk is about those patterns, the decisions and trade-offs you face when turning these building blocks into developer workflows.
I'll demonstrate one workflow live to make it concrete: an alert fires, an agent assembles context, produces a root-cause report, helps a developer troubleshoot, and proposes a fix the platform executes.
From that one workflow, I'll pull out patterns you can reuse anywhere, treating the platform as the agent's context source, letting agents express intent while the platform reconciles deterministically, scoping permissions to bound blast radius, and placing the human handoffs that keep people in the loop.
These are repeatable patterns I learned building OpenChoreo, a CNCF project and they apply to any platform, not just this one.
Bio
Sameera Jayasoma is VP & Distinguished Engineer at WSO2, Creator and Lead Architect of OpenChoreo—an open-source internal developer platform built on Kubernetes and Backstage—and Lead Architect for the Ballerina Language Platform.
With over 18 years at WSO2, Sameera has spearheaded major platform initiatives, including co-designing the Ballerina language spec, leading compiler architecture, and driving WSO2’s strategic adoption of Go for backend infrastructure. Based in the Greater Seattle Area, he is a frequent speaker, open-source maintainer, and engineering mentor.
Girish Konda
Microsoft
A Thousand Alerts a Day, and One Real Incident Hiding In Them
Abstract
A monitoring stack that fires a thousand alerts a day does not have a monitoring problem. It has a reading problem. Nearly all of it is benign and repeats the same twenty shapes - and somewhere in there is a real incident that looks exactly like the rest until someone digs.
The obvious fix is to point an agent at the queue and let it triage. We tried that. It is confidently wrong, and the confidence is the dangerous part: it returns a fluent, well-argued verdict on every alert at the same level of certainty, whether it pulled real evidence or just pattern-matched the title. At a thousand a day, even a small confidently-wrong rate is dozens of silent misfilings - and the ones it buries are the unfamiliar ones, which is exactly where the real incidents live.
I am a Principal Engineer in Microsoft Fabric, where I am the tech lead and architect for capacity management. What finally worked was refusing to let one agent do both jobs.
Triage got cut back until routing is all it can do. It groups by signature and hands off. It is not allowed to conclude, to close, or to declare anything benign - that is the judgement it is worst at and the one with the worst downside.
Investigation was then split across dedicated agents, one per failure domain. Each knows only its own telemetry, its own known-benign patterns, and its own escalation bar. That narrow hypothesis space is what makes "I could not find evidence for this" an answer the agent will actually give, instead of a plausible story. It also means each one can be tested on its own against past incidents, which a single generalist triager never can be.
The metric that matters is recall on the rare real incident, not how much of the queue got closed.
I will go through the routing design, how the investigation agents are kept honest, what is escalated to a human on purpose, and the real problems we only caught because a specialist refused to answer.
Bio
Girish Konda is a Principal Software Engineer at Microsoft with over eight years of experience building scalable distributed systems and machine learning platforms. He has led large scale engineering projects and backend teams, with previous experience at Amazon and AWS.
An Phan
Hippo Harvest
When the Site Is Actually a Site: Rethinking Site Reliability for the Physical World
Abstract
What does Site Reliability Engineering look like when the "site" is an actual physical place?
At a physical site, software is only one part of the system. Networks can disappear, machines have their own state, sensors provide incomplete evidence, and environmental conditions keep changing regardless of whether our infrastructure is healthy. A reliable cloud service does not necessarily mean a reliable site.
Drawing from engineering systems that span cloud infrastructure, on-site computing, machines, sensors, and physical processes, this talk explores where the traditional reliability boundary starts to break down. It proposes a broader way to think about observability, failure, recovery, and system health when the physical site itself becomes part of the system we are responsible for keeping reliable.
Bio
An Phan is a Senior Data Infrastructure Engineer at Hippo Harvest, where he works on infrastructure spanning cloud systems and physical production sites. He has more than a decade of experience building software, data, and AI systems, with recent work focused on the challenges that emerge when computing meets the physical world. He regularly speaks about data infrastructure, Physical AI, and real-world production systems.
Hazel Mahajan
Cisco
Chaos Engineering for AI Agents: What Happens When Your Tools Fail, Time Out, or Vanish
Abstract
Chaos engineering has been standard practice in distributed systems for over a decade - kill a node, inject latency, watch what breaks, fix it before production does. Agentic AI systems now chain LLM calls to external tools (via MCP or similar), but there's no equivalent discipline yet for what happens when those tools misbehave. In this 30-minute session, I'll apply classic chaos-engineering principles - fault injection, steady-state hypothesis, blast-radius containment - to an actual agent pipeline: tool calls that time out, return malformed data, or silently disappear mid-chain. I'll walk through what broke loudly vs. what failed silently and went undetected, share the one number from testing that surprised me most, and leave the audience with a minimal "chaos toolkit for agents" checklist they can apply to their own systems.
Bio
Hazel Mahajan is a software engineer specializing in backend systems, distributed architectures, and cloud infrastructure. She has worked with technologies including Java, Python, Go, Kafka, AWS, Docker, and Kubernetes, with experience building scalable APIs and event-driven systems. She recently completed an MS in Computer Science and Engineering at the University at Buffalo and previously worked as a Software Engineer at Delhivery.
Matt Schillerstrom
Harness
KeynoteYour Customers Are Already Building Your Roadmap
Abstract
What if the best feature request isn't a ticket but something your customer has already built?
AI agents are making that possible. Customers can now create specialized workflows on a platform using their own context, tools, and operational knowledge. At Harness, we are seeing this firsthand with Worker Agents.
In this talk and demo, we’ll show how customers are extending the platform, what those agents reveal through usage and telemetry, and how that gives product teams a new way to discover what customers need next.
The strongest signal is not just what customers ask for. It's what they actually build and continue to use. The feature request of the future may be something the customer has already built.
Bio
Matt Schillerstrom is a product and technology leader specializing in DevOps, with a focus on Site Reliability Engineering and Chaos Engineering. He currently leads product management and product marketing at Harness, where he works on solutions that help teams deliver software more reliably and efficiently. With a background spanning engineering, product, and resiliency strategy at companies like Target and Gremlin, Matt brings a deep technical perspective to building and scaling modern software systems. He’s also an active community leader and mentor, organizing Chaos Engineering events and supporting others in the tech industry.
Arjun Iyer
Signadot
Keynote10x the PRs, Same Error Budget: Validating Agent-Written Code Before Prod
Abstract
Coding agents changed the math on software delivery. Teams that shipped ten PRs a day now see a hundred, and most of that code was never run by a human before it hit CI. The error budget didn't grow with it.
Change is still the leading cause of incidents. When change volume goes up an order of magnitude and validation stays flat, one of two things happens. The pipeline backs up, or teams cut corners to clear it. Either way, the risk lands in production and on the people who run it.
The usual fixes (more CI runners, more mocks, test in prod) don't hold at this scale. In a microservices system, the failures that reach production are rarely unit-level. Integration failures across services, queues, and data stores are the ones that slip into production, and the only way to catch them before merge is to run the change in an environment against real dependencies.
Teams build them two ways today: shared or fully duplicated stacks. I'll show why neither can absorb 10x the change, then walk through a third approach that gives every PR its own isolated environment without duplicating the stack.
You'll leave with a framework for what to verify before merge, what to leave to production, and how to keep the error budget intact when agents write most of the code.
Bio
Arjun Iyer is the CEO and Co-founder of Signadot. Signadot helps engineering teams verify code before it merges, at the volume coding agents now produce it. Before Signadot, he spent nearly six years at AppDynamics, where he launched the BusinessIQ business unit and led the company's data science organization. He holds an MS in Computer Science from the University of Illinois Urbana-Champaign.
John Jamie & Dharani Vijayakumar
StackGen
KeynoteThe Reliability Factory: What 109,000 Incidents Say Comes Next for SRE
Abstract
We analyzed 109,000 incidents from the public status pages of 392 companies, along with hundreds of published post-mortems. We started with resolution times and then went under the hood, classifying each incident by failure mode, root cause and the remediation the team applied. The result is one of the most detailed public pictures of how production systems fail and how teams recover them. It also shows that a team's own response practice explains three times more of its resolution time than its industry does.
The data also shows how AI is changing incidents. AI-related incidents grew from under 2% of the total in 2023 to more than 10% in 2026. Incidents caused by AI agents acting on production systems are rising sharply this year. With AI coding pushing change volume higher every quarter, SRE teams will face much more change than current practice was designed to handle.
In the second half of the talk, we describe where we believe SRE goes next. Faster incident response on its own won't keep pace. The focus has to shift from mitigation to prevention, and the mitigation that remains has to become autonomous. Reliability becomes part of an operations factory, in which agentic systems check each change against known failure modes before it ships and mitigate what still breaks within the team's SLOs.
Agents can only do that with a world model of the systems they run and a safe place to test a fix before applying it. Today's service and context graphs map what is deployed and how it connects. A world model also needs intended state, change history, SLOs, upstream providers and the failure patterns from our research. The SRE becomes a reliability architect, who designs this system and decides what agents are allowed to act on. We'll also propose an autonomy index to track how much operational work agents carry and how well they do it.
Attendees will leave with insights from the incident research and a roadmap toward autonomous SRE.
Bio
John Jamie is VP of Marketing at StackGen, where he leads the State of Reliability research program and the launch of StackGen's autonomous SRE capabilities. He has spent about a decade in cloud infrastructure and operations, including roles at Platform9 and at Sedai. StackGen is his second company working on autonomous operations. He presented the State of Reliability findings on LinkedIn Live in August, has spoken on SRE topics at SREday, and is a multi-time SREcon attendee.
Dharani Vijayakumar is Director of Product, Technical GTM at StackGen, where he leads product marketing and technical go-to-market for the Aiden agentic platform. Before rejoining StackGen this year, he worked on AI agents as Executive Director of Product Management for the Data and AI/ML platform at JPMorganChase. In an earlier role at StackGen, he launched a developer self-service AI agent for infrastructure. He has also held product management and engineering roles at HashiCorp, AWS and Cisco, including product work on HashiCorp Cloud Platform and AWS Lambda. He holds an MS in Computer Science and Engineering from Penn State and an MBA from Santa Clara University.