SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.
Companies presenting:
Altinity, Antimetal, AWS, Bank Of JuliusBaer, Cleric, Datadog, DocuSOR, Extraterrestrial Incorporated, FUSSMOBILE, Game Plan Tech, Gatling, GitGuardian, LoopStudio, Palo Alto Networks, Providence, Seismic, Teladoc Health
Agentic tools and AI-assisted code review have made shipping faster than ever. The reliability practices that keep those systems running are under pressure to match.
The systems themselves are changing too. Most teams are already running multiple models in production. When something breaks, the cause is often a rate limit, a prompt update, or a model that changed upstream rather than anything that would show up in a deploy log.
Ajuna Kyaruzi, Manager of SRE and Platform Advocacy at Datadog, will share how we can keep reliability in step with development velocity. The observability signals that help give us the complete picture, how incident response changes when systems can drift without a deployment, and what teams operating at scale are learning.... Read more
Your app works great on your laptop, in the dev environment. Then production hits 10x expected traffic during a marketing campaign and everything falls apart. Or maybe not, instead six months of data accumulates, causing the response times to be painfully slow. Load testing, stress testing, soak testing, and spike testing, they all sound similar, but address completely different problems, and most teams only do one, if at all.
This talk breaks down the essential types of performance testing every app and web developer should understand. Learn when to use each approach, what problems they uncover, and how to integrate them into your development workflow without drowning in complexity. We'll cover real-world scenarios where each testing type saves production systems, helping you choose the right weapon for your performance battles.
What you'll learn:
- The differences between load, stress, soak, and spike testing, and when each matters
- Which performance testing types reveal which production problems, before they happen
- How to integrate performance testing into CI/CD without slowing down development and deployments
- Practical criteria for deciding which tests your application actually needs... Read more
Antimetal builds production engineering agents that operate across hundreds of customer software systems. To be effective, these agents require the same context an experienced engineer has: what depends on what, what changed recently, how failures propagate, what each piece means to the team that runs it. Most of that information is not explicitly documented anywhere.
To solve this problem, we’ve built a world model: a unified, machine-legible representation of any customer’s software systems.
This talk presents the engineering and process behind constructing it, including a provider-agnostic ontology, linking runtime to code, streaming updates in, causal reasoning and inference, among much more. We close with where the model breaks down today and the open problems we are still working through.... Read more
Over the past decade, four trends in enterprise software have reduced the need for large, specialized teams and placed operational capabilities directly in the hands of individual engineers: the shift from monoliths to micro-services, Kubernetes, Site Reliability Engineering, and AI-assisted development. Each trend follows the same pattern. What once required a dedicated specialist or team became a tool or practice that any engineer could own. AI-assisted development accelerates this pattern most dramatically, compressing capabilities previously owned by specialized teams directly into the development workflow. Think about QE and E2E test generation and maintenance, security analysis, error budget analysis, and performance testing - these capabilities historically required large amounts of human and technical capital to execute properly. As AI agents begin to handle increasingly critical aspects of engineering, the trajectory is clear: the agentic software development lifecycle is one where specialized expertise shifts left, resulting in each individual engineer becoming the focal point of shipping products.... Read more
So you finally got your organization to invest in OpenTelemetry. You carefully evaluated observability backends and picked the perfect one. Everything is awesome. Then twelve months later, your costs have skyrocketed and you can’t explain why. What happened? This talk examines how to emit meaningful telemetry while keeping costs under control, by exploring the following: - What to actually instrument - Which metrics to focus on - Pipeline efficiency with OTel Arrow Applying sampling, filtering, and intentional instrumentation to cut down on noise Schema management and validation with tools like Weaver We’ll review the ingredients of a mature observability implementation with OpenTelemetry: one that grows with you instead of overwhelming you. You’ll learn how to apply cost-effective techniques to achieve meaningful observability. Speaker Notes (visible to organizers only) As more organizations embrace OpenTelemetry and mature their Observability practice, we find ourselves coming out of that Observability and OpenTelemetry "honeymoon period". We've gone from, “We’re using OpenTelemetry, therefore we have Observability” to "How do we actually make this work for us?" This talk will equip organizations to build a sustainable and long-lasting observability practice built on OpenTelemetry.... Read more
“Treat servers like cattle, not pets” captured one of the biggest shifts in how we run infrastructure. Going from servers we named with masking tape on the case to infrastructure-as-code to deploy thousands of processes changed the mental model around authentication and authorization. Now, AI agents are forcing another mental shift, as we anthropomorphize these nondeterministic autonomous processes. Many people in the industry are now contemplating how we should handle delegation, security, and identity for these agentic systems. We will look to answer why human and non-human identity management diverged and what that means to various parts of our organizations. This talk will look at the state of authentication and authorization across services today and how we got here. The audience will walk away with a much better sense of where we are headed with machine-to-machine authentication, covering: - SPIFFE/SPIRE - WIMSE and Workload Identity Tokens (WIT) - Security Token Services (STS) - Federated Identities on Cloud Providers... Read more
Product Managers are expected to balance strategy, customer needs, stakeholder communication, prioritization, analytics, documentation, and delivery, often while operating under constant time pressure. As AI capabilities mature, they offer a powerful opportunity to reduce administrative overhead and accelerate many of the information-heavy tasks that consume a PM’s day.
In this session, Greg Spektor explores how AI can augment the product management lifecycle across customer discovery, product strategy, prioritization, documentation, analytics, and delivery coordination. Through practical examples, attendees will learn how AI can synthesize research, analyze customer feedback, generate product documentation, automate reporting, identify risks, and support better decision-making.
The talk also examines the limitations of AI, including hallucinations, bias, governance concerns, and the areas where human judgment remains irreplaceable. Attendees will leave with a practical framework for adopting AI within product organizations and a roadmap for moving from quick wins to more strategic AI-enabled workflows.... Read more
RLHF rewards confident, complete answers. Production incident investigation requires the opposite: withholding judgment, maintaining competing hypotheses, and seeking disconfirming evidence. The training incentive points directly away from the task. At Cleric, we built an AI agent that investigates production incidents. When a cascading failure produces dozens of correlated anomalies, the agent latches onto the loudest signal and builds a coherent narrative around it. The reasoning is sound. The answer is wrong. We taxonomized our agent's production failures and found that premature convergence (the agent finding a plausible answer and stopping) is one of the most common failure modes. Red herrings in production aren't random noise. They're symptoms that look like causes: a downstream service timing out because an upstream database is slow. The timeout is the loudest signal. The database is the root cause. This talk covers the architectural patterns we use to counteract it: forcing multiple competing hypotheses before commitment, requiring evidence that distinguishes between hypotheses rather than confirms the leading one, and a devil's advocate layer that argues against the top hypothesis. We also cover causal inference techniques that make the problem tractable, cascading failures propagate with measurable latency through dependency graphs, and dozens of anomalies collapse into a short causal chain when you follow timestamps instead of severity.... Read more
Most Kubernetes reliability discussions assume stable cloud networking, elastic infrastructure, and always-available control planes. But the real operational challenge starts when Kubernetes has to run closer to the edge: in constrained, bandwidth-limited, intermittently connected, or partially air-gapped environments where standard cloud assumptions no longer hold.
This talk shares practical SRE lessons from operating managed Kubernetes environments across edge and hybrid infrastructure using technologies such as GKE Enterprise, EKS Anywhere, Bottlerocket OS, GitOps workflows, observability tooling, and production-grade operational runbooks. I will cover what changes when clusters are no longer “just in the cloud”: upgrade planning, image and artifact distribution, node OS lifecycle, observability under constrained bandwidth, incident response, storage behavior, and the tradeoffs between automation and safe human control.
The session is intended to be candid and technical, focused on lessons learned rather than theory. Attendees will walk away with a practical mental model for designing and operating Kubernetes platforms in environments where connectivity is imperfect, upgrades require choreography, observability must be selective, and reliability depends as much on operational discipline as it does on tooling.... Read more
An agent is made up of just two parts: the LLM and the harness. But what exactly is a harness? Is it the core “ReAct loop,” tool interface, and context management? Or the instructions, memory layer, skills, and context that you provide? Spoiler alert: the answer is “yes.”
This talk is for anyone who wants more out of their agents, whether you're building agentic applications, using coding agents like Claude Code or Codex, or all of the above. More predictability and reliability. Higher quality output. How to be the engineer in the loop, not the agent babysitter.
As many in the industry move toward “thin” harnesses, this talk argues for a strong “outer” harness: a coordination and control layer outside the core agent itself. Let the model reason and judge. Let the harness make the work repeatable, observable, and safe.... Read more
16:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
In fintech systems, milliseconds don’t just impact performance — they directly translate into financial risk. AI-driven decisions such as fraud detection, credit scoring, and transaction authorization operate under strict latency budgets, often under 100 milliseconds. When these systems slow down, the consequences go beyond user experience: delayed decisions can increase exposure, impact compliance, and lead to measurable financial loss. This talk dives into the engineering challenges of running low-latency AI systems reliably in production. We’ll explore how to define meaningful SLOs for AI workloads, manage tail latency (p95/p99), and design resilient inference pipelines on Kubernetes. Through real-world scenarios, we’ll unpack trade-offs between model accuracy and response time, how bottlenecks emerge under scale, and what typically fails first in high-pressure environments. We’ll also cover observability patterns for AI systems — from detecting latency regressions early to maintaining consistent performance under fluctuating load. If you build or operate systems where delays are unacceptable, this session will provide practical insights to design AI platforms that are not only intelligent, but reliably fast under real-world conditions.... Read more
####Free Kiro Access for Attendees####
Kiro is a new agentic IDE from AWS that turns your ideas into working software through specs, not just prompts. In this hands-on workshop, you'll use Kiro to build a website from scratch using spec-driven development - where requirements, design, and tasks are generated before a single line of code is written. Then you'll use Kiro's agent skills to generate a polished, tailored resume in minutes. No prior Kiro experience needed. Every attendee gets a free Kiro Pro coupon... Read more
We know Ephemeral Environments (EEs) reduce unexpected errors in production, but what happens when an issue inevitably slips through? In this follow-up to “Making Life Easier for SREs with EEs”, we shift our focus from prevention to the high-pressure world of incident response and post-mortems.
This session explores how EEs can help us become faster firefighters and more effective investigators. We will walk through a live Kubernetes microservice demo to experience how EEs serve as a dual-purpose tool when facing real production failures: first, as a firefighting asset to quickly reproduce and mitigate production failures eliminating the “it only breaks in production” problem; and second, as a time machine to rebuild the exact configuration and code state that triggered the failure for deep debugging and root cause analysis leading to permanent fixes. We will also touch upon common concerns, including how to manage the cost and persistence of these temporary environments in a real-world SRE workflow.... Read more
AI agents (MCP servers, LangChain tools, autonomous coding assistants) are becoming production workloads that SREs own, but nobody’s treating their identities, secrets, or privilege escalation paths with the same rigor as human operators. Walk through real patterns: Conjur-brokered credentials for AI agents, JIT privilege grants scoped to agent tasks, and what an incident looks like when an agent’s session token gets over-permissioned.... Read more
Maintaining an open source project is hard. It requires managing a group of people who are largely working for free to build something that other people profit off of, usually distributed across the globe, with limited resources. The whole time you’re doing this, you’re receiving demands from users and businesses alike for features or bug fixes on a timeline that works for them, not you and your (possibly very limited) group of contributors that you can’t exactly order around, since they aren’t being paid. It’s stressful, and it can be overwhelming. When one of these projects is the victim of an attack that takes advantage of the fact that there are only one or two maintainers, or eventually has to shut down due to rising technical debt and falling contributor numbers, the public blame falls on us, not on the businesses that didn’t offer contributors in time.... Read more
Internal Developer Platforms aim to create golden paths, reduce cognitive load, and standardize how teams build and operate software. But there’s a blind spot: documentation. READMEs, onboarding guides, and operational runbooks are often treated as static artifacts. Over time they drift from reality, even as the platform itself evolves. The result is slower onboarding, brittle self-service workflows, and increased support burden on platform teams. In this talk, I’ll explore documentation as an unverified surface within Internal Developer Platforms, and introduce a practical approach to making it executable. By treating markdown instructions and runbooks as workflows that can be executed and validated inside CI or ephemeral environments, platform teams can detect drift automatically. Setup steps, service bootstrapping commands, and health checks become verifiable contracts rather than informal guidance. We’ll cover: • Why documentation drift increases cognitive load and support tickets • How to integrate executable documentation into platform pipelines • How this approach strengthens golden paths and reduces onboarding time • What changes in platform team workflows when documentation becomes testable Document drift is a universal problem, and a costly one. It's time we implmented an actual fix.... Read more
Most AI agents in production process each incident from scratch. No memory of what worked last time, no feedback loops, no adaptation. Stateless tools pretending to be smart. At Cleric, we build an autonomous AI SRE. Getting the agent to diagnose incidents was the easy part. Getting it to retain and apply what it learned from previous ones is where the real engineering problems lie. When one engineer figures out that an OOM spike is always the Redis sidecar, the agent should know that too. And it should know it across teams, across services, across time. We built a three-layer operational memory architecture — semantic, episodic, and procedural — that enables the agent to retain context across investigations and to improve over time. Semantic memory captures what the agent knows about infrastructure and relationships. Episodic memory records specific investigations and their outcomes. Procedural memory encodes the patterns that worked and when to apply them. But captured knowledge decays. The runbook from six months ago references a service that has since been decomposed into three microservices. The fix that worked in Q3 causes a different failure in Q1 because traffic patterns shifted. This talk covers how we detect and handle staleness, the feedback loops that update agent behavior based on resolution outcomes, and what we've learned about building agents that actually get better at their job over time.... Read more
This presentation provides a comprehensive and engaging overview of Service Level Objectives (SLOs) and Service Level Agreements (SLAs), using a scroller game built in with HTML Canvas and Vanilla JS to illustrate the concepts. The three sections of the scroller game cover availability, latency, and error rate. For each metric, the accompanied presentation explains the math behind the metrics in an accessible way. It also reasons why certain percentiles or thresholds may be set based on situation. The talk ends with the exploration of case studies that illustrate how SLOs/SLAs have helped support the core values and product values of companies (ex, through supporting customer-first development and the delivery of high-quality results).... Read more
16:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
The New York Times Building
5620 8th Avenue, 45th Floor,
New York, NY 10018, United States
Sponsors & Partners
Want to become a sponsor? Get in touch!
Heather Thacker
Gatling
Choose Your Weapon: The Performance Testing Arsenal
Abstract
Your app works great on your laptop, in the dev environment. Then production hits 10x expected traffic during a marketing campaign and everything falls apart. Or maybe not, instead six months of data accumulates, causing the response times to be painfully slow. Load testing, stress testing, soak testing, and spike testing, they all sound similar, but address completely different problems, and most teams only do one, if at all.
This talk breaks down the essential types of performance testing every app and web developer should understand. Learn when to use each approach, what problems they uncover, and how to integrate them into your development workflow without drowning in complexity. We'll cover real-world scenarios where each testing type saves production systems, helping you choose the right weapon for your performance battles.
What you'll learn:
- The differences between load, stress, soak, and spike testing, and when each matters
- Which performance testing types reveal which production problems, before they happen
- How to integrate performance testing into CI/CD without slowing down development and deployments
- Practical criteria for deciding which tests your application actually needs
Bio
Developer Advocate with a background in software engineering.
Shreyas Iyer
Antimetal
Building a World Model for Production
Abstract
Antimetal builds production engineering agents that operate across hundreds of customer software systems. To be effective, these agents require the same context an experienced engineer has: what depends on what, what changed recently, how failures propagate, what each piece means to the team that runs it. Most of that information is not explicitly documented anywhere.
To solve this problem, we’ve built a world model: a unified, machine-legible representation of any customer’s software systems.
This talk presents the engineering and process behind constructing it, including a provider-agnostic ontology, linking runtime to code, streaming updates in, causal reasoning and inference, among much more. We close with where the model breaks down today and the open problems we are still working through.
Bio
Shreyas Iyer is the co-founder and CTO of Antimetal, building autonomous agents to run production. He has spent his career in DevOps and SRE at companies big and small. He was most recently a systems engineer at Facebook on their data center infrastructure team.
Ian Miller
Seismic
Shifting Left - Evolution of the Software Development Life Cycle
Abstract
Over the past decade, four trends in enterprise software have reduced the need for large, specialized teams and placed operational capabilities directly in the hands of individual engineers: the shift from monoliths to micro-services, Kubernetes, Site Reliability Engineering, and AI-assisted development. Each trend follows the same pattern. What once required a dedicated specialist or team became a tool or practice that any engineer could own. AI-assisted development accelerates this pattern most dramatically, compressing capabilities previously owned by specialized teams directly into the development workflow. Think about QE and E2E test generation and maintenance, security analysis, error budget analysis, and performance testing - these capabilities historically required large amounts of human and technical capital to execute properly. As AI agents begin to handle increasingly critical aspects of engineering, the trajectory is clear: the agentic software development lifecycle is one where specialized expertise shifts left, resulting in each individual engineer becoming the focal point of shipping products.
Bio
Ian Miller is an Engineering Manager on the Site Reliability Engineering team at Seismic, where he leads multi-cloud infrastructure and reliability initiatives. His career spans renewable energy, software education, and engineering roles at several technology companies.
Josh Lee
Altinity
OpenTelemetry: Playtime Is Over
Abstract
So you finally got your organization to invest in OpenTelemetry. You carefully evaluated observability backends and picked the perfect one. Everything is awesome. Then twelve months later, your costs have skyrocketed and you can’t explain why. What happened? This talk examines how to emit meaningful telemetry while keeping costs under control, by exploring the following: - What to actually instrument - Which metrics to focus on - Pipeline efficiency with OTel Arrow Applying sampling, filtering, and intentional instrumentation to cut down on noise Schema management and validation with tools like Weaver We’ll review the ingredients of a mature observability implementation with OpenTelemetry: one that grows with you instead of overwhelming you. You’ll learn how to apply cost-effective techniques to achieve meaningful observability. Speaker Notes (visible to organizers only) As more organizations embrace OpenTelemetry and mature their Observability practice, we find ourselves coming out of that Observability and OpenTelemetry "honeymoon period". We've gone from, “We’re using OpenTelemetry, therefore we have Observability” to "How do we actually make this work for us?" This talk will equip organizations to build a sustainable and long-lasting observability practice built on OpenTelemetry.
Bio
Josh Lee is an Open Source Advocate at Altinity, focused on ClickHouse, Kubernetes, and OpenTelemetry. With more than 20 years of software engineering experience, he has worked across developer advocacy, product management, observability, and full stack development. Josh is also an organizer of the Open Source Analytics Conference and a frequent speaker in the cloud native and observability community.
Dwayne McDaniel
GitGuardian
From Pets To Cattle To Agents: Evolving Identity And Security For Workloads
Abstract
“Treat servers like cattle, not pets” captured one of the biggest shifts in how we run infrastructure. Going from servers we named with masking tape on the case to infrastructure-as-code to deploy thousands of processes changed the mental model around authentication and authorization. Now, AI agents are forcing another mental shift, as we anthropomorphize these nondeterministic autonomous processes. Many people in the industry are now contemplating how we should handle delegation, security, and identity for these agentic systems. We will look to answer why human and non-human identity management diverged and what that means to various parts of our organizations. This talk will look at the state of authentication and authorization across services today and how we got here. The audience will walk away with a much better sense of where we are headed with machine-to-machine authentication, covering: - SPIFFE/SPIRE - WIMSE and Workload Identity Tokens (WIT) - Security Token Services (STS) - Federated Identities on Cloud Providers
Bio
Dwayne McDaniel is a Principal Developer Advocate who has been on a mission to "help people figure stuff out" for over a decade. At GitGuardian, he specializes in secrets security and non-human identity governance across cloud and DevOps environments. A frequent speaker at events like DevOpsDays and BSides, he helps security and engineering teams better understand complex issues.
Greg Spektor
Extraterrestrial Incorporated
AI for Product Management: Working Faster, Smarter, and More Strategically
Abstract
Product Managers are expected to balance strategy, customer needs, stakeholder communication, prioritization, analytics, documentation, and delivery, often while operating under constant time pressure. As AI capabilities mature, they offer a powerful opportunity to reduce administrative overhead and accelerate many of the information-heavy tasks that consume a PM’s day.
In this session, Greg Spektor explores how AI can augment the product management lifecycle across customer discovery, product strategy, prioritization, documentation, analytics, and delivery coordination. Through practical examples, attendees will learn how AI can synthesize research, analyze customer feedback, generate product documentation, automate reporting, identify risks, and support better decision-making.
The talk also examines the limitations of AI, including hallucinations, bias, governance concerns, and the areas where human judgment remains irreplaceable. Attendees will leave with a practical framework for adopting AI within product organizations and a roadmap for moving from quick wins to more strategic AI-enabled workflows.
Bio
Greg Spektor is an AI delivery leader and product strategist specializing in enterprise AI implementation, governance, and operationalization. He has led AI and product initiatives across healthcare, manufacturing, and technology, helping organizations scale AI from pilot projects to production systems that deliver business value.
Willem Pienaar
Cleric
Perfect Reasoning, Wrong Answer
Abstract
RLHF rewards confident, complete answers. Production incident investigation requires the opposite: withholding judgment, maintaining competing hypotheses, and seeking disconfirming evidence. The training incentive points directly away from the task. At Cleric, we built an AI agent that investigates production incidents. When a cascading failure produces dozens of correlated anomalies, the agent latches onto the loudest signal and builds a coherent narrative around it. The reasoning is sound. The answer is wrong. We taxonomized our agent's production failures and found that premature convergence (the agent finding a plausible answer and stopping) is one of the most common failure modes. Red herrings in production aren't random noise. They're symptoms that look like causes: a downstream service timing out because an upstream database is slow. The timeout is the loudest signal. The database is the root cause. This talk covers the architectural patterns we use to counteract it: forcing multiple competing hypotheses before commitment, requiring evidence that distinguishes between hypotheses rather than confirms the leading one, and a devil's advocate layer that argues against the top hypothesis. We also cover causal inference techniques that make the problem tractable, cascading failures propagate with measurable latency through dependency graphs, and dozens of anomalies collapse into a short causal chain when you follow timestamps instead of severity.
Bio
Willem Pienaar is the Co-Founder and CTO of Cleric, where he builds autonomous AI agents that operate inside production infrastructure. He created Feast, the widely adopted open source feature store powering ML systems at companies including Cloudflare, Discord, Robinhood, Salesforce, and Shopify. Before founding Cleric, Willem led open source engineering at Tecton and built the ML platform team at Gojek, one of Southeast Asia's largest ride-hailing companies. He holds an MS in Computer Science from Georgia Tech.
Edward Rodriguez
FUSSMOBILE
Kubernetes at the Edge: SRE Lessons from Disconnected Clusters
Abstract
Most Kubernetes reliability discussions assume stable cloud networking, elastic infrastructure, and always-available control planes. But the real operational challenge starts when Kubernetes has to run closer to the edge: in constrained, bandwidth-limited, intermittently connected, or partially air-gapped environments where standard cloud assumptions no longer hold.
This talk shares practical SRE lessons from operating managed Kubernetes environments across edge and hybrid infrastructure using technologies such as GKE Enterprise, EKS Anywhere, Bottlerocket OS, GitOps workflows, observability tooling, and production-grade operational runbooks. I will cover what changes when clusters are no longer “just in the cloud”: upgrade planning, image and artifact distribution, node OS lifecycle, observability under constrained bandwidth, incident response, storage behavior, and the tradeoffs between automation and safe human control.
The session is intended to be candid and technical, focused on lessons learned rather than theory. Attendees will walk away with a practical mental model for designing and operating Kubernetes platforms in environments where connectivity is imperfect, upgrades require choreography, observability must be selective, and reliability depends as much on operational discipline as it does on tooling.
Bio
Edward Rodriguez is the founder of FUSSMOBILE, a Kubernetes and cloud managed services provider focused on production-grade platform operations, SRE practices, and hybrid infrastructure. He and his team work with enterprise Kubernetes environments across cloud and edge deployments, including EKS, GKE Enterprise, Rancher, EKS Anywhere, Container based OS, GitOps, observability, and operational automation. His work focuses on helping organizations run reliable Kubernetes platforms in complex real-world environments where availability, upgrade safety, and operational discipline matter.
Paul Caplan
Teladoc Health
Agent Harnesses: From Slot Machines to Safety Nets
Abstract
An agent is made up of just two parts: the LLM and the harness. But what exactly is a harness? Is it the core “ReAct loop,” tool interface, and context management? Or the instructions, memory layer, skills, and context that you provide? Spoiler alert: the answer is “yes.”
This talk is for anyone who wants more out of their agents, whether you're building agentic applications, using coding agents like Claude Code or Codex, or all of the above. More predictability and reliability. Higher quality output. How to be the engineer in the loop, not the agent babysitter.
As many in the industry move toward “thin” harnesses, this talk argues for a strong “outer” harness: a coordination and control layer outside the core agent itself. Let the model reason and judge. Let the harness make the work repeatable, observable, and safe.
Bio
Paul Caplan is a Principal Engineer on the Developer Experience team at Teladoc Health, where he builds SDLC agents and helps drive adoption of AI tooling across the engineering organization. By night, he is building Codagent, an open-source harness toolkit for AI coding agents, and writing the “Agents, Harnessed” newsletter.
Anisha Manoharan
Bank Of JuliusBaer
When Milliseconds Cost Millions: Designing Low-Latency AI Systems in Fintech
Abstract
In fintech systems, milliseconds don’t just impact performance — they directly translate into financial risk. AI-driven decisions such as fraud detection, credit scoring, and transaction authorization operate under strict latency budgets, often under 100 milliseconds. When these systems slow down, the consequences go beyond user experience: delayed decisions can increase exposure, impact compliance, and lead to measurable financial loss. This talk dives into the engineering challenges of running low-latency AI systems reliably in production. We’ll explore how to define meaningful SLOs for AI workloads, manage tail latency (p95/p99), and design resilient inference pipelines on Kubernetes. Through real-world scenarios, we’ll unpack trade-offs between model accuracy and response time, how bottlenecks emerge under scale, and what typically fails first in high-pressure environments. We’ll also cover observability patterns for AI systems — from detecting latency regressions early to maintaining consistent performance under fluctuating load. If you build or operate systems where delays are unacceptable, this session will provide practical insights to design AI platforms that are not only intelligent, but reliably fast under real-world conditions.
Bio
Anisha Manoharan is a Senior Platform Engineer specializing in cloud-native systems, AI infrastructure, and scalable platform engineering. She has extensive experience building and operating distributed systems using Kubernetes, enabling teams to move from experimental models to reliable, production-grade AI platforms. With a background in SRE and DevOps, Anisha has worked on automation, monitoring, and resilient infrastructure in regulated environments, bringing a strong focus on reliability, scalability, and real-world impact. She is passionate about bridging the gap between machine learning and production systems, sharing practical insights on how modern engineering teams can successfully deploy and scale AI in real-world applications.
Aaron Hunter & Saurabh Rob Dahal
AWS
1h Hands-On with Kiro: Use AI Agents to Build Your Website and Optimize Your Resume
Abstract
Free Kiro Access for Attendees####
Kiro is a new agentic IDE from AWS that turns your ideas into working software through specs, not just prompts. In this hands-on workshop, you'll use Kiro to build a website from scratch using spec-driven development - where requirements, design, and tasks are generated before a single line of code is written. Then you'll use Kiro's agent skills to generate a polished, tailored resume in minutes. No prior Kiro experience needed. Every attendee gets a free Kiro Pro coupon
Bio
Aaron Hunter is a Principal Developer Advocate at AWS based in Frisco, TX. With over 15 years of experience spanning system administration, networking, and training, he brings more than a decade of cloud expertise to help Engineers, Developers, Builders, and tech enthusiasts master AWS technologies. His philosophy is simple: always be learning something new, and have fun while doing it! Aaron shares his knowledge through workshops, online courses, mentoring, and live streaming – making complex cloud concepts accessible and enjoyable. When he's not building in the cloud, you'll find him exploring craft beer scenes in new cities or teaming up with friends in Marvel Rivals on his PlayStation, because even superheroes need good teammates!
Saurabh Rob Dahal: career change from Biomedical field to Tech. Worked as a Solutions Engineer at Oracle, Software Developer and Coding Bootcamp instructor. Currently working as a Developer Advocate at AWS specializing in AI Engineering, software development, and Agents.
Marcos Novelli Harispe
LoopStudio
Beyond Prevention: Mastering Incident Response and Post-Mortems with Ephemeral Environments
Abstract
We know Ephemeral Environments (EEs) reduce unexpected errors in production, but what happens when an issue inevitably slips through? In this follow-up to “Making Life Easier for SREs with EEs”, we shift our focus from prevention to the high-pressure world of incident response and post-mortems.
This session explores how EEs can help us become faster firefighters and more effective investigators. We will walk through a live Kubernetes microservice demo to experience how EEs serve as a dual-purpose tool when facing real production failures: first, as a firefighting asset to quickly reproduce and mitigate production failures eliminating the “it only breaks in production” problem; and second, as a time machine to rebuild the exact configuration and code state that triggered the failure for deep debugging and root cause analysis leading to permanent fixes. We will also touch upon common concerns, including how to manage the cost and persistence of these temporary environments in a real-world SRE workflow.
Bio
Marcos Novelli Harispe is a software engineer specializing in AWS, serverless architectures, and data-driven solutions. He has delivered scalable notification systems, analytics platforms, and automation tools across freelance and full-time roles, with a focus on reliability, efficiency, and measurable business impact.
Zachary Gruenberg
Palo Alto Networks
Who Watches the AI Agents? SRE for Non-Human Identity at Scale
Abstract
AI agents (MCP servers, LangChain tools, autonomous coding assistants) are becoming production workloads that SREs own, but nobody’s treating their identities, secrets, or privilege escalation paths with the same rigor as human operators. Walk through real patterns: Conjur-brokered credentials for AI agents, JIT privilege grants scoped to agent tasks, and what an incident looks like when an agent’s session token gets over-permissioned.
Bio
Zach is an AI Subject Matter Expert at Palo Alto via CyberArk, where he works at the intersection of identity security and site reliability — focusing on non-human identity governance, AI agent privilege controls, and secrets management for cloud-native workloads. Day-to-day, he builds secure credential brokering for AI agents, integrates machine identity into Kubernetes platforms using SPIFFE and Conjur, and helps organizations treat identity as a reliability problem, not just a security checkbox.
Christopher Tineo
Game Plan Tech
Free Software Isn't Gratis: why companies should treat Open Source as part of their Infrastructure
Abstract
Maintaining an open source project is hard. It requires managing a group of people who are largely working for free to build something that other people profit off of, usually distributed across the globe, with limited resources. The whole time you’re doing this, you’re receiving demands from users and businesses alike for features or bug fixes on a timeline that works for them, not you and your (possibly very limited) group of contributors that you can’t exactly order around, since they aren’t being paid. It’s stressful, and it can be overwhelming. When one of these projects is the victim of an attack that takes advantage of the fact that there are only one or two maintainers, or eventually has to shut down due to rising technical debt and falling contributor numbers, the public blame falls on us, not on the businesses that didn’t offer contributors in time.
Bio
I'm passionate about everything related to the Open Source and Cloud Native ecosystem. I'm results-driven and passionate about this industry. I truly believe in the value of giving back to others, so when I'm not working or studying, you will see me volunteering for the communities around my area or Latin America.
Omari Gaskins Jr
DocuSOR
Your Outdated Docs are Costly: Why You should be writing tests for your docs
Abstract
Internal Developer Platforms aim to create golden paths, reduce cognitive load, and standardize how teams build and operate software. But there’s a blind spot: documentation. READMEs, onboarding guides, and operational runbooks are often treated as static artifacts. Over time they drift from reality, even as the platform itself evolves. The result is slower onboarding, brittle self-service workflows, and increased support burden on platform teams. In this talk, I’ll explore documentation as an unverified surface within Internal Developer Platforms, and introduce a practical approach to making it executable. By treating markdown instructions and runbooks as workflows that can be executed and validated inside CI or ephemeral environments, platform teams can detect drift automatically. Setup steps, service bootstrapping commands, and health checks become verifiable contracts rather than informal guidance. We’ll cover: • Why documentation drift increases cognitive load and support tickets • How to integrate executable documentation into platform pipelines • How this approach strengthens golden paths and reduces onboarding time • What changes in platform team workflows when documentation becomes testable Document drift is a universal problem, and a costly one. It's time we implmented an actual fix.
Bio
Omari Gaskins Jr. is the creator of DocuSOR, the first platform that executes and validates documentation to prevent drift and accelerate developer onboarding. As a Software Engineer at JPMorgan Chase for the last 5 years, he has worked on hundreds of different applications, with on prem and public cloud integrations, managing the scale that comes with working for a large global enterprise.
Shahram Anver
Cleric
Agentic Learning Patterns for SRE
Abstract
Most AI agents in production process each incident from scratch. No memory of what worked last time, no feedback loops, no adaptation. Stateless tools pretending to be smart. At Cleric, we build an autonomous AI SRE. Getting the agent to diagnose incidents was the easy part. Getting it to retain and apply what it learned from previous ones is where the real engineering problems lie. When one engineer figures out that an OOM spike is always the Redis sidecar, the agent should know that too. And it should know it across teams, across services, across time. We built a three-layer operational memory architecture — semantic, episodic, and procedural — that enables the agent to retain context across investigations and to improve over time. Semantic memory captures what the agent knows about infrastructure and relationships. Episodic memory records specific investigations and their outcomes. Procedural memory encodes the patterns that worked and when to apply them. But captured knowledge decays. The runbook from six months ago references a service that has since been decomposed into three microservices. The fix that worked in Q3 causes a different failure in Q1 because traffic patterns shifted. This talk covers how we detect and handle staleness, the feedback loops that update agent behavior based on resolution outcomes, and what we've learned about building agents that actually get better at their job over time.
Bio
Shahram Anver is the Co-Founder and CEO of Cleric, where he's building an autonomous AI SRE that investigates and resolves production incidents 24/7. Before Cleric, Shahram led engineering for MLOps, container deployment, and FinOps platforms at Gojek, Southeast Asia's largest super-app, managing infrastructure handling millions of daily transactions across hundreds of microservices. He previously built TripAdvisor's first ML automated bidding system, scaling it to manage tens of millions in annual ad spend, and co-founded DataCue, an ML-driven e-commerce personalization platform.
Vivek Shah
Providence
99.9% Fun: A Game-Based Guide to SLOs and SLAs
Abstract
This presentation provides a comprehensive and engaging overview of Service Level Objectives (SLOs) and Service Level Agreements (SLAs), using a scroller game built in with HTML Canvas and Vanilla JS to illustrate the concepts. The three sections of the scroller game cover availability, latency, and error rate. For each metric, the accompanied presentation explains the math behind the metrics in an accessible way. It also reasons why certain percentiles or thresholds may be set based on situation. The talk ends with the exploration of case studies that illustrate how SLOs/SLAs have helped support the core values and product values of companies (ex, through supporting customer-first development and the delivery of high-quality results).
Bio
Vivek Shah is a Senior Software Engineer at Providence Health with over a decade of experience building full-stack applications across healthcare, retail, and industrial technology. His work spans LLM-powered features, cloud infrastructure, CI/CD pipelines, and scalable web platforms. Previously, he held engineering roles at Starbucks, Falkonry, and GE, contributing to large-scale infrastructure, IoT dashboards, and predictive SaaS systems.
Intro
KeynoteIntro by Mark Pawlikowski
Abstract
Walking everyone through the agenda, warming up for a great event !
Ajuna Kyaruzi
Datadog
KeynoteObservability and SRE in the AI Era
Abstract
Agentic tools and AI-assisted code review have made shipping faster than ever. The reliability practices that keep those systems running are under pressure to match.
The systems themselves are changing too. Most teams are already running multiple models in production. When something breaks, the cause is often a rate limit, a prompt update, or a model that changed upstream rather than anything that would show up in a deploy log.
Ajuna Kyaruzi, Manager of SRE and Platform Advocacy at Datadog, will share how we can keep reliability in step with development velocity. The observability signals that help give us the complete picture, how incident response changes when systems can drift without a deployment, and what teams operating at scale are learning.
Bio
Ajuna Kyaruzi leads the SRE & Platform Advocacy team at Datadog. She cares about using software to help people sustainably run large-scale systems, focusing on Incident Response and SLOs. She loves community building and volunteers with multiple mentorship programs aimed at helping early career folks break into tech, and ensuring they have successful careers. Previously she worked at Google as a Software Engineer on Google Maps and as a Site Reliability Engineer on Google Cloud.