SREday

Site Reliability, DevOps and Cloud

February 26, 2026 Harness, New York, US

1
Day
10+
Speakers
1
Track
80+
Attendees

SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.

Companies presenting:

Causely, Chainguard, Cribl, Harness, ilert, Imply, LoopStudio, MetalBear, NeuBird AI, PagerDuty, Redis, Resizes, solo.io, Traversal, Wild Moose, Xurrent

Topics so far:

This is a past event, what's next?

Schedule

February 26, 2026 • single track • 9AM - 6:30PM • NYC, in-person
view as table
main room • Track 1

09:00

Dewan Ahmed

KeynoteSecure by Default: Building Confidence in AI-Driven Delivery

HarnessWatch
The fastest way to break trust in DevSecOps is to automate insecurity at scale. As AI takes a central role in our pipelines, it is time to rethink what "secure by default" really means. In this keynote, Dewan Ahmed will challenge the audience to look beyond vulnerability scanners and compliance gates. He will share a vision for intelligent security by design, where native intelligence within the delivery platform detects not only vulnerable code but risky delivery behavior such as misconfigured environments, suspicious artifact provenance, and drift between source and runtime. You will walk away with a framework for balancing automation with human oversight and examples from Harness’ work on building verifiable, auditable, AI-native delivery systems. In the new world of DevSecOps, safety is not a step; it is an outcome we continuously learn to improve.... Read more

09:30

Birol Yildiz

KeynoteWhen Incidents Fix Themselves: AI SRE in action

ilertWatch
The next evolution of incident response isn’t faster alerts, it’s autonomous resolution. Join ilert CEO Birol Yildiz as he shows how AI SRE agents now diagnose and remediate outages without waking anyone up. Learn how these systems combine observability data, deployment context, and code intelligence to restore services in minutes and hand over clean incident reports instead of 3 a.m. pages.... Read more

10:00

Ben Hopp

KeynoteDecoupling Observability for Incident Response at Scale

ImplyWatch
On-call incidents don’t fail because teams lack dashboards. They fail because observability systems slow down under real investigative load. As telemetry volumes grow and retention windows expand, SRE teams are being asked to run deeper, broader investigations—often under time pressure—on platforms that were designed for steady-state monitoring, not bursty incident response. Tightly coupled observability stacks bind storage, compute, and query together, forcing teams to overprovision infrastructure, limit retention, or accept degraded performance during incidents. In this talk, we’ll explore why decoupling observability architectures is becoming essential for SRE teams operating at scale. Using a real incident investigation workflow, we’ll break down how separating data storage, compute, and interaction layers allows teams to keep fast, reliable monitoring while elastically scaling investigations when incidents occur. We’ll connect these patterns to lessons learned in other data-intensive systems, but stay grounded in the day-to-day realities of on-call life: faster root cause analysis, fewer tradeoffs during incidents, and observability systems that hold up when you need them most.... Read more

10:30

AJ Mejorado

KeynoteThe Best CVE Is the One You Never Patch: Removing Vulnerability Toil for SREs

ChainguardWatch
For most SRE teams, CVE remediation has become a never-ending source of toil—rebuilds, backports, emergency releases, and late-night firefighting. This talk explores how Chainguard flips that model by removing vulnerability remediation from the SRE workload entirely. We’ll look at how secure-by-default images, continuous rebuilds, and automated patch pipelines eliminate the need for teams to chase CVEs in production. Instead of reacting to vulnerabilities, SREs inherit artifacts that are already patched, minimal, and ready to run—freeing them to focus on reliability, performance, and scale. This session is about shifting CVE response left, shrinking the attack surface, and giving SRE teams their time back—because patching shouldn’t be an on-call responsibility.... Read more

11:00

Michael Levan

KeynoteObserving Agentic/MCP Traffic & Keeping Cost Low

solo.ioWatch
You use an LLM, ask it some questions, have it write a little code for you, and before you know it, you're hit with a massive bill. Out of the box, there's zero way to see token usage, cost, and an overall ability to control it. In this talk, you'll learn exactly how to do that.... Read more

11:30

Endre Sara

KeynoteThe Missing Layer in AI-Driven Reliability

Causely Watch
AI agents are entering incident response with powerful capabilities—detecting anomalies, summarizing logs, and suggesting fixes. Some even execute changes autonomously. The promise is compelling: fully autonomous resolution of production incidents in complex, constantly changing distributed systems. But before we hand control to agents, we must ask a fundamental question: What kind of understanding must exist before we let agents act? This talk argues that the real bottleneck is not model intelligence, but the absence of a continuously updated causal model of the system itself. ... Read more

12:00

Francois Martel

KeynoteThe AI SRE Landscape: From LLMs to Multi-Agent Ecosystems

NeuBird AIWatch
With more than 15 vendors now claiming AI powered SRE capabilities, engineering leaders are facing a deafening amount of noise. Teams are asking hard questions. Is Datadog Bits AI the same thing as an AI agent? AWS just launched a DevOps Agent. Do we still need a vendor? Our ServiceNow rep says their AI can handle it all. How do you cut through the hype? This talk provides a practical framework for understanding the AI systems transforming Site Reliability Engineering, including foundation LLMs, RAG systems, chat based tools, agentic AI, and multi agent ecosystems. We map the current vendor landscape across five categories: Incident Coordinators, AI Investigators, Observability Native AI, Cloud Provider Agents, and Coding Agents. You will see what each category can actually do today and why context engineering, not model intelligence, is the real differentiator. The session includes a live walkthrough of a multi agent workflow handling a production incident end to end. It starts with intelligent detection and root cause analysis, continues through automated remediation using a coding agent, and finishes with verified resolution. Humans stay in the loop only where it matters. The workflow persists for hours, coordinates across tools, and retains full context throughout. You will leave with a clear mental model for evaluating AI SRE tools and a phased adoption roadmap. The key takeaways are straightforward. Know what type of AI you are buying. Invest in context architecture over model hype. Build a layered stack rather than a monolith.... Read more

12:30

Lunch & networking

Main lobby

13:30

Yasmin Dunsky & Avner Yaacov

What It Takes To Build Reliability as a System

Wild Moose & RedisWatch
In this talk, Redis and Wild Moose share how they're approaching AI SRE as something that can be designed, automated, and continuously improved. We'll explore how to build AI SRE systems that teams can actually trust in production by applying the same standards used for observability systems: fast, testable, configurable, and transparent. Using a real-world example from Redis, we'll show how debugging agents turn tribal knowledge into structured investigation workflows and automate critical parts of the operational flow—creating a reliability system that improves with every incident. We'll also share why treating agents like production code, with regression testing and validation, is essential for long-term scale. You'll walk away with an understanding of the cultural shift needed to fully leverage AI for incident management.... Read more

14:00

Mandi Walls

Reevaluating Post-Incident Reviews

PagerDutyWatch
Post-Incident Review. Postmortem. Incident Report. Whatever you call them, they take time and resources, and sometimes you’re not even sure if anyone reads them. The Post-Incident Review is a story. Maybe it’s a bit of a mystery, maybe it’s a feel-good story of redemption. Maybe it’s a buddy comedy. We produce these reports so that other folks in our organization can learn from the things we learned and hopefully not repeat our same mistakes. Collecting the data and creating the story are work. We’ll talk about how PagerDuty has changed and adapted our methodology over time to produce better reviews that help more engineers get more out of the process and the assets that are produced.... Read more

14:30

Jason Schechner

What Makes an AI SRE Trustworthy? An SRE’s Take

TraversalWatch
On-call is stressful enough without worrying whether AI hallucinated the root cause. In this talk, an SRE who's been handling production incidents since before the role had a name will share what it's like to build an AI system that triages incidents and monitors system health — and why it's not as simple as throwing telemetry at an LLM. He'll walk through how he worked with AI engineers to transfer decades of troubleshooting instincts — and where the AI deliberately diverges (like parallelizing hypothesis exploration). He'll also discuss how we build trust through evaluation — comparing accuracy, latency, and reasoning against ground truth from SREs. This is a practitioner's view of what it takes to build an AI investigator that engineers can rely on when things are on fire at 3am.... Read more

15:00

Maria Garcia Garcia & Lucia Lopez Barrero

Responsible Kubernetes: clusters that don’t kill the planet (or your budget)

ResizesWatch
This talk focuses on applying FinOps and sustainability (GreenOps) principles in Kubernetes, addressing very common challenges such as cluster overprovisioning, inefficient resource usage, lack of cost visibility, and the environmental impact of always-on infrastructure. We share real-world practices, common mistakes, and practical lessons on how to run Kubernetes in a more efficient, cost-aware, and sustainable way. We are two engineers who started working in Platform Engineering about a year ago, and we believe this perspective brings real value: we speak from recent, hands-on experience, from what we’ve encountered operating real clusters, what didn’t work, what did, and what we’ve learned along the way while trying to make Kubernetes more responsible in day-to-day operations.... Read more

15:30

Bryan Patton

The "Supply Chain" Mindset: Why I hire Logistical Engineers, Not Just SREs

XurrentWatch
This talk covers how we hire and run SRE, DevOps, and TechOps teams by treating infrastructure like a supply chain. Looking at systems this way makes bottlenecks, risk, and capacity limits easier to see before they cause outages. I’ll explain what we look for when hiring, how we expect engineers to operate day to day, and how this approach helps us define the right metrics early instead of after incidents.... Read more

16:00

Networking and sponsor crawl

Main lobby

16:30

Adna Zujo Lakisic

Honey, the Audience Broke My App: Reproduce & Fix Live in Kubernetes with mirrord

MetalBearWatch
A room of engineers breaks my live app and we fix it together. They scan a QR code which triggers a real failure in a Kubernetes staging environment. Together we reproduce, trace, patch, and validate the fix with kubectl debug and mirrord. A real breakage repaired live.... Read more

17:00

Leon Adato

Looking back at a lifetime of poor tech choices

CriblWatch
Watching your technology choice go down the drain - because of broken promises, vendor implosion, or obsolescence - can feel like a career-ending experience. But in my experience there's usually a lot of good that comes out of those seemingly bad tech choices.... Read more

17:30

Marcos Novelli Harispe

Making Life Easier for SREs with Ephemeral Environments

LoopStudioWatch
SREs often work under constant pressure, dealing with incidents and operational overhead. This talk explores how ephemeral environments can help make day-to-day SRE work easier, starting small and without disruptive changes.... Read more

18:00

Happy Hour by Imply - grab a beer!

Main lobby

18:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time main room
09:00 Keynote: Secure by Default: Building Confidence in AI-Driven Delivery
Dewan Ahmed • Harness
09:30 Keynote: When Incidents Fix Themselves: AI SRE in action
Birol Yildiz • ilert
10:00 Keynote: Decoupling Observability for Incident Response at Scale
Ben Hopp • Imply
10:30 Keynote: The Best CVE Is the One You Never Patch: Removing Vulnerability Toil for SREs
AJ Mejorado • Chainguard
11:00 Keynote: Observing Agentic/MCP Traffic & Keeping Cost Low
Michael Levan • solo.io
11:30 Keynote: The Missing Layer in AI-Driven Reliability
Endre Sara • Causely
12:00 Keynote: The AI SRE Landscape: From LLMs to Multi-Agent Ecosystems
Francois Martel • NeuBird AI
12:30 Lunch & networking
13:30 What It Takes To Build Reliability as a System
Yasmin Dunsky & Avner Yaacov • Wild Moose & Redis
14:00 Reevaluating Post-Incident Reviews
Mandi Walls • PagerDuty
14:30 What Makes an AI SRE Trustworthy? An SRE’s Take
Jason Schechner • Traversal
15:00 Responsible Kubernetes: clusters that don’t kill the planet (or your budget)
Maria Garcia Garcia & Lucia Lopez Barrero • Resizes
15:30 The "Supply Chain" Mindset: Why I hire Logistical Engineers, Not Just SREs
Bryan Patton • Xurrent
16:00 Networking and sponsor crawl
16:30 Honey, the Audience Broke My App: Reproduce & Fix Live in Kubernetes with mirrord
Adna Zujo Lakisic • MetalBear
17:00 Looking back at a lifetime of poor tech choices
Leon Adato • Cribl
17:30 Making Life Easier for SREs with Ephemeral Environments
Marcos Novelli Harispe • LoopStudio
18:00 Happy Hour by Imply - grab a beer!
18:30 Wrap up

Speakers

Adna Zujo Lakisic
MetalBear
AJ Mejorado
Chainguard
Ben Hopp
Imply
Birol Yildiz
ilert
Bryan Patton
Xurrent
Dewan Ahmed
Harness
Endre Sara
Causely
Francois Martel
NeuBird AI
Jason Schechner
Traversal
Leon Adato
Cribl
Mandi Walls
PagerDuty
Marcos Novelli Harispe
LoopStudio
Maria Garcia Garcia
& Lucia Lopez Barrero
Resizes
Michael Levan
solo.io
Yasmin Dunsky
& Avner Yaacov
Wild Moose & Redis

Venue

Harness.io - New York

45 W 45th St, 7th floor, New York,
NY 10036, US

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one