SREday is a worldwide series of community events for engineers who build, ship and run modern software systems. Across cities around the world, we bring together people working in reliability, cloud, DevOps, observability and production engineering to share real-world experience, connect with their local community and explore how these disciplines are evolving in the age of AI.
The fastest way to break trust in DevSecOps is to automate insecurity at scale. As AI takes a central role in our pipelines, it is time to rethink what "secure by default" really means.
In this keynote, Dewan Ahmed will challenge the audience to look beyond vulnerability scanners and compliance gates. He will share a vision for intelligent security by design, where native intelligence within the delivery platform detects not only vulnerable code but risky delivery behavior such as misconfigured environments, suspicious artifact provenance, and drift between source and runtime.
You will walk away with a framework for balancing automation with human oversight and examples from Harness’ work on building verifiable, auditable, AI-native delivery systems. In the new world of DevSecOps, safety is not a step; it is an outcome we continuously learn to improve.... Read more
The next evolution of incident response isn’t faster alerts, it’s autonomous resolution. Join ilert CEO Birol Yildiz as he shows how AI SRE agents now diagnose and remediate outages without waking anyone up. Learn how these systems combine observability data, deployment context, and code intelligence to restore services in minutes and hand over clean incident reports instead of 3 a.m. pages.... Read more
On-call incidents don’t fail because teams lack dashboards. They fail because observability systems slow down under real investigative load.
As telemetry volumes grow and retention windows expand, SRE teams are being asked to run deeper, broader investigations—often under time pressure—on platforms that were designed for steady-state monitoring, not bursty incident response. Tightly coupled observability stacks bind storage, compute, and query together, forcing teams to overprovision infrastructure, limit retention, or accept degraded performance during incidents.
In this talk, we’ll explore why decoupling observability architectures is becoming essential for SRE teams operating at scale. Using a real incident investigation workflow, we’ll break down how separating data storage, compute, and interaction layers allows teams to keep fast, reliable monitoring while elastically scaling investigations when incidents occur.
We’ll connect these patterns to lessons learned in other data-intensive systems, but stay grounded in the day-to-day realities of on-call life: faster root cause analysis, fewer tradeoffs during incidents, and observability systems that hold up when you need them most.... Read more
For most SRE teams, CVE remediation has become a never-ending source of toil—rebuilds, backports, emergency releases, and late-night firefighting. This talk explores how Chainguard flips that model by removing vulnerability remediation from the SRE workload entirely.
We’ll look at how secure-by-default images, continuous rebuilds, and automated patch pipelines eliminate the need for teams to chase CVEs in production. Instead of reacting to vulnerabilities, SREs inherit artifacts that are already patched, minimal, and ready to run—freeing them to focus on reliability, performance, and scale.
This session is about shifting CVE response left, shrinking the attack surface, and giving SRE teams their time back—because patching shouldn’t be an on-call responsibility.... Read more
You use an LLM, ask it some questions, have it write a little code for you, and before you know it, you're hit with a massive bill. Out of the box, there's zero way to see token usage, cost, and an overall ability to control it. In this talk, you'll learn exactly how to do that.... Read more
AI agents are entering incident response with powerful capabilities—detecting anomalies, summarizing logs, and suggesting fixes. Some even execute changes autonomously. The promise is compelling: fully autonomous resolution of production incidents in complex, constantly changing distributed systems.
But before we hand control to agents, we must ask a fundamental question: What kind of understanding must exist before we let agents act?
This talk argues that the real bottleneck is not model intelligence, but the absence of a continuously updated causal model of the system itself. ... Read more
With more than 15 vendors now claiming AI powered SRE capabilities, engineering leaders are facing a deafening amount of noise. Teams are asking hard questions. Is Datadog Bits AI the same thing as an AI agent? AWS just launched a DevOps Agent. Do we still need a vendor? Our ServiceNow rep says their AI can handle it all. How do you cut through the hype?
This talk provides a practical framework for understanding the AI systems transforming Site Reliability Engineering, including foundation LLMs, RAG systems, chat based tools, agentic AI, and multi agent ecosystems. We map the current vendor landscape across five categories: Incident Coordinators, AI Investigators, Observability Native AI, Cloud Provider Agents, and Coding Agents. You will see what each category can actually do today and why context engineering, not model intelligence, is the real differentiator.
The session includes a live walkthrough of a multi agent workflow handling a production incident end to end. It starts with intelligent detection and root cause analysis, continues through automated remediation using a coding agent, and finishes with verified resolution. Humans stay in the loop only where it matters. The workflow persists for hours, coordinates across tools, and retains full context throughout.
You will leave with a clear mental model for evaluating AI SRE tools and a phased adoption roadmap. The key takeaways are straightforward. Know what type of AI you are buying. Invest in context architecture over model hype. Build a layered stack rather than a monolith.... Read more
In this talk, Redis and Wild Moose share how they're approaching AI SRE as something that can be designed, automated, and continuously improved. We'll explore how to build AI SRE systems that teams can actually trust in production by applying the same standards used for observability systems: fast, testable, configurable, and transparent.
Using a real-world example from Redis, we'll show how debugging agents turn tribal knowledge into structured investigation workflows and automate critical parts of the operational flow—creating a reliability system that improves with every incident. We'll also share why treating agents like production code, with regression testing and validation, is essential for long-term scale. You'll walk away with an understanding of the cultural shift needed to fully leverage AI for incident management.... Read more
Post-Incident Review. Postmortem. Incident Report. Whatever you call them, they take time and resources, and sometimes you’re not even sure if anyone reads them.
The Post-Incident Review is a story. Maybe it’s a bit of a mystery, maybe it’s a feel-good story of redemption. Maybe it’s a buddy comedy. We produce these reports so that other folks in our organization can learn from the things we learned and hopefully not repeat our same mistakes.
Collecting the data and creating the story are work. We’ll talk about how PagerDuty has changed and adapted our methodology over time to produce better reviews that help more engineers get more out of the process and the assets that are produced.... Read more
On-call is stressful enough without worrying whether AI hallucinated the root cause.
In this talk, an SRE who's been handling production incidents since before the role had a name will share what it's like to build an AI system that triages incidents and monitors system health — and why it's not as simple as throwing telemetry at an LLM.
He'll walk through how he worked with AI engineers to transfer decades of troubleshooting instincts — and where the AI deliberately diverges (like parallelizing hypothesis exploration). He'll also discuss how we build trust through evaluation — comparing accuracy, latency, and reasoning against ground truth from SREs.
This is a practitioner's view of what it takes to build an AI investigator that engineers can rely on when things are on fire at 3am.... Read more
This talk focuses on applying FinOps and sustainability (GreenOps) principles in Kubernetes, addressing very common challenges such as cluster overprovisioning, inefficient resource usage, lack of cost visibility, and the environmental impact of always-on infrastructure. We share real-world practices, common mistakes, and practical lessons on how to run Kubernetes in a more efficient, cost-aware, and sustainable way.
We are two engineers who started working in Platform Engineering about a year ago, and we believe this perspective brings real value: we speak from recent, hands-on experience, from what we’ve encountered operating real clusters, what didn’t work, what did, and what we’ve learned along the way while trying to make Kubernetes more responsible in day-to-day operations.... Read more
This talk covers how we hire and run SRE, DevOps, and TechOps teams by treating infrastructure like a supply chain. Looking at systems this way makes bottlenecks, risk, and capacity limits easier to see before they cause outages.
I’ll explain what we look for when hiring, how we expect engineers to operate day to day, and how this approach helps us define the right metrics early instead of after incidents.... Read more
A room of engineers breaks my live app and we fix it together. They scan a QR code which triggers a real failure in a Kubernetes staging environment. Together we reproduce, trace, patch, and validate the fix with kubectl debug and mirrord. A real breakage repaired live.... Read more
Watching your technology choice go down the drain - because of broken promises, vendor implosion, or obsolescence - can feel like a career-ending experience. But in my experience there's usually a lot of good that comes out of those seemingly bad tech choices.... Read more
SREs often work under constant pressure, dealing with incidents and operational overhead. This talk explores how ephemeral environments can help make day-to-day SRE work easier, starting small and without disruptive changes.... Read more
18:00
Happy Hour by Imply - grab a beer!
Main lobby
18:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
In this talk, Redis and Wild Moose share how they're approaching AI SRE as something that can be designed, automated, and continuously improved. We'll explore how to build AI SRE systems that teams can actually trust in production by applying the same standards used for observability systems: fast, testable, configurable, and transparent.
Using a real-world example from Redis, we'll show how debugging agents turn tribal knowledge into structured investigation workflows and automate critical parts of the operational flow—creating a reliability system that improves with every incident. We'll also share why treating agents like production code, with regression testing and validation, is essential for long-term scale. You'll walk away with an understanding of the cultural shift needed to fully leverage AI for incident management.
Bio
Yasmin Dunsky is the Co-Founder & CEO of Wild Moose (YC W23), an AI-first SRE platform that helps engineering teams investigate production incidents in under a minute by turning tribal debugging knowledge into reliable, testable AI agents. Prior to founding Wild Moose, she led and scaled one of Israel’s largest nonprofits for computer science education, partnering with Google and Microsoft to build programs that reached thousands of students nationwide. She holds an MBA from Stanford and previously worked as a software engineer at Google.
Avner Yaacov is Director of Cloud Operations at Redis, where he leads global SRE, Cloud Operations, and Software Engineering teams responsible for the reliability, scalability, and operability of Redis Cloud. He manages the largest multi-cloud deployment of Redis, overseeing complex production environments across cloud providers and regions. With over a decade of experience in DevOps and SRE roles, Avner has built and operated large-scale distributed systems across cloud and high-performance environments. Over the past year, he has led Redis' shift toward AI-driven operations, focusing on moving from reactive operations to AI-assisted troubleshooting and automated remediation, while establishing clear standards for how automation is written, tested, rolled out, and trusted in production.
Mandi Walls
PagerDuty
Reevaluating Post-Incident Reviews
Abstract
Post-Incident Review. Postmortem. Incident Report. Whatever you call them, they take time and resources, and sometimes you’re not even sure if anyone reads them.
The Post-Incident Review is a story. Maybe it’s a bit of a mystery, maybe it’s a feel-good story of redemption. Maybe it’s a buddy comedy. We produce these reports so that other folks in our organization can learn from the things we learned and hopefully not repeat our same mistakes.
Collecting the data and creating the story are work. We’ll talk about how PagerDuty has changed and adapted our methodology over time to produce better reviews that help more engineers get more out of the process and the assets that are produced.
Bio
Mandi Walls is a DevOps Advocate at PagerDuty and an Advisory Board Member for the Customer Experience Executive Program at Ithaca College. With more than a decade of experience across operations, automation, and community leadership - including nearly nine years in technical and community roles at Chef Software - she brings deep expertise in incident response, service reliability, and practitioner education. A frequent speaker, Mandi focuses on practical DevOps practices and building healthier engineering cultures.
Jason Schechner
Traversal
What Makes an AI SRE Trustworthy? An SRE’s Take
Abstract
On-call is stressful enough without worrying whether AI hallucinated the root cause.
In this talk, an SRE who's been handling production incidents since before the role had a name will share what it's like to build an AI system that triages incidents and monitors system health — and why it's not as simple as throwing telemetry at an LLM.
He'll walk through how he worked with AI engineers to transfer decades of troubleshooting instincts — and where the AI deliberately diverges (like parallelizing hypothesis exploration). He'll also discuss how we build trust through evaluation — comparing accuracy, latency, and reasoning against ground truth from SREs.
This is a practitioner's view of what it takes to build an AI investigator that engineers can rely on when things are on fire at 3am.
Bio
Jason Schechner is a seasoned Site Reliability Engineer and technical leader with deep expertise across infrastructure, including compute, storage, networking, security, and automation. He currently serves as a Member of Technical Staff at Traversal, where he works on AI-driven systems for diagnosing and resolving real-world SRE incidents. Previously, he held senior SRE and production engineering roles at firms such as Teladoc Health, Citadel, The D. E. Shaw Group, and Goldman Sachs, leading large-scale, high-availability systems and infrastructure initiatives. Known for blending hands-on engineering with team leadership, he has a strong track record of building resilient platforms and guiding complex technical projects.
Maria Garcia Garcia & Lucia Lopez Barrero
Resizes
Responsible Kubernetes: clusters that don’t kill the planet (or your budget)
Abstract
This talk focuses on applying FinOps and sustainability (GreenOps) principles in Kubernetes, addressing very common challenges such as cluster overprovisioning, inefficient resource usage, lack of cost visibility, and the environmental impact of always-on infrastructure. We share real-world practices, common mistakes, and practical lessons on how to run Kubernetes in a more efficient, cost-aware, and sustainable way.
We are two engineers who started working in Platform Engineering about a year ago, and we believe this perspective brings real value: we speak from recent, hands-on experience, from what we’ve encountered operating real clusters, what didn’t work, what did, and what we’ve learned along the way while trying to make Kubernetes more responsible in day-to-day operations.
Bio
María García García is a junior platform engineer currently completing an apprenticeship in Gijón, Asturias. She is building her foundations in platform engineering while gaining hands-on experience in real environments. Previously, she worked as an application developer intern at ABAMobile, where she focused on multiplatform development, and as a web developer intern at Bittia, contributing to real-world projects and strengthening her teamwork skills. María has a background in technical training from TuniverS Formación and brings a practical, detail-oriented mindset shaped by both technical and non-technical roles.
Lucía López Barrero is a computer engineering student specializing in information technologies at the University of Oviedo. She is currently completing an internship, gaining practical experience alongside her academic studies. Lucía is in the final stage of her degree and is focused on developing a solid technical foundation while applying her knowledge in a professional setting.
Bryan Patton
Xurrent
The "Supply Chain" Mindset: Why I hire Logistical Engineers, Not Just SREs
Abstract
This talk covers how we hire and run SRE, DevOps, and TechOps teams by treating infrastructure like a supply chain. Looking at systems this way makes bottlenecks, risk, and capacity limits easier to see before they cause outages.
I’ll explain what we look for when hiring, how we expect engineers to operate day to day, and how this approach helps us define the right metrics early instead of after incidents.
Bio
Bryan Patton is Director of Site Reliability and TechOps at Xurrent, where he leads global reliability engineering, operational excellence including architecture and performance monitoring, and incident response practices for modern cloud environments.
Bryan previously directed a high-performance ecosystem engineered to handle massive surges in real-time sports data, consistently managing traffic loads exceeding millions of requests per second. He oversaw the scaling of cloud-native systems and low-latency databases to ensure 99.9% uptime during peak global sporting events, where millisecond-level precision is critical for user engagement. By implementing robust automated scaling protocols and optimizing distributed cache layers, they successfully balanced extreme throughput with system resilience and cost-efficiency in a high-stakes, data-intensive environment.
Adna Zujo Lakisic
MetalBear
Honey, the Audience Broke My App: Reproduce & Fix Live in Kubernetes with mirrord
Abstract
A room of engineers breaks my live app and we fix it together. They scan a QR code which triggers a real failure in a Kubernetes staging environment. Together we reproduce, trace, patch, and validate the fix with kubectl debug and mirrord. A real breakage repaired live.
Bio
Adna Zujo Lakisic is a Solutions Engineer based in North America, currently working at MetalBear. She brings a strong engineering background spanning software development, technical consulting, and solutions engineering across media, marketing, and technology companies. Prior to MetalBear, she spent several years as a Senior Engineer at Haymarket Media US, where she supported large scale digital platforms. Her experience also includes technical consulting roles at Omega IT and solutions engineering at Midan Marketing, with hands on work in areas such as Terraform, Python, Kubernetes, and cloud infrastructure. Earlier in her career, Adna helped shape and fund a mobile tour guide platform focused on promoting tourism in Mostar through curated self guided experiences.
Leon Adato
Cribl
Looking back at a lifetime of poor tech choices
Abstract
Watching your technology choice go down the drain - because of broken promises, vendor implosion, or obsolescence - can feel like a career-ending experience. But in my experience there's usually a lot of good that comes out of those seemingly bad tech choices.
In my sordid career, I have been an actor, bug exterminator and wild-animal remover (nothing crazy like pumas or wildebeests. Just skunks, snakes, and raccoons.), electrician, carpenter, stage-combat instructor, ASL interpreter, and Sunday school teacher. Oh, yeah, I've also worked with computers.
While my first keyboard was an IBM selectric, and my first digital experience was on an Atari 400, my professional work in tech started in 1989 (when you got Windows 286 for free on twelve 5¼” when you bought Excel 1.0). Since then I've worked as a classroom instructor, courseware designer, helpdesk operator, desktop support staff, sysadmin, network engineer, and software distribution technician.
Then, about 25 years ago, I got involved with monitoring. I've worked with a wide range of tools: Tivoli, BMC, OpenView, janky perl scripts, Nagios, SolarWinds, DOS batch files, Zabbix, Grafana, New Relic, and other assorted nightmare fuel. I've designed solutions for companies that were modest (~10 systems), significant (5,000 systems), and ludicrous (250,000 systems). In that time, I've learned a lot about monitoring and observability in all it's many and splendid forms.
Marcos Novelli Harispe
LoopStudio
Making Life Easier for SREs with Ephemeral Environments
Abstract
SREs often work under constant pressure, dealing with incidents and operational overhead. This talk explores how ephemeral environments can help make day-to-day SRE work easier, starting small and without disruptive changes.
Bio
Marcos Novelli Harispe is a software engineer specializing in AWS, serverless architectures, and data-driven solutions. He has delivered scalable notification systems, analytics platforms, and automation tools across freelance and full-time roles, with a focus on reliability, efficiency, and measurable business impact.
Dewan Ahmed
Harness
KeynoteSecure by Default: Building Confidence in AI-Driven Delivery
Abstract
The fastest way to break trust in DevSecOps is to automate insecurity at scale. As AI takes a central role in our pipelines, it is time to rethink what "secure by default" really means.
In this keynote, Dewan Ahmed will challenge the audience to look beyond vulnerability scanners and compliance gates. He will share a vision for intelligent security by design, where native intelligence within the delivery platform detects not only vulnerable code but risky delivery behavior such as misconfigured environments, suspicious artifact provenance, and drift between source and runtime.
You will walk away with a framework for balancing automation with human oversight and examples from Harness’ work on building verifiable, auditable, AI-native delivery systems. In the new world of DevSecOps, safety is not a step; it is an outcome we continuously learn to improve.
Bio
Dewan Ahmed is a Principal Developer Advocate at Harness and a Governing Board General Member Representative at the Continuous Delivery Foundation. He focuses on DevRel and content strategy across CI/CD, DevOps, and open source, with deep expertise in software supply chain security and developer experience.
Birol Yildiz
ilert
KeynoteWhen Incidents Fix Themselves: AI SRE in action
Abstract
The next evolution of incident response isn’t faster alerts, it’s autonomous resolution. Join ilert CEO Birol Yildiz as he shows how AI SRE agents now diagnose and remediate outages without waking anyone up. Learn how these systems combine observability data, deployment context, and code intelligence to restore services in minutes and hand over clean incident reports instead of 3 a.m. pages.
Bio
Birol Yildiz is the Co-founder and CEO of ilert, adeptly steering the company with a rare combination of technical and product expertise. His prior experience includes a significant role as Chief Product Owner for Big Data products at REWE Digital. With a strong foundation in computer science, Birol bridges the gap between developer and product strategist, constantly striving to innovate and provide customer-centric solutions at ilert.
Ben Hopp
Imply
KeynoteDecoupling Observability for Incident Response at Scale
Abstract
On-call incidents don’t fail because teams lack dashboards. They fail because observability systems slow down under real investigative load.
As telemetry volumes grow and retention windows expand, SRE teams are being asked to run deeper, broader investigations—often under time pressure—on platforms that were designed for steady-state monitoring, not bursty incident response. Tightly coupled observability stacks bind storage, compute, and query together, forcing teams to overprovision infrastructure, limit retention, or accept degraded performance during incidents.
In this talk, we’ll explore why decoupling observability architectures is becoming essential for SRE teams operating at scale. Using a real incident investigation workflow, we’ll break down how separating data storage, compute, and interaction layers allows teams to keep fast, reliable monitoring while elastically scaling investigations when incidents occur.
We’ll connect these patterns to lessons learned in other data-intensive systems, but stay grounded in the day-to-day realities of on-call life: faster root cause analysis, fewer tradeoffs during incidents, and observability systems that hold up when you need them most.
Bio
Ben Hopp is an Architect at Imply, where he helps organizations decouple their observability stack with Imply Lumi. Based in Denver, Ben brings an extensive background in real-time data systems and observability, having worked on some of the world's largest data platforms. He specializes in helping teams break free from legacy observability stacks, enabling them to retain flexibility and control over their observability data while reducing costs and complexity.
AJ Mejorado
Chainguard
KeynoteThe Best CVE Is the One You Never Patch: Removing Vulnerability Toil for SREs
Abstract
For most SRE teams, CVE remediation has become a never-ending source of toil—rebuilds, backports, emergency releases, and late-night firefighting. This talk explores how Chainguard flips that model by removing vulnerability remediation from the SRE workload entirely.
We’ll look at how secure-by-default images, continuous rebuilds, and automated patch pipelines eliminate the need for teams to chase CVEs in production. Instead of reacting to vulnerabilities, SREs inherit artifacts that are already patched, minimal, and ready to run—freeing them to focus on reliability, performance, and scale.
This session is about shifting CVE response left, shrinking the attack surface, and giving SRE teams their time back—because patching shouldn’t be an on-call responsibility.
Bio
AJ Mejorado is a Solutions Engineer at Chainguard with a background in DevOps and CI/CD. He focuses on helping teams improve software delivery practices, scale pipelines safely, and reduce risk across the development lifecycle. AJ has held engineering, sales engineering, and product education roles, including helping build certification programs and developer enablement initiatives.
You use an LLM, ask it some questions, have it write a little code for you, and before you know it, you're hit with a massive bill. Out of the box, there's zero way to see token usage, cost, and an overall ability to control it. In this talk, you'll learn exactly how to do that.
Bio
Michael Levan translates technical complexity into practical value. He is a seasoned engineer, advisor/solutions engineer, and content creator in the AI and Platform Engineering space who spends his time working with organizations around the globe on technical implementation and strategy. Michael is also a Microsoft MVP, 4x published author, podcast host, international public speaker, CNCF Ambassador, and was part of the Kubernetes v1.28 Release Team.
Endre Sara
Causely
KeynoteThe Missing Layer in AI-Driven Reliability
Abstract
AI agents are entering incident response with powerful capabilities—detecting anomalies, summarizing logs, and suggesting fixes. Some even execute changes autonomously. The promise is compelling: fully autonomous resolution of production incidents in complex, constantly changing distributed systems.
But before we hand control to agents, we must ask a fundamental question: What kind of understanding must exist before we let agents act?
This talk argues that the real bottleneck is not model intelligence, but the absence of a continuously updated causal model of the system itself.
Bio
Co-Founder of Causely and former Engineering Director at Turbonomic (acquired by IBM), Endre has over 20 years of experience designing distributed data and runtime platforms. Before joining Turbonomic, he served as Vice President of Enterprise Systems Management at Goldman Sachs, leading large-scale infrastructure and monitoring initiatives. Endre holds a Ph.D. in Electrical Engineering from Stevens Institute of Technology and an M.S. in Electrical Engineering from the Budapest University of Technology and Economics.
Francois Martel
NeuBird AI
KeynoteThe AI SRE Landscape: From LLMs to Multi-Agent Ecosystems
Abstract
With more than 15 vendors now claiming AI powered SRE capabilities, engineering leaders are facing a deafening amount of noise. Teams are asking hard questions. Is Datadog Bits AI the same thing as an AI agent? AWS just launched a DevOps Agent. Do we still need a vendor? Our ServiceNow rep says their AI can handle it all. How do you cut through the hype?
This talk provides a practical framework for understanding the AI systems transforming Site Reliability Engineering, including foundation LLMs, RAG systems, chat based tools, agentic AI, and multi agent ecosystems. We map the current vendor landscape across five categories: Incident Coordinators, AI Investigators, Observability Native AI, Cloud Provider Agents, and Coding Agents. You will see what each category can actually do today and why context engineering, not model intelligence, is the real differentiator.
The session includes a live walkthrough of a multi agent workflow handling a production incident end to end. It starts with intelligent detection and root cause analysis, continues through automated remediation using a coding agent, and finishes with verified resolution. Humans stay in the loop only where it matters. The workflow persists for hours, coordinates across tools, and retains full context throughout.
You will leave with a clear mental model for evaluating AI SRE tools and a phased adoption roadmap. The key takeaways are straightforward. Know what type of AI you are buying. Invest in context architecture over model hype. Build a layered stack rather than a monolith.
Bio
Francois Martel is the Field CTO at Neubird.ai, where he works with SRE and platform teams to apply generative AI to real world reliability and operations challenges. He focuses on helping engineering leaders move beyond hype to practical, production ready AI systems for incident response, observability, and operations at scale.