
Your API monitoring was green. Dashboards calm. Then a quiet spike: cost per task up 40%, grounded answer rate down 8%, and users start regenerating responses twice as often. Infra metrics say “all good” , but the model silently shifted behavior after a prompt tweak plus a vendor embedding update. Non-AI-adopted SRE doesn’t page you here. AISRE would.
As AI-powered systems move into production, many teams discover that traditional Site Reliability Engineering metrics,latency, availability, and error rates are no longer sufficient to describe real system health.
AI isn’t just predictable APIs anymore. We’re shipping probabilistic systems: prompts → retrieval → model decoding → agents → filters → feedback loops. Every layer can drift independently… and still return a 200 OK.
In this talk, I introduce AI Site Reliability Engineering (AISRE): an extension of SRE principles tailored specifically for AI-driven systems. I explore how reliability must expand to include semantic correctness, grounding quality, safe tool execution, economic efficiency, and controlled behavioral drift.
Ehsan Khodadadi is a Senior Site Reliability Engineer at ING, with extensive experience leading and building reliability practices across large-scale systems. Before rejoining ING, he led the Site Reliability Engineering team at LeasePlan, where he focused on system stability, team management, and operational excellence. His background includes roles at Techspire and BMW Group, where he combined deep technical expertise in DevOps and Linux systems with a pragmatic, hands-on approach to problem solving. Ehsan is known for creating strong engineering teams and improving service reliability through thoughtful automation and collaboration