
AI agents are showing up in on-call workflows. They triage alerts, pull runbooks, correlate metrics and forget everything the moment an incident closes. That's a problem. SRE agents need more than fast inference. They need context that persists across incidents, spans multiple tools, and doesn't bloat every prompt with the entire history of your system. That's memory management, and most teams aren't thinking about it yet. In this talk, I'll walk through how memory works (and breaks) in agentic SRE workflows, the difference between short-term context, episodic recall across incidents, and shared memory across a team of agent. You'll leave with a mental model for designing memory into your SRE agents from day one, not bolting it on after your agent confidently recommends the same wrong fix three incidents in a row.
Dipesh has been building software for over 14 years across Enterprises and startups. He was head of engineering at a $500M company for 5 years before starting his own venture in monitoring space.