
Most AI agents in production process each incident from scratch. No memory of what worked last time, no feedback loops, no adaptation. Stateless tools pretending to be smart. At Cleric, we build an autonomous AI SRE. Getting the agent to diagnose incidents was the easy part. Getting it to retain and apply what it learned from previous ones is where the real engineering problems lie. When one engineer figures out that an OOM spike is always the Redis sidecar, the agent should know that too. And it should know it across teams, across services, across time. We built a three-layer operational memory architecture — semantic, episodic, and procedural — that enables the agent to retain context across investigations and to improve over time. Semantic memory captures what the agent knows about infrastructure and relationships. Episodic memory records specific investigations and their outcomes. Procedural memory encodes the patterns that worked and when to apply them. But captured knowledge decays. The runbook from six months ago references a service that has since been decomposed into three microservices. The fix that worked in Q3 causes a different failure in Q1 because traffic patterns shifted. This talk covers how we detect and handle staleness, the feedback loops that update agent behavior based on resolution outcomes, and what we've learned about building agents that actually get better at their job over time.
Shahram Anver is the Co-Founder and CEO of Cleric, where he's building an autonomous AI SRE that investigates and resolves production incidents 24/7. Before Cleric, Shahram led engineering for MLOps, container deployment, and FinOps platforms at Gojek, Southeast Asia's largest super-app, managing infrastructure handling millions of daily transactions across hundreds of microservices. He previously built TripAdvisor's first ML automated bidding system, scaling it to manage tens of millions in annual ad spend, and co-founded DataCue, an ML-driven e-commerce personalization platform.