You've seen the demo. The agent read the alert, walked the traces, found the misconfigured deployment, and proposed a fix — all in ninety seconds, on stage. Now you have to decide whether it goes anywhere near your production cluster. The published evidence should make you cautious. On IBM's Kubernetes incident benchmark, the best model today scores 56% under a metric that awards zero for a single missed entity. On CUJBench, agents given the full toolset performed worse than agents restricted to browser evidence — 20% versus 28% — and one frontier model collapsed from 52% to 12% accuracy when handed more tools, burning 92 tool calls and 4 million tokens per failed run without ever submitting an answer. Across 1,675 root-cause analysis runs, hallucinated interpretation of data showed up in 71% of them, at every model tier. None of that means agents are useless. It means the demo told you almost nothing, and the questions you'd normally ask a vendor don't cover the failure modes that actually matter here. This talk is six questions, each grounded in what ITBench, AIOpsLab, SREGym and CUJBench actually measured — and each with a way to answer it on your own cluster rather than taking a number on faith. We'll cover why more context can make diagnosis worse, why an agent that retrieves the right evidence still names the wrong service, why a scenario catalog is not a benchmark, and what happens when an agent "resolves" an incident by deleting the fault injector that caused it. You'll leave with the six questions, and a harness design for answering them yourself.
Diego Braga is CTO at Krateo PlatformOps. Diego has spent the last 10 years building open-source architectures for customers, from embedded devices to large-scale distributed systems. Most recently he has been focused on the open cloud infrastructure space, and on emerging patterns for cloud-native applications.