From Reactive to Intelligent: Site Reliability Engineering is Moving Towards Autonomy. As organisations deploy increasingly complex and distributed architectures, traditional methods of incident response based on reactive practices fall behind. The need arises for smarter automation which would allow for quicker detection of anomalies, faster context analysis, and timely interventions within defined constraints. In this sense, agentic SRE is emerging as a discipline in which AI agents learn to analyze telemetry data, establish relationships among various system components, and help SREs with triage, diagnostics, and troubleshooting while respecting certain boundaries. In this talk, I will discuss the principles of building a robust and dependable platform for agentic SRE, the key components that constitute the architecture, and the observability layer which provides foundational data. I will also examine the necessary safeguards required to ensure autonomy while at the same time keeping operations secure and auditable. Finally, through examples of practical use cases, such as alert triaging, incident copiloting, root cause analysis, and bounded automated remediation, we will explore how AI can be integrated into operations to improve reliability.
I help platform engineers and DevOps teams understand and adopt cloud-native infrastructure through talks, demos, and community building. I organise the CNCF Gandhinagar, HashiCorp UG Gandhinagar, and Google Cloud Gandhinagar communities, and I have spoken at 40+ events, including KubeCon India, PlatformCon, HashiTalks, OpenSearchCon Korea, FOSS ASIA, and LinuxFest Northwest. My core focus areas are Kubernetes, GitOps, IaC with Terraform/Vault, AI-powered DevOps tooling, and observability pipelines. I've mentored 15+ hackathons and open-source programmes. Certifications: GCP DevOps Professional, Azure AZ-104, FinOps Certified Engineer, Terraform Associate, Vault Associate.