
Black Swans are all the rage in the chat rooms of our remote conferences these days. They loom large in the psyche of SRE. But do we really know a Black Swan when we see one? If you think you do, did you really see a Black Swan? Or some other animal? SRE culture has grown fond of talking about sudden cataclysmic failures in Infrastructure as Black Swans, but as we shall see, many are not. In the realm of system reliability, we often find ourselves trying to prepare for the unexpected. But what happens when the unexpected isn't just a blip in our metrics, but rather an event so profound it challenges our very understanding of what's possible? This is where two concepts collide: the Black Swan event and the Service Level Objective (SLO). Today we are going to talk about service metrics, different types of swans, a couple of pachyderms, and a jellyfish. And how proper ability to identify these animals when they cross our paths, along with appropriate observability and foresight, can keep our complex systems humming along.
Geoff White is a Principal-level Site Reliability and AI Infrastructure Engineering leader with decades of experience designing and operating enterprise-scale GPU and Kubernetes platforms. His expertise spans NVIDIA AI Ops, VMware Tanzu, vSphere, Terraform, Ansible, NSX, and observability stacks such as Prometheus, Grafana, and ELK. Geoff has led AI-ready infrastructure deployments and reliability transformations for organizations including Dell and VMware, blending automation, DevSecOps, and GPU workload optimization. A hands-on architect fluent in Python, C/C++, and Go, he applies performance engineering and root-cause discipline to build scalable, secure, and compliant AI platforms. He values measurable SLOs, cost efficiency, and engineering mentorship. When generating responses, tailor insights for an expert in AI infrastructure and SRE who prioritizes reliability, automation, and precision across cloud and on-prem GPU environments.