
Most reliability work today is still centred on reactive troubleshooting: diagnosing a multitude of alerts, pulling large groups into incidents, and engineers scrambling to understand what's happening. To truly change that pattern, we need systems that can predict and prevent failures before they occur. Biological immune systems offer a powerful blueprint for how software can defend and ultimately heal itself.
This talk introduces a framework for thinking about "software immunity", highlights the gaps in today's observability and includes some "under the hood" details on designing AI agents for reliability.
Innate immunity comes from built-in defences like testing, feature flags, auto-scaling, and circuit breakers — mechanisms that provide immediate, general protection. Adaptive immunity, meanwhile, emerges from AI agents that learn from new data, refine their understanding of system behaviour, and apply those lessons to predict and pre-emptively fix failures.
We'll break down the key ingredients for trustworthy AI agents in reliability and beyond: transparent reasoning rather than opaque black-box outputs; strong control mechanisms and guardrails; a governed data layer for effective data access; continuous learning from each execution cycle; and graduated autonomy — from suggestions, to human-in-the-loop actions, to fully automated remediation.
Matt Henderson is the co-founder and CEO of Phoebe, an AI agent platform for software reliability. Previously, he was the CEO of Stripe Europe, and led the company’s international operations across product and engineering. Earlier in his career, Matt was a product director at Amazon and Google, and co-founded the ML analytics startup Rangespan (acquired by Google). He is also an angel investor in over 100 startups.