Modern observability is very good at telling us what is unhealthy. It is much worse at telling us why.
A database can be on fire because it caused an incident—or because something several hops away made it the place where the damage accumulated. An LLM can read all of the telemetry and still make the same mistake. The problem is not always a lack of data or intelligence. Sometimes what is missing is an understanding of how this particular system behaves when things go wrong.
While building TruffleRoot, we’ve been exploring a different approach: give machines an explicit, evolving model of an environment—how it is connected, what changes over time, how failures can propagate, and what evidence supports or contradicts an explanation.
This talk is about that shift from observability to causality: why generic assumptions about failure break down in real systems, what it means for a system to learn how your system fails, and where deterministic reasoning and LLMs each fit into that picture.
The goal isn’t an AI that tells a better story about an outage. It’s a system that learns how your system fails and can show its work.
Smita Pasumarthi builds enterprise AI applications at truffleroot.ai, drawing on extensive experience in enterprise integration and automation. Her expertise spans generative AI, LLMs, AI infrastructure, MuleSoft, Boomi, Workato, and Python.
Chandan Chilumula is a technology professional based in the San Francisco Bay Area, with experience in software solutions and enterprise technology.