SREday

Site Reliability, DevOps and Cloud

Nov 8, 2024 San Francisco, CA, USA

1
Day
16+
Speakers
1
Track
100
Attendees

Reliable AI Needs More Than Observability

Joel Lipton
FoxCommand

We know how to observe conventional software: requests, traces, errors, latency, resource utilization. But an AI system can return 200 OK, stay within its latency budget, and still make the wrong decision. That creates a different reliability problem. What instructions and policies governed an execution? What did the model actually receive? Which decisions were probabilistic and which were deterministic? What changed between two runs? Can we replay an execution, test a proposed change against historical cases, and produce evidence of exactly what happened? This talk explores the architecture that emerged while building AI for real operational workflows: shrinking the responsibility of the model, moving policy and authority into deterministic software, and treating execution itself as a governed, replayable artifact. The goal isn’t to make probabilistic models deterministic. It’s to make the systems around them controllable, testable, and accountable.

Joel Lipton is the founder of FoxCommand and a former Product Designer at Amazon, where he worked on Agent AI integration into customer facing and operations products. His background spans product design, AI systems, and complex operational workflows. He is currently building FoxCommand, an execution and governance layer for AI systems that lets teams define deterministic contracts and policies around probabilistic models, then replay, compare, simulate, and audit how those systems behave. His work focuses on a practical engineering problem: how do we get the flexibility of probabilistic models without giving up the control, testability, and accountability we expect from production software?

Sponsors & Partners

Want to become a sponsor? Get in touch!