
As generative AI moves from the cloud to the edge, we are colliding with a hard reality: physics. Deploying 500GB-plus models into disconnected, resource-constrained environments—such as air-gapped critical infrastructure—exposes fundamental limits in today’s containerization and delivery patterns. Docker layers time out, container registries collapse under load, and application startup times stretch beyond 45 minutes, rendering traditional approaches impractical. In this talk, I’ll walk through a battle-tested architecture designed to break this “physics versus physics” deadlock. The approach reframes how we package and deliver large language models at the edge by decoupling lightweight inference engines from massive model weights using OCI Artifacts, enabling far more flexible deployment strategies. I’ll also introduce a sideloading ingress pattern that bypasses container registries entirely, injecting model data directly from object storage to eliminate registry bottlenecks during large, concurrent transfers. Finally, I’ll share a practical quantization decision matrix that makes it possible to run 70B-parameter models within the constraints of a single A100 GPU. This is not a theoretical exploration. The architecture was validated in production during Exercise Mobility Guardian 2025, where it powered GenAI workloads on a GDC appliance in fully disconnected, zero-internet environments. Attendees will leave with concrete patterns and lessons learned for deploying large-scale generative models where bandwidth, connectivity, and startup time are non-negotiable constraints.
I specialize in AI infrastructure and edge reliability at Google based on SF Bay Area, currently focusing on solving the "Sovereignty Paradox" for secure environments. My work sits at the intersection of SRE, platform engineering, and critical infrastructure resilience.