A cheaper model does not always mean a cheaper run. Routing between models, trimming context, splitting work across agents, or adding tools and skills mid-run can look like smart optimizations. But when those changes disrupt prompt-cache reuse, they can turn discounted, reusable context into input you pay to process again—undermining the savings you expected. OpenAI Developers This talk explores why prompt caching deserves to be an architectural consideration, not an afterthought, in production LLM systems. We’ll examine how caching interacts with model routing, context management, subagents, and dynamic tool loading. Through concrete examples, we’ll compare preserving a warm cache with starting fresh, explore when breaking the cache is worth the trade-off, and show how to make these decisions based on the economics of an entire run rather than the price of an individual request. Attendees will leave with practical patterns for cache-aware workflows and an observability checklist covering cache reuse, cache-read and write costs, latency, and cost per successfully completed task. The takeaway: don’t optimize away your biggest savings. Optimize the whole run—not just the token count.
Dor Amir is the founder of Nadir, an LLM infrastructure company focused on making AI systems more efficient. Previously, he worked on machine learning and personalization at Dropbox, Amazon, Guesty, and Fiverr. His work focuses on production AI systems, LLM routing, agents, and reducing the cost and latency of running AI at scale.