Categories
Articles Videos

AI / LLM Context Caching: Mechanics, Economics, Storage and Data-Retention Risk

How prompt caching works, why cached input is billed at a fraction of the input price, where the cache physically lives, and what it means for data retention. 2 October 2026.

Executive summary

  1. What is cached is not text — it is internal model state. The provider stores the attention key/value (KV) tensors computed during prefill for a prompt prefix and reuses them when a later request begins with that identical prefix. The raw prompt does not need to live in any shared text store for caching to work.
  2. Cached input is cheap because a cache hit converts a compute-bound job into a memory-read job. Recomputing the KV state for a long prefix is the dominant cost of serving that prompt; reusing it costs a fraction of that. Cached reads have settled around 10% of the full input price across major providers, with 1.25–2× write premiums where cache creation is charged, and best-effort discounts where it is free.
  3. The cache lives on the provider’s serving infrastructure, short-lived and tiered: GPU/HBM first, spilling to GPU-local storage for extended retention (up to 24 h on OpenAI and Azure), host RAM or disk tiers in distributed setups, and — uniquely — durable distributed disk arrays on DeepSeek, cleared within hours to days. It is scoped to the customer (organisation/workspace/project/account) and is never directly customer-accessible.
  4. Data-retention risk is real but bounded, and heavily provider-specific: cached state is derived customer data retained for minutes up to 24 h (or days on DeepSeek), expiry is not immediate erasure, isolation is organisation-level rather than per-user-level, and cache-hit timing side channels are demonstrated in academic research. Customers with strict zero-data-retention requirements face specific model-level conflicts that must be checked explicitly (Section 5).