Categories
Articles Videos

AI / LLM Context Caching: Mechanics, Economics, Storage and Data-Retention Risk

How prompt caching works, why cached input is billed at a fraction of the input price, where the cache physically lives, and what it means for data retention. 2 October 2026.

Executive summary

  1. What is cached is not text — it is internal model state. The provider stores the attention key/value (KV) tensors computed during prefill for a prompt prefix and reuses them when a later request begins with that identical prefix. The raw prompt does not need to live in any shared text store for caching to work.
  2. Cached input is cheap because a cache hit converts a compute-bound job into a memory-read job. Recomputing the KV state for a long prefix is the dominant cost of serving that prompt; reusing it costs a fraction of that. Cached reads have settled around 10% of the full input price across major providers, with 1.25–2× write premiums where cache creation is charged, and best-effort discounts where it is free.
  3. The cache lives on the provider’s serving infrastructure, short-lived and tiered: GPU/HBM first, spilling to GPU-local storage for extended retention (up to 24 h on OpenAI and Azure), host RAM or disk tiers in distributed setups, and — uniquely — durable distributed disk arrays on DeepSeek, cleared within hours to days. It is scoped to the customer (organisation/workspace/project/account) and is never directly customer-accessible.
  4. Data-retention risk is real but bounded, and heavily provider-specific: cached state is derived customer data retained for minutes up to 24 h (or days on DeepSeek), expiry is not immediate erasure, isolation is organisation-level rather than per-user-level, and cache-hit timing side channels are demonstrated in academic research. Customers with strict zero-data-retention requirements face specific model-level conflicts that must be checked explicitly (Section 5).

1. Scope, method and evidence grading

This brief answers four questions: how context/prompt caching works in LLM inference, why cached input is materially cheaper, where the cache is stored, and whether provider-side caching carries data-retention risk. It was produced on 2 October 2026 from live web research across provider technical documentation, product pages, pricing trackers and security literature, with each material claim cross-checked across independent sources where possible.

Evidence grading, used throughout: – (provider docs) — the vendor’s own technical or product documentation. Highest band used in this brief. – (third-party) — pricing aggregators, third-party trackers and product blogs. Directional; verify against the provider pricing page before using in a business case. – (academic) — peer-reviewed or preprint security research. – (reported) — figures from secondary sources where the primary document was not directly reachable.

Rates, TTLs and defaults change frequently (within the last week OpenAI changed the default cache-retention mode for organisations without ZDR). Per-model rates must be taken from the provider pricing pages at decision time (provider docs for structure; third-party for current numbers).

2. How it works (mechanics)

Transformer inference has two phases, and the cost difference between them is what the entire discount is built on:

  • Prefill — the model processes all input tokens at once and computes, at every layer, the key/value (KV) vectors each token needs for attention. This is dense, compute-heavy work that scales roughly linearly with prompt length, and it is the dominant cost of serving long prompts.
  • Decode — the model generates output tokens incrementally, repeatedly attending over the stored KV state. This is memory-bandwidth-bound.

Without caching, a second request whose prompt shares a long prefix with the first recomputes the entire KV state from zero — even if 90% of the tokens are identical.

What prompt caching does: after a request, the serving system retains the KV tensors for the cached token span of the prompt prefix. The next request’s prompt is hashed or fingerprinted and matched against stored prefixes; on a hit, the provider loads the stored KV state and computes prefill only for the new suffix. Decoding proceeds as normal — caching never skips generation.

  • Exact-prefix, start-anchored. The match is the longest identical prefix starting at token 0. OpenAI requires at least 1,024 tokens and grows the match in 128-token increments from there (provider docs, Prompt Caching in the API). DeepSeek states that only prefixes identical from the 0th token count, with no matching in the middle of the input (provider docs, news announcement). A change anywhere near the start invalidates everything after it.
  • Where breakpoints are set differs by provider. OpenAI sets breakpoints implicitly on older models, with explicit caller-set breakpoints on the newest family. Anthropic exposes explicit cache_control breakpoints (up to four) plus automatic caching (provider docs). Gemini offers implicit caching (automatic, best-effort, on Gemini 2.5+ models) and explicit caching (a managed CachedContent resource the caller creates, with guaranteed savings) (provider docs). AWS Bedrock supports implicit and explicit cache checkpoints per model (provider docs).
  • Hits are best-effort and routing-dependent. KV state sits on particular serving machines, so a request routed elsewhere misses even if the prefix was computed recently. OpenAI exposes a prompt_cache_key that increases routing stickiness — requests sharing a prefix are more likely to land on the engine holding the state — but a hit is never guaranteed (provider docs, Prompt Caching 201).
  • TTLs and refresh. Lifetime is measured from the last use; a hit refreshes the entry, usually without re-charging the write. Typical values: 5–10 minutes baseline on OpenAI/Azure model families and Bedrock; 5 minutes default with an optional 1 h TTL on Claude; 1 h default for explicit Gemini caches; a 30-minute minimum with an optional 24 h extended retention on the newest OpenAI models; hours to days on DeepSeek’s disk cache.

The practical corollary for prompt design: stable content first (system instructions, tool definitions, reference documents), variable content last, and nothing user-specific above the last cache breakpoint.

3. Why cached input is materially cheaper

  • Compute saved. A hit skips prefill for the cached span — the GPU work for that span is simply not spent. The remaining work (fetching KV state from memory, prefilling the suffix, decoding) is far cheaper than recomputation, so the provider’s marginal cost of a cached token is a small fraction of an uncached one.
  • Capacity is freed. Cached prefixes raise fleet throughput (more concurrent requests per GPU-hour), part of which providers pass through as the discount.
  • Latency follows cost. OpenAI cites latency reductions of up to roughly 80% for prompts over 10,000 tokens (provider docs, Prompt Caching 101).
  • But caching is not free for the provider. Retention consumes HBM, RAM and disk, and routing adds overhead. That is why some providers charge a write premium (paid when the cache is created), read at a discount, and why automatic caches come with no-savings-guarantee language (Gemini implicit; DeepSeek best-effort).

Pricing structure by provider (structure per provider docs; current per-model numbers per third-party trackers — verify before quoting):

ProviderMechanismCache writeCache read (cached input)Default TTL
OpenAIAutomatic (implicit breakpoints); explicit on newest family1.25× input on newest-family models; free on older models~0.1× input (example from provider pricing: $2.00 → $0.10 per 1M)30 min minimum on new models; 24 h optional extended retention; 5–10 min historical
AnthropicAutomatic plus explicit cache_control (up to 4 breakpoints)1.25× input (5 min TTL) / 2× input (1 h TTL)0.1× input5 min, optionally 1 h
Google GeminiImplicit (auto, best-effort) / explicit CachedContent resourceStandard input price on cache creation, plus storage billed per hour of retentionDiscounted — roughly 0.25× of input per Gemini/Vertex docs; guaranteed on explicit1 h explicit (configurable); 24 h or less implicit on Vertex
Azure OpenAIFollows OpenAI model behaviourNo separate write fee on standard deploymentsDiscounted; up to 100% off input on Provisioned (PTU) deployments5–10 min historical; up to 24 h extended retention (GPU-local)
AWS BedrockImplicit and explicit checkpoints (per model)Write premium, roughly 1.5× input per AWS docs — verify per model~0.1× input5 min default; 1 h on select models
DeepSeekOn-by-default disk-based prefix cachingNone~0.1× of uncached at launch ($0.014 vs $0.14 per 1M); smaller fractions on current models (third-party tracker)Cleared within a few hours to a few days

Break-even logic: if a prefix of P tokens is reused R times within its TTL, the caching customer pays approximately P × (1.25 + 0.1 × (R − 1)) input-token-equivalents instead of P × R. Caching pays from the second use on the 5-minute tiers; the same arithmetic says a 1 h Anthropic write (2×) pays if the prefix survives to its second use.

4. Where the cache is stored

The cached object is the KV tensor block plus a cache key/hash for lookup — not the prompt text — and it is never customer-visible or customer-addressable, except for Gemini’s explicit CachedContent resources, which the caller can enumerate and delete. Retrieval is automatic on an exact-prefix match. Stated storage media and scopes, per provider documentation:

ProviderStated storage mediumStated retentionScope and isolationPersisted at rest?
OpenAIGPU/HBM first; extended retention writes encrypted tensors to local GPU machines (documented as “application state”)5–10 min historical in memory; 30 min minimum on newest models; up to 24 h extendedOrganisation; not shared across organisations or regional-processing boundariesYes — GPU-local storage for extended retention
AnthropicKV representations and hashes held in memory5 min default; 1 h optionalOrganisation; workspace-isolated within an organisation on Claude API, Claude Platform on AWS, Microsoft FoundryNo — “not stored at rest”
GoogleVertex default performance cache: in memory, up to 24 h; explicit cache: managed resource per project and location24 h or less (Vertex default); 1 h explicit default, configurableProject + location (Vertex); project-level isolation for the default cacheNot at rest for the Vertex default cache; explicit resources persist for their TTL
Azure OpenAIIn memory; extended retention offloads KV tensors to GPU-local storage when memory fills5–10 min historical; up to 24 h extendedNot shared between Azure subscriptions; data-zone or region boundary for extended retentionYes — GPU-local storage for extended tier
AWS BedrockInternal model state on AWS infrastructure5 min default; 1 h on select modelsAWS account + region; may share across requests within that scopeDurable storage not stated in docs
DeepSeekDistributed disk array (Context Caching on Disk)A few hours to a few daysPer user — “logically invisible to others” (vendor claim)Yes — durable disk

5. Data-retention risk on the provider side

Yes — the cache is customer-derived data retained on the provider’s side, and a privacy or DPA review should count it as such for its TTL. The risk profile is nonetheless materially better than “the provider stores your prompts”. The six points that matter:

  1. It is derived state, not raw text — but treat it as data anyway. KV tensors are not readable prompt text, yet they are derived from it and can leak information about it (see point 5). OpenAI frames cached state as “application state” and may store encrypted tensors on GPU machines; Azure persists it on GPU-local storage; DeepSeek on disk arrays. Count the TTL-window retention in the privacy review rather than assuming “in-memory cache = nothing retained”.
  2. ZDR interactions are provider-specific and model-specific. Anthropic documents prompt caching as ZDR-eligible (in-memory only, deleted promptly after TTL). Google’s ZDR guidance says explicit CachedContent is incompatible with an “absolute zero-data footprint” — do not use explicit caching where a zero-footprint is required. OpenAI offers caching with ZDR and defaults ZDR organisations to in-memory retention — but its newest models do not support the in_memory option, so the extended (24 h) state can be unavoidable on them. Bedrock’s general “no durable storage” retention mode does not explicitly answer the cache question. DeepSeek is the material outlier: disk-backed caching retained for days, and its privacy policy states personal data is processed and stored in China.
  3. Deletion is expiry, not purge. Cleanup is TTL-based (“promptly, though not immediately” — Anthropic). Only Gemini’s explicit cache offers customer-triggered deletion; no provider guarantees immediate erasure or zeroisation of cache state. Contractually acceptable for short TTLs, but it should be a conscious decision, not an oversight.
  4. Isolation is organisation-level, not user-level. Within your own org, workspace, project or account, different users or applications can hit each other’s cached prefixes. That is usually a feature (shared system prompts), but it is a design hazard in multi-tenant products: a tenant-specific prefix cached in a shared workspace can be reused — and probed — across tenants inside the boundary. AWS explicitly recommends tenant-specific prefixes where application tenants must not share cached state.
  5. The demonstrated attack class is timing side channels, and it is real research. Cache-hit responses are measurably faster, so an attacker who can reach a shared cache can probe candidate prefixes and infer whether specific text exists. Documented work includes The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving Systems (KV-cache hit/miss classification and end-to-end prompt-stealing in a vulnerable serving stack — Academic); Auditing Prompt Caching in Language Model APIs (hit-versus-miss latency leakage at commercial APIs — Academic); and I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM Serving (PROMPTPEEK) (prompt reconstruction from side-channel signals — Academic). A second, separate mechanism — semantic cache poisoning (When Cache Poisoning Meets LLM Systems, NDSS 2026 — Academic) — applies to semantic/response caches that return answers for similar rather than identical queries, not to exact-prefix KV caches. Do not conflate the two.
  6. Cache TTL is not the only retention window. Abuse and safety monitoring logs (for example, up to 30 days for non-ZDR accounts on OpenAI), files, and conversation state are retained separately. A 5-minute cache tells you nothing about the rest of the retention surface.

6. Mitigations, in order of leverage

  1. Keep secrets, personal data and tenant-specific records out of long-lived, shared prefixes; put per-request sensitive content after the cache breakpoint — it is never cached.
  2. Use the narrowest isolation boundary available — separate projects, workspaces or subscriptions per security tenant, rather than relying on one shared organisation.
  3. For strict ZDR or regulated data, choose the provider-feature combination deliberately (Anthropic in-memory caching; Vertex default cache, which can be disabled at project level; OpenAI in_memory retention where the model supports it), and get a written provider answer for the specific model — this is where the exceptions live.
  4. Prefer explicit cache objects where offered (Gemini) when deletion control and guaranteed savings matter; prefer implicit or automatic caching when frictionless cost reduction matters and variable hit rates are acceptable.
  5. Use tenant-specific cache keys or prefixes in multi-tenant products; monitor for unexpected hit patterns and latency-probing behaviour.
  6. For residency, confirm the cache’s location boundary, not the endpoint name: OpenAI caches cannot cross regional-processing boundaries; Azure keeps extended caches within data-zone or regional boundaries; Bedrock cross-region inference profiles can process in another region within the selected geography.

7. Caveats and how to read this brief

  • Per-model cached rates, write premiums and TTLs change frequently — the structural statements in the pricing table are documented by providers, but the numbers quoted are directional (third-party tier) and must be verified on provider pricing pages before entering a business case.
  • The security research is cited from the papers’ reported content as gathered in this research pass; the papers’ full texts were not retrieved, so treat attack specifics as needing the original papers before design decisions.
  • Provider isolation claims are exactly that — vendor statements. The brief reports them as documented; it does not independently verify enforcement.

Sources

Provider documentation: – OpenAI: Prompt Caching in the API; Prompt caching guide; Prompt Caching 101; Prompt Caching 201; API changelog; Better prompt caching for GPT-6; API pricing – Anthropic: Prompt caching — Claude docs – Google: Context caching — Gemini API; Gemini 2.5 implicit caching; Vertex AI context caching – Microsoft: Prompt caching with Azure OpenAI in Microsoft Foundry – AWS: Bedrock prompt caching; Bedrock prompt caching product page; Effectively use prompt caching on Amazon Bedrock – DeepSeek: Context Caching — API docs; Context Caching on Disk announcement

Security research (titles as reported in the research pass; individual paper URLs were not retrievable from this environment’s search tooling): – The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving Systems – Auditing Prompt Caching in Language Model APIs – I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM Serving (PROMPTPEEK) – When Cache Poisoning Meets LLM Systems: Semantic Cache Poisoning and Its Countermeasures (NDSS 2026)

Leave a Reply

Your email address will not be published. Required fields are marked *