Tldr
Three related mechanisms are often called caching, but they operate at different scopes:
KV cache β reuses attention state within one generation.
Prefix caching β reuses previously computed prefix state across requests in a serving engine.
Prompt caching β provider-managed prefix caching exposed through an API, usually with cache-control, TTL, usage, and pricing semantics.
All three avoid recomputing tokens the model has already processed. Assuming the cached state is valid and reused exactly, they do not change model generation semantics.
The relationship is roughly:
KV cache is the underlying state β prefix caching reuses that state across requests β prompt caching exposes that reuse as a managed API feature.
Comparison
| KV cache | Prefix caching | Prompt caching | |
|---|---|---|---|
| Whatβs cached | Attention K/V tensors for previous tokens | Cached model state / KV blocks for a shared prefix | Provider-managed state for a reusable prompt prefix |
| Scope | Within one generation | Across requests | Across API requests |
| Main phase saved | Decode | Prefill | Prefill |
| Typical layer | Model runtime | Serving engine | Hosted model provider |
| Examples | vLLM, llama.cpp, TensorRT-LLM | vLLM prefix caching, SGLang RadixAttention | OpenAI, Anthropic, DeepSeek |
| Control | Usually automatic | Engine/config dependent | Provider dependent |
| Primary benefit | Faster decode | Lower TTFT + higher throughput | Lower latency and/or input cost |
| Cache miss | No reusable prior state | Prefix is recomputed | Prefix is processed at normal/cache-write rates |
| Changes model semantics? | No | No | No |
The key relationship
1. KV cache: reuse inside one request
During prefill, the model processes the prompt and computes attention keys and values for its tokens.
Prompt
β
βΌ
Prefill
β
βΌ
KV cache
β
βββ generate token 1
βββ append its K/V
βββ generate token 2
βββ append its K/V
βββ ...Without a KV cache, every decode step would need to recompute attention state for tokens the model has already processed.
With it, each new token computes its new state and attends to the previously stored state.
So:
KV caching primarily accelerates decode.
2. Prefix caching: reuse across requests
Suppose many requests begin with the same large prefix:
[system prompt]
[tool definitions]
[few-shot examples]
[user-specific message]Request A computes:
shared prefix ββprefillβββΊ cached model stateLater, Request B starts with the same prefix:
shared prefix ββββββββββββΊ reuse cached state
new tokens ββprefillβββΊ compute uncached suffixInstead of rerunning the entire shared prefix, the serving engine restores or references its cached state and continues from there.
So:
Prefix caching primarily accelerates prefill.
This can improve:
- TTFT β time to first token
- GPU throughput
- compute efficiency
- cost per request in self-hosted systems
Prefix reuse ends when the cacheable representation of the requests diverges.
For example:
A: System + Tools + Few-shot + User A
B: System + Tools + Few-shot + User B
β²
reusable prefix ends hereA change near the end loses little reuse.
A change near the beginning can destroy most of the cacheable prefix.
3. Prompt caching: managed prefix caching
Hosted providers expose versions of this optimization as prompt caching or context caching.
Conceptually:
[large stable prefix]
[small changing suffix]becomes:
stable prefix βββΊ cache hit
changing suffix βββΊ normal prefill
β
βΌ
decodeDepending on the provider, the API may expose:
- cached-input token counts
- separate cache-write and cache-read prices
- automatic prefix detection
- explicit cache breakpoints
- TTL controls
- cache-routing keys
- cache-hit/miss telemetry
So the useful mental model is:
Prompt caching is provider-managed prefix reuse with an API and billing contract around it.
It is closely related to prefix caching, but you should not assume every provider implements storage, matching, eviction, or cache boundaries identically.
Prefill vs. decode
This distinction makes all three mechanisms easier to remember.
REQUEST
Prompt tokens
β
βΌ
βββββββββββ
β PREFILL β β Prefix / prompt caching helps here
ββββββ¬βββββ
β
βΌ
KV state
β
βΌ
ββββββββββ
β DECODE β β KV cache helps here
ββββββββββ
β
βΌ
Generated tokensPrefill
The model processes the input context and constructs the model state needed for generation.
Long prompts make prefill expensive.
Prefix/prompt caching reduces repeated prefill work.
Decode
The model then generates tokens autoregressively:
token 1 β token 2 β token 3 β ...Each generated token needs access to previous context.
KV caching prevents the model from rebuilding the previous K/V state at every step.
Practical prompt layout
To maximize prefix-cache reuse, put the most stable content first:
1. System/developer instructions
2. Tool definitions
3. Stable policies/context
4. Few-shot examples
5. Conversation history
6. Current user message
7. Volatile/request-specific metadataPrefer:
[stable 20k-token prefix]
...
Current time: 10:32
User request: ...over:
Current time: 10:32
[stable 20k-token prefix]
...A volatile value near the beginning causes the request to diverge early and can greatly reduce cache reuse.
Important caveats
A cache miss is primarily a performance/cost problem
If prefix caching misses, the model normally processes the prefix again.
cache hit β reuse previous computation
cache miss β perform prefill againThe cache optimization itself is not supposed to alter how the model generates the answer.
βOne changed byte invalidates everythingβ is too simplistic
A better rule is:
Once the cacheable representation diverges, later state generally cannot reuse that previous prefix.
Providers operate on rendered/tokenized model input and may also incorporate relevant request settings, tools, images, reasoning configuration, or provider-defined cache boundaries.
So changing whitespace, tool schemas, message structure, metadata, model settings, or content can affect reuse even when two prompts appear conceptually equivalent.
KV-cache lifetime is implementation-dependent
A normal generation uses K/V state locally while decoding, but βKV cache = request-local foreverβ is not a fundamental property.
Cross-request prefix caching works precisely because inference systems can preserve or otherwise reuse model state beyond the immediate decode loop.
Provider comparison
| OpenAI GPT-5.6+ | Anthropic | DeepSeek | |
|---|---|---|---|
| Enabled | Automatically | Developer opts in with cache_control | Automatically |
| Automatic prefix management | Yes | Yes, after enabling automatic caching | Yes |
| Explicit breakpoints | Yes | Yes | Not the primary API model |
| Read price | 0.1Γ input | Usually 0.1Γ input | Model-specific hit price |
| Write price | 1.25Γ input | 1.25Γ / 2Γ for 5m / 1h | Not exposed as a separate cache-write SKU |
| TTL | 30m minimum for GPT-5.6+ | 5m default or 1h | Best-effort; typically hoursβdays before unused cache clears |
| Minimum length | 1,024 visible tokens on GPT-5.6+ | Model-dependent | Provider-managed persisted prefix units |
| Usage telemetry | cached + cache-write tokens | cache-read + cache-creation tokens | cache-hit + cache-miss tokens |
| Core abstraction | Managed KV-prefix cache | Explicit/automatic cache breakpoints | Automatic disk context cache |
The main takeaway is that βprompt cachingβ is not one standardized API feature.
All three providers are trying to avoid repeated prefill, but the product contracts are different:
OpenAI:
automatic/explicit prefix cache
+ cache-write/read pricing
+ 30m sliding minimum lifetime
Anthropic:
developer enables caching
+ automatic or explicit breakpoints
+ selectable 5m / 1h TTL
+ write/read pricing
DeepSeek:
automatic disk context cache
+ provider-managed persisted prefix units
+ cache-hit / cache-miss pricingSo avoid writing generic infrastructure code that assumes one providerβs cache semantics apply to another.
Donβt confuse this with response caching
Response caching is a different optimization.
Prompt caching:
prompt
β
reuse model state
β
model still generates a new answer
Response caching:
prompt/query
β
cache lookup
β
return an existing answerSemantic response caches may even reuse an answer for a merely similar request.
That introduces a different class of correctness risk:
- stale answers
- incorrect semantic matches
- user/context leakage
- bypassing fresh model inference
KV, prefix, and prompt caching reuse computation.
Response caching reuses the answer itself.
One-line definitions
KV cache: Donβt recompute previous-token attention state while generating the next token.
Prefix caching: Donβt recompute the same prompt prefix when another request has already processed it.
Prompt caching: Provider-managed prefix caching exposed through API controls, TTLs, telemetry, and pricing.
Response caching: Donβt run the model if an acceptable answer is already cached.