What actually gets cached
Three different things get called “the cache” in these conversations, and conflating them is how teams end up with a stale-answer bug they cannot trace. They have different keys, different hit-rate ceilings, and different invalidation costs.
These are not competing options: prompt caching makes the model call cheaper and faster, answer caching makes it not happen. The answer-caching variant is called cache-augmented generation, or CAG — what the Streamlit + vLLM + context caching walkthrough builds. It is RAG’s sibling, not its enemy: I front-load knowledge into a cached prefix instead of retrieving at query time.
The three ways to share one cache across users
Method 1 — Local inference with vLLM Automatic Prefix Caching
When I host an open-source model locally or on my own cloud GPUs, vLLM automatically caches the Key-Value (KV) blocks of shared prompts, such as long system instructions or base documents, across different incoming user requests. I enable it at startup with the--enable-prefix-caching flag:
1 / replicas.
The practice that decides whether any of this works: keep the shared text (system prompt, few-shot examples, or reference data) completely identical and at the very beginning of the prompt string for every request, as this caching walkthrough stresses. Prefix caching is prefix-matched, not substring-matched. One changed character — a timestamp, a request ID, a user name interpolated early — and every user pays the full prefill again. It is the most common way I see teams silently get zero benefit.
Method 2 — Cloud managed APIs with a shared cache ID
With managed APIs such as Google Vertex AI, I create a central cache object once and distribute its reference identifier to many user calls, as described in this Gemini context caching walkthrough:- Create a central cache: upload the large files or long system prompts to the API to generate a
cache_id. - Share the cache ID: pass that same
cache_idin the initialisation code for all subsequent user queries.
cache_id read from the same cloud-stored token memory, reducing cost and latency without duplicating the upload per user. Two properties I plan for: the cache is billed for storage while it exists, and it has a TTL I must refresh. If the renewal job fails, traffic does not error — it quietly reverts to full-price prefill, so I alert on cached-token ratio, not on cache errors.
Method 3 — Application-level exact and semantic caching
If I want to share previous responses and conversation prompts at the application layer rather than the raw GPU layer, I use a caching proxy or a database like Redis. The strategies are summarised in this LLM caching strategies overview and in Azure’s semantic caching guide. Exact-match caching hashes the incoming query; semantic caching runs a similarity search over previously asked questions and returns a stored answer when the distance is close enough. User A’s common question is saved to Redis, and User B’s matching or semantically identical question gets that cached answer with no model call (same ordering idea).The lookup path I actually implement
Ordered by cost per hit, cheapest first, and every level records which one served the request:1
Exact answer cache
Normalised question plus corpus version as the key. Sub-millisecond, no model call, and the only level that also skips retrieval.
2
Semantic answer cache
Query embedding compared against recent questions. One embedding call, no LLM call. I start at 0.92 cosine and tune against labelled pairs.
3
Cached prefix, live generation
The context sits in a cache object or an APC block, so prefill is skipped and decode still happens.
4
RAG fallback
Retrieve top-k chunks and generate. The home for long-tail questions.
5
Write-through
Store the answer and its embedding under the corpus version that produced it, so a document update can evict it.
Hit-rate math
Leth be the fraction of requests served without a model call. Cost per query becomes E[cost] = h × lookup_cost + (1 − h) × pipeline_cost, and latency behaves the same way. So the whole business case collapses onto one number: how repetitive is my traffic?
FAQ traffic is far more repetitive than people expect. A rough Zipf assumption — the top 20% of distinct questions carrying about 80% of volume — gives me a planning figure of h ≈ 0.7 for exact match on a hotel-support bot, with semantic match buying another 5–10 points. Normalising the question (strip punctuation, lower-case, collapse whitespace, move the user name and property ID into metadata instead of the query text) is what makes the exact key hit at all.
Now the arithmetic on numbers I use elsewhere in this repo: 71,220 daily active users, 5,000 unique input tokens and 1,500 output tokens per query, GPT-5.4 at 15.00 per million output (User usage at scale):
Two conclusions. Cached context is 6x more expensive per query than a 5,000-token RAG payload on the same model — caching does not make a huge context free, it makes it merely survivable. Answer caching changes the shape of the curve: at
h = 0.5 the effective RAG per-query cost is 0.0070, thirty times cheaper than a single 0.7M-token cached read. If questions repeat, caching the answer beats caching the context by a wide margin.
Latency splits the same way. A Redis answer hit is a sub-millisecond read plus network. A cached-prefix hit skips prefill. A RAG round trip is an ANN lookup — a scalable vector store is expected to answer in sub-100 milliseconds, per Scalable vector database — plus a short generate. An uncached long context is the worst case: 10 to 30 seconds of prefill at 1M tokens, 2 to 5 minutes at 10M, as recorded in Why need RAG if gemini can answer.
Invalidation and staleness
This is where cached-context systems fail in production, and none of it is model-specific.- Version the key. Put a corpus version (or a content hash of the retrieved context) into the cache key, so a document move evicts the answers that depended on it.
- Prefix mutation is a wipe. Editing the cached system prompt or reference set invalidates every derived KV block. Publish the change as a new
cache_idand roll traffic over, rather than mutating in place and serving two versions at once. - Mind the TTL and the renewal job. Managed caches expire; APC blocks are evicted under memory pressure. Both failures are silent and both look like “the model got slow and the bill doubled”.
- Beware semantic poisoning. A semantically matched answer served to a different question is worse than a miss —
What time is check-in?andWhat time is check-out?are neighbours. Keep the threshold high, store the question text with the answer, and require agreement on the discriminative tokens before serving a hit. - Never cache personalisation. If the answer depends on who is asking, the key must include the user, plan tier, and entitlements — at which point the hit rate collapses and answer caching stops paying.
- Never cache the unverifiable. Anything time-sensitive (availability, price, weather, room inventory) is excluded by policy, not by TTL luck.
When this beats RAG
I choose cached context or cached answers over a retrieval pipeline when most of these hold:
The rule I use: if the corpus fits the window and the traffic is repetitive, cache; if the corpus is bigger than the window or the questions are novel, retrieve; and put answer caching on top of RAG rather than instead of it. For one analyst asking one question of a codebase, dumping everything into a cached context is genuinely better — no index to build, no chunking to tune, no recall to lose. For 71,220 users a day the same architecture runs to $448,686 a month, which is why
/user-usage slices the 700k payload with vector search and loads about 20k tokens per prompt.