How to Build a Semantic Cache for LLM Apps

Exact-match caching misses paraphrases, so LLM bills stay high. Here is how to build a semantic cache with embeddings, a similarity threshold, and its traps.

Two users ask your app the same thing an hour apart. One types “how do I reset the FNAL transfer queue?” and the other types “what’s the way to clear the transfer queue at Fermilab?” The answer is the same. A normal cache, though, treats them as two different keys, misses on the second, and pays for a full model call it did not need. Do that a few thousand times a day and the bill is real. So is the second or two each user waits while the model regenerates a paragraph it already wrote for someone else.

Exact-match caching barely helps because people rarely phrase the same question the same way. A cache keyed on the literal string only fires when two requests match character for character, and for free-form questions that almost never happens. What you want is a cache that fires when two questions mean the same thing. That is a semantic cache.

This post is for engineers running an LLM feature in production who have watched the token bill climb and want to stop paying twice for the same answer. It covers the request flow, a small implementation, and the one parameter that decides whether the cache helps or hurts. It also covers the failure mode that will serve a wrong answer with full confidence if you are not careful.

Why an exact-match cache doesn’t fire

The standard move is to hash the prompt and use the hash as a key in Redis or a dict. That works perfectly for deterministic inputs, which is why it caches database queries and rendered templates so well. LLM prompts are the opposite. The “key” is a sentence a human wrote, and humans write the same intent a dozen ways: different word order, a synonym, a typo, an extra “please.” Every variation hashes differently, so every variation is a miss.

A related feature needs separating out here, because the names collide. Providers offer prompt caching: they cache a long shared prefix (a system prompt, a big document) on their side, so you are not billed full price to re-process it on every call. Anthropic and OpenAI both do this. It is useful and you should turn it on, but it is a different thing. Prompt caching discounts the tokens you send, while a semantic cache skips the call entirely when someone has effectively asked before. The two stack.

How a semantic cache matches by meaning

The trick is to stop comparing strings and start comparing meaning. You represent each query as an embedding: a vector, produced by a model, in which sentences with similar meanings land close together. You then look up each new query by nearest-neighbour search instead of an exact key match. If the closest previous query is close enough, you serve its stored answer. If nothing is close enough, you call the model and store the new query and answer as a pair.

Request flow of a semantic cache: an incoming query is embedded to a vector, the nearest cached query is found by vector search, and a cosine similarity threshold decides between a cache hit that returns the stored response in milliseconds and a cache miss that calls the model and stores the new pair

Embeddings are the same tool behind retrieval in a RAG system. If you have read my posts on chunking or HNSW vector search, the machinery is identical: encode text to a vector, index the vectors, query by nearest neighbour. A semantic cache points that machinery at your own past queries instead of a document corpus.

A minimal implementation

Start with the smallest thing that shows the idea. The class below encodes each query with a small sentence embedding model and normalizes the vector to unit length, so a dot product equals cosine similarity (how closely two vectors point in the same direction). It keeps the stored answer next to each vector, at the same index. get scores a new query against every cached one and returns the best match only if it clears the threshold.

from sentence_transformers import SentenceTransformer
import numpy as np

encoder = SentenceTransformer("all-MiniLM-L6-v2")

class SemanticCache:
    def __init__(self, threshold=0.92):
        self.threshold = threshold
        self.keys = []      # query embeddings, unit length
        self.values = []    # stored responses, same index as keys

    def _embed(self, text):
        v = encoder.encode(text)
        return v / np.linalg.norm(v)   # so dot product == cosine similarity

    def get(self, prompt):
        if not self.keys:
            return None
        q = self._embed(prompt)
        sims = np.stack(self.keys) @ q       # cosine to every cached query
        i = int(sims.argmax())
        if sims[i] >= self.threshold:
            return self.values[i]
        return None

    def put(self, prompt, response):
        self.keys.append(self._embed(prompt))
        self.values.append(response)

Then the cache wraps the model call, and nothing else in your code has to know it is there:

def ask(prompt, cache):
    hit = cache.get(prompt)
    if hit is not None:
        return hit                    # milliseconds, no tokens spent
    answer = call_llm(prompt)         # the expensive path
    cache.put(prompt, answer)
    return answer

The linear scan with argmax is fine for a few thousand entries, and it is honest about what is happening: score the new query against every cached one and take the best. Past that it gets slow, and you swap the Python list for a real vector index. In production that is a job for Redis with its vector search, an HNSW index, or a purpose-built layer like the open-source GPTCache, which packages exactly this pattern with pluggable stores and eviction. The logic above is the whole idea; those tools handle scale and persistence.

The similarity threshold decides everything

The threshold=0.92 in that constructor is not a detail. It is the one number that decides whether the cache saves you money or quietly corrupts your answers, and no default is right for everyone.

The similarity threshold as a tradeoff line: too low a threshold gives a high hit rate but false hits where opposite questions match, too high a threshold eliminates false hits but drops the hit rate to near zero, and a working range in the middle must be tuned on real query logs

Set it too low and everything looks like a hit. “How do I enable the site whitelist?” matches “how do I disable the site whitelist?”, because by embedding those sentences are more than 90% similar. You then serve the wrong instruction to a user who asked the opposite question. Set it too high and the cache only fires on near-identical rewordings. Your hit rate collapses, and you pay for the embedding call on top of the model call you failed to avoid.

The right number depends on your embedding model and on how tightly your real queries cluster, so you measure it instead of guessing:

  1. Pull a few hundred real query pairs from your logs.
  2. Label which pairs are genuinely the same question.
  3. Find the threshold that catches the true matches without letting the near-misses through.

This is the same “where do I draw the line” problem I ran into tuning retrieval quality for Archi, and the answer is the same: a small labeled set beats an educated guess every time.

The failure mode that bites: confident false hits

Misses are not the dangerous errors. A miss just costs you a model call you hoped to skip. The costly error is the false hit: two questions sit close in embedding space but need different answers, the cache confidently returns the wrong one, and nobody sees a stack trace.

Three shapes of this show up over and over:

  • Negation. Think “enable” and “disable,” “allow” and “block,” “include” and “exclude.” Embedding models are famously weak at negation because the sentences share almost every word, so the vectors sit very close while the correct answers are opposites. This is the one to fear most.
  • Named entities. “reset the FNAL queue” and “reset the CERN queue” differ by one token that carries the entire meaning. High similarity, different answer.
  • Numbers and parameters. “retry after 30 seconds” and “retry after 300 seconds” are the same sentence with one changed value, and the value is the whole point.

None of these throw an error. They return a plausible answer that happens to be wrong, which is worse than a crash, because a crash at least tells you something broke. The mitigation is partly the threshold and partly scope: do not rely on similarity alone to separate meanings that hinge on a single token. Wherever a wrong-but-confident answer causes real harm, either raise the threshold hard or keep those routes out of the semantic cache entirely.

Scope the keys, and expire stale answers

Two more things separate a toy from something you would run.

Scope the keys. A cache key is more than the query text. If two users are in different projects, or you run two model versions, or answers are personalized, a hit from the wrong bucket is a correctness bug and sometimes a data-leak bug. Namespace the cache by whatever changes the correct answer: model version, tenant, user, prompt-template revision. A cheap way to do it is a separate vector store per namespace, so a lookup can only ever match within its own bucket.

Expire stale answers. Semantic caches have the classic caching problem in a sharper form: the world moves and the cached answer does not. A regular cache keyed on a database row can be invalidated when the row changes. A semantic cache keyed on a question has no idea that the answer to “what is the current recommended retry policy?” changed last week when someone updated the runbook. Put a TTL (time to live) on entries so answers expire, keep it short for anything time-sensitive, and clear the cache when the underlying knowledge changes.

This is the same coupling as the cache-versus-source problem I studied years ago in a cache size and miss-rate experiment. A cache is only as good as its agreement with the source of truth, and a stale hit is just a fast wrong answer.

When it pays off, and what I would do differently

A semantic cache earns its keep when your traffic has real repetition and a wrong hit costs little: FAQ-style support, documentation Q&A, anything where the same handful of intents come up constantly. It earns nothing when every query is unique, because you pay for an embedding on every request and almost never hit. On flows where a subtly wrong answer causes harm, it is actively dangerous, so leave it off there.

If I were setting one up again, I would ship it in shadow mode first. Run the lookup on every request and log what the cache would have returned and how similar the match was, but keep serving live model calls. A day of that log tells you your real hit rate. More importantly, it surfaces the near-miss pairs around your threshold, which is exactly where the false hits hide. I would rather find “enable” matching “disable” in a log than in a support ticket.

I would also start the threshold deliberately high, closer to 0.95, and loosen it only once the shadow data shows I am leaving safe hits on the table. A conservative cache saves less and breaks nothing. A loose one saves more, until the day it serves the opposite of what someone asked.

Where a semantic cache fits

Most of the cost and latency in the LLM tools I work on comes from one place: calling the model when you did not have to. A semantic cache is one of the cheaper ways to cut that. It sits naturally next to the rate-limit and backoff handling I wrote about earlier, since a request served from cache is a request you never have to retry. The token accounting in LLM DevMate and the retrieval pipeline behind Archi, the RAG copilot I built for CMS computing operations at CERN, both live and die on not repeating expensive work. Caching the answer to a question someone already asked is the most direct version of that.

The catch is that a cache which occasionally returns the opposite of what was asked is worse than no cache at all. That is why this post spends more words on the threshold and the false hits than on the code. Get the threshold and the scoping right first; the savings are the easy part.

Image credit: diagrams by M. Hassan Ahmed, created for this post and released under CC0. Example queries reference the Archi project.