Build a Semantic Cache for LLM Apps

Semantic caching for LLM apps: cache answers by embedding similarity in FastAPI, tune the cutoff, and avoid false cache hits. Working code and failure modes.

Two operators open the copilot within a minute of each other. One types “why is the transfer to T1USFNAL failing,” and the other types “T1 US FNAL transfers keep failing, what’s wrong.” It is the same question in different words. Both requests go out to the retrieval pipeline and the model, and each costs a few cents and a few seconds. Both come back with the answer that was already sitting in memory from the first one. A plain response cache never caught the second query, because the strings did not match.

A semantic cache fills that gap. Instead of keying on the exact request, it keys on what the request means, so the near-duplicate hits the cache and skips the expensive path. I built one into Archi, the retrieval copilot I worked on for CMS computing operations at CERN, where recurring incidents mean recurring questions.

This post is for engineers running an LLM or RAG backend who are watching their token bill and their p95 latency (the time within which 95% of requests finish). You probably already know that a real chunk of your traffic is the same handful of questions asked slightly differently. It covers the read path, a working FastAPI implementation, how to pick the similarity cutoff, and the ways a semantic cache quietly serves the wrong answer.

One thing to clear up first, because the names collide. This is not the KV cache that lives on the GPU during a single generation. It is also not the provider-side prompt caching that discounts a repeated prefix. It is an application-level cache of whole answers that sits in front of your model, and it is yours to build and tune.

Why exact-match caching leaves most of the wins on the table

The obvious first move is a dictionary or a Redis key on the normalized request string: lowercase it, strip whitespace, hash it, done. That catches the literal reload, such as the same user pressing enter twice or a retry after a dropped connection. It is worth having as a cheap first layer.

It just does not catch much. Natural-language questions almost never repeat verbatim across users. People change the word order, add “please,” paste a different error line, or abbreviate. Under exact matching every one of those is a miss, even though the answer is identical. In the operations traffic Archi sees, nearly all of the interesting overlap is near-duplicates rather than exact ones, so an exact-match cache leaves most of the savings on the table.

Semantic caching closes that gap by comparing meaning. You turn the question into an embedding, the same kind of vector you already compute for retrieval, in which questions with similar meanings land close together. A new question is a hit if some previously cached question sits close enough to it in vector space. “Close enough” means cosine similarity above a cutoff you choose. Choosing that cutoff well is the entire job, and I get to it below.

The read path: embed, look up, fall through

Here is the whole flow on one page. A request comes in and you embed it. You look for the nearest cached question, and you fall through to the model only if nothing near enough exists.

The semantic cache read path: the incoming question is embedded, a nearest-neighbour search runs against the cache vector index, and if the top match clears the similarity threshold the cached answer is returned in about ten milliseconds with no model call; otherwise the LLM or RAG pipeline runs, its answer is stored keyed by the question embedding, and the fresh answer is returned.

The shape is deliberately the same as a retrieval index, because it is one. Your cache is a small vector store of (question embedding → answer) pairs. A lookup is a top-1 nearest-neighbour query (find the single closest stored question) with a distance gate in front. If you have read the post on how HNSW makes vector search fast, you already know the machinery. The cache is just a second, smaller index used for a different purpose.

A minimal implementation in FastAPI

Here is a working cache in about thirty lines, using sentence-transformers for the embedding and cosine similarity for the match. I kept it in-process and in-memory to show the logic; the production notes below say what changes. get returns the answer stored for the closest cached question if it clears the threshold, and None otherwise. put stores a new question and its answer.

import numpy as np
from sentence_transformers import SentenceTransformer

embedder = SentenceTransformer("all-MiniLM-L6-v2")

class SemanticCache:
    def __init__(self, threshold: float = 0.92):
        self.threshold = threshold
        self.vectors: list[np.ndarray] = []
        self.answers: list[str] = []

    def _embed(self, text: str) -> np.ndarray:
        v = embedder.encode(text, normalize_embeddings=True)
        return np.asarray(v, dtype=np.float32)

    def get(self, question: str) -> str | None:
        if not self.vectors:
            return None
        q = self._embed(question)
        sims = np.dot(np.vstack(self.vectors), q)  # cosine, vectors are unit-norm
        best = int(np.argmax(sims))
        if sims[best] >= self.threshold:
            return self.answers[best]
        return None

    def put(self, question: str, answer: str) -> None:
        self.vectors.append(self._embed(question))
        self.answers.append(answer)

Because the embeddings are normalized to unit length, a dot product is the cosine similarity, which keeps the lookup to a single matrix-vector multiply.

Wiring the cache into an endpoint is the part that matters. It has to sit before the expensive call and record the result after it:

from fastapi import FastAPI

app = FastAPI()
cache = SemanticCache(threshold=0.92)

@app.post("/ask")
async def ask(question: str):
    hit = cache.get(question)
    if hit is not None:
        return {"answer": hit, "cached": True}

    answer = await run_rag_pipeline(question)  # the slow, paid path
    cache.put(question, answer)
    return {"answer": answer, "cached": False}

That is the entire idea. The embedding call is cheap and local next to a model round trip. On a miss you have added a few milliseconds to a path that was already going to take a couple of seconds. On a hit you skip the whole thing.

Two production notes:

  • Don’t block the event loop. SentenceTransformer.encode is synchronous CPU or GPU work, and calling it straight from an async def handler blocks the event loop. That is the exact mistake I wrote up in FastAPI’s blocking-the-event-loop trap. Run it in a thread pool with run_in_executor, or hand embeddings to a batched worker.
  • Use a real vector store. A Python list is fine for a demo and hopeless past a few thousand entries. Back the store with FAISS, a managed vector database, or a purpose-built LLM cache like GPTCache or RedisVL, so the lookup stays sub-linear and survives a restart.

Choosing the similarity cutoff

Everything above is plumbing. The one number that decides whether the cache helps or hurts is the similarity cutoff, and it trades two failure modes against each other.

A chart of the two competing rates against the cosine-similarity cutoff. As the cutoff loosens toward 0.80 the false-hit rate (returning a cached answer to a question that only looks similar) climbs steeply, while cache misses fall. As the cutoff tightens toward 0.98 false hits vanish but misses climb, meaning more paid model calls. A workable band sits in the middle, and the caption notes the right cutoff depends on real traffic and should be measured, not guessed.

Set the cutoff too loose and you serve confident wrong answers. “How do I restart the transfer agent” and “how do I stop the transfer agent” are one word apart and embed close together, but the correct responses are opposites. A cache that returns the first answer for the second question is worse than no cache, because the user trusts it.

Set the cutoff too strict and almost nothing clears the bar. Near-duplicates fall through to the model, and you have paid for an embedding index that rarely fires.

There is no universal right number. It depends on the embedding model, the domain, and how varied your traffic is. To find it:

  1. Log real questions and embed them.
  2. Look at the similarity distribution of pairs you know share an intent, and of pairs you know differ.
  3. Pick the cutoff that separates the two, and lean strict when a wrong answer is expensive.

For most sentence-embedding models on short questions I start around 0.9 and adjust from there. Treat that as a starting point to measure against, not a default to ship.

Where a semantic cache breaks

The read path is short. The interesting engineering is in the ways a semantic cache quietly serves the wrong thing.

Negations and antonyms collapse in embedding space. “Include failed jobs” and “exclude failed jobs” are near neighbours to most embedding models, even though they ask for opposite results. This is the sharpest edge, and no threshold fully removes it. If negations carry weight in your domain, either raise the cutoff hard or add a cheap lexical guard that refuses a hit when the two questions disagree on a negation token.

Context makes two identical questions different. “What changed last night” depends on who is asking and when. If the answer is a function of the user, their permissions, the current time, or the site they operate, then the question string alone is the wrong key. Fold the parts of the context that change the answer into the cache key: namespace the cache per user or per site, or bucket by time, so a shared question does not leak a personalized answer.

Cached answers go stale. An LLM answer is only as fresh as the retrieval behind it. When the underlying documents change, the cached answer becomes wrong, and the cache keeps serving it, fast and confident. You need an eviction policy: a time-to-live (TTL) on entries, or invalidation tied to the same indexing pipeline that updates your documents. A stale hit is a false hit that used to be true.

Changing the embedding model invalidates the whole cache. Similarities are only comparable within one model’s vector space. Swap all-MiniLM-L6-v2 for a larger model and every stored vector is meaningless against new queries. Version the cache by embedding model and rebuild it on a change, the same discipline you would apply to your main retrieval index.

A cache hit skips your safety and logging path too. If your normal request flow runs moderation, redaction, or audit logging around the model call, returning early from the cache can bypass all of it. Make sure the early return still records the interaction and still applies whatever guards a fresh answer would have gone through.

Tradeoffs, and what I would actually cache

A semantic cache earns its place when a real fraction of your traffic is genuine near-duplicates and the answers stay stable for a while: recurring operational questions, documentation lookups, the same handful of “how do I” queries. It is the wrong tool in three situations:

  • almost every request is unique;
  • answers depend heavily on live state;
  • a single wrong answer is costly enough that the false-hit risk outweighs the savings.

In those cases the exact-match layer is still worth keeping, and the semantic layer is not.

The honest tradeoff is that you are approximating equality. Approximate equality of meaning is a genuinely hard problem, and embeddings only mostly solve it. The failure mode is not a slow answer but a confident wrong one, which is worse. So I keep the cutoff on the strict side, namespace aggressively by anything that changes the answer, and give entries a TTL rather than trusting them forever.

For Archi the fit is good. Operations questions cluster hard around recurring incidents, and the corpus changes on a known schedule I can tie invalidation to. The same instinct runs through the FastAPI streaming and retrieval work across the rest of these posts, and through projects like Gemini Alchemy. Find the part of the request path that is doing repeated, redundant work, and stop doing it. Do it carefully, though, because the shortcut is only worth it if the answer it returns is still right.


Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.