RAG Citations: Make the LLM Cite Its Sources

How to make a RAG system cite its sources: force the model to reference chunks by ID, then verify each citation against the source text before trusting it.

A retrieval-augmented answer that sounds right is worth very little if the person reading it can’t check where it came from. On Archi, the RAG copilot I worked on for CMS computing operations, the questions look like “what do I do when this workflow keeps failing at the merge step?” The answer might be correct. But an operator on shift at 3am will not act on a confident paragraph from a model unless they can click through to the runbook, JIRA ticket, or logbook entry it came from. No source, no trust, no action.

That is the real job of citations in a RAG system. They let a human verify the machine, and they are the difference between a demo and something people rely on. This post is for engineers who already have retrieval working and now need the generated answers to point back at real sources, with citations that are true rather than merely plausible.

It covers why “just tell the model to cite its sources” fails, how to force citations by ID, and the verification step that catches the citations the model invents. There is code, and a section on the ways this breaks in practice, because it breaks in specific and repeatable ways.

Why asking nicely doesn’t work

The naive version is a one-line prompt addition: “cite your sources.” Try it and you get citations. They look great, and some of them are even correct. The problem is the ones that aren’t, and you can’t tell which is which by looking.

Language models generate citations the same way they generate everything else: by predicting plausible tokens. A model that has seen thousands of documents with [3] after a factual sentence will happily produce [3] after its own factual sentence, whether or not source 3 says anything of the kind. The citation is a stylistic pattern the model reproduces, not a lookup it performs. This is the same failure behind fabricated legal citations and invented DOIs: the shape is right, but the referent is fiction.

Two distinct things go wrong, and it helps to name them separately:

  • Fabrication. The model cites a source ID that supports nothing it said, or a source that doesn’t exist in the retrieved set at all.
  • Misattribution. The model says something true, drawn from chunk 2, and cites chunk 5. The claim is fine; the pointer is wrong. A reader who follows it lands in the wrong place and loses faith in every other citation on the page.

Both are common enough that you cannot ship citations you generated but never checked. The academic benchmark work makes this concrete. The ALCE evaluation measures citation precision and recall separately from answer correctness, precisely because models routinely get one right and the other wrong.

So the design has two halves: make the model cite in a form you can machine-check, then check it. The diagram below shows the whole pipeline.

A horizontal pipeline. A retriever produces top-k chunks, which become chunks with stable IDs like runbook#p4, JIRA-8123, logbook#59. Those feed an LLM that returns an answer with inline bracket-n markers. The answer flows into a verifier, and a dashed feedback path shows the verifier re-reading the cited chunk text rather than trusting the model's claim. Verified claims flow to a final answer where each claim links to a real source. A caption notes the dashed verification path is what catches an invented citation before the user sees it.

Step one: cite by ID, not by name

The first fix is to stop letting the model write free-form citations and give it a closed set of IDs to choose from. You control the context you send it, so you control the identifiers.

When you assemble the prompt, tag every chunk with a short stable ID and print that ID right next to its text. The model’s job becomes “pick from these labels,” which is a far easier and more checkable task than “recall a reference.”

In the code below, format_context numbers each chunk and records its source_id, and the system prompt tells the model to cite those numbers after every factual sentence:

def format_context(chunks: list[Chunk]) -> str:
    """Render retrieved chunks with the exact IDs the model must cite."""
    blocks = []
    for i, chunk in enumerate(chunks, start=1):
        # [1], [2], ... are what the model cites; source_id is how we
        # resolve a citation back to a real URL after generation.
        blocks.append(f"[{i}] (source_id={chunk.source_id})\n{chunk.text}")
    return "\n\n".join(blocks)


SYSTEM = """You answer questions for CMS computing operators using only the \
sources provided. After every sentence that states a fact, cite the source it \
came from using its bracket number, like [1] or [2][3]. If the sources do not \
contain the answer, say so. Do not use any knowledge that is not in the sources."""

Two details matter here.

The labels are short and local. The [1], [2] labels are sequential and exist only for this one request. That keeps them short and keeps the model from confusing them with numbers that appear inside the text.

The real IDs stay on your side. Each label maps to a source_id that you keep, so after the model answers, you can turn [2] back into a link to JIRA-8123 or runbook#p4. Never ask the model to emit the real URL or ticket number directly. It will get characters wrong, and then you have a broken link that looks authoritative.

The instruction to refuse when the sources don’t cover the question is doing real work too. Without it, a model that retrieved nothing useful will still answer from its parametric memory (what it learned in training) and cite whatever chunk is closest. That is misattribution by construction. “I don’t have a source for that” is a feature.

Step two: parse the citations out

Inline [n] markers are easy to extract, and it is easy to check that they point at chunks that exist. The function below splits the cited numbers into valid ones and “phantom” ones, meaning numbers outside the range of chunks you actually sent:

import re

CITE = re.compile(r"\[(\d+)\]")

def extract_citations(answer: str, n_chunks: int) -> tuple[set[int], set[int]]:
    """Return (valid_ids, phantom_ids) referenced in the answer."""
    cited = {int(m) for m in CITE.findall(answer)}
    valid = {i for i in cited if 1 <= i <= n_chunks}
    phantom = cited - valid          # e.g. [7] when only 5 chunks were sent
    return valid, phantom

A non-empty phantom set is an immediate, cheap signal that the model is fabricating: it cited a number you never gave it. In practice, I treat any phantom citation as grounds to regenerate the answer. A model confident enough to invent [7] is not one I trust on the citations that happen to land in range.

If you want something sturdier than regex over prose, ask for structured output instead: a JSON array of {claim, source_ids} objects. That moves you onto the ground I covered in getting reliable JSON out of LLMs. It also makes the next step, verification, cleaner, because each claim already arrives paired with its sources.

Some providers now expose this natively. Anthropic’s Citations feature returns cited spans with character offsets into the documents you passed, which removes the parsing problem entirely for that API. When it’s available, use it. The verification logic below still applies either way.

Step three: verify, because the model will lie to you

Extracting a citation only tells you the pointer is in range. It does not tell you the cited chunk actually supports the claim. That is the misattribution case, and catching it is the part people skip and then regret.

The check is a per-claim gate:

  1. Split the answer into claims.
  2. Take each claim’s cited chunk.
  3. Ask a direct question: does this chunk’s text support this sentence?

Read the source, not the model’s summary of it. The diagram below shows the three possible outcomes.

A decision diagram. A box holds a claim and its citation, for example restart the agent on ECAL errors, cited as bracket 2. An arrow labeled overlap, NLI, or LLM judge leads into a diamond asking whether chunk 2 entails the claim. Three outcomes branch out: supported means keep the claim and link the citation to its source; wrong source means re-cite by searching the other chunks for a real match; unsupported means drop or flag the claim and never show it as fact. A caption notes the verifier reads the source text directly, so the model saying bracket 2 is never enough on its own.

There are three ways to implement this entailment check (deciding whether one text supports another), in increasing order of cost and accuracy.

String overlap is the cheap baseline. Does the claim share enough content with the cited chunk, by token overlap or a fuzzy match, to be plausibly grounded? It is fast, needs no model call, and catches gross fabrication, such as a chunk about disk quotas cited for a claim about network timeouts. It misses paraphrase, so treat it as a floor, not a verdict.

Natural language inference (NLI) is the middle ground. A small NLI model classifies the (chunk, claim) pair as entailment, neutral, or contradiction. It handles paraphrase, runs locally, and is cheap enough to call on every claim. This is what the faithfulness metrics in evaluation frameworks lean on. RAGAS computes faithfulness by checking whether each claim in an answer can be inferred from the retrieved context, which is exactly this gate applied as an offline metric.

An LLM judge is the most capable and the most expensive. You hand a model the claim and the cited text and ask for a supported/unsupported ruling with a reason. This is the LLM-as-a-judge pattern pointed at citations specifically, and the same caveats from that post apply: keep the judge’s job narrow, give it a rubric, and don’t ask it to be creative.

Here is the verification loop. The check is a pluggable supports function, so you can start with overlap and swap in something stronger without touching the control flow. Each claim comes out as supported, misattributed (it has valid citations, but none back it), or unsupported (no valid citation at all):

def verify_answer(claims: list[Claim], chunks: dict[int, Chunk],
                  supports) -> list[VerifiedClaim]:
    results = []
    for claim in claims:
        cited = [c for c in claim.source_ids if c in chunks]
        # A claim survives only if at least one cited chunk backs it.
        grounded = any(supports(claim.text, chunks[c].text) for c in cited)
        if grounded:
            results.append(VerifiedClaim(claim, status="supported"))
        elif cited:
            results.append(VerifiedClaim(claim, status="misattributed"))
        else:
            results.append(VerifiedClaim(claim, status="unsupported"))
    return results

What you do with a failed claim is a product decision, not a technical one:

  • Drop it silently. This keeps the output clean but can gut an answer.
  • Flag it (“this part could not be sourced”). This is more honest and, for an operations tool where a wrong instruction has consequences, usually the right call.
  • Repoint a misattributed citation. The cheapest fix is to re-run the overlap check against all the retrieved chunks, not just the cited one, and repoint the citation to the chunk that genuinely matches. Often the model had the right fact and simply grabbed the wrong label.

Where it breaks

The pipeline is straightforward. The failures are where the time goes.

Lost in the middle. Models attend unevenly across a long context, favoring the beginning and end and skimming the middle, a bias documented in the Lost in the Middle study. For citations, this means a claim’s real support might sit in chunk 6 of 10 while the model cites chunk 1, because that’s where it was paying attention. Verification catches this as misattribution, which is one more reason not to skip it, and it argues for retrieving fewer, better chunks rather than stuffing the context.

ID drift after reranking. If you assign the [n] labels and a reranker then reorders the chunks, the numbers the model cites no longer line up with the sources you thought you sent. Assign IDs after every reordering step, immediately before you build the prompt, and freeze them. This bites hardest when you add a cross-encoder reranker to an existing pipeline and forget that it changed the ordering the labels were built on.

Over-citation. Some models cite every chunk after every sentence, [1][2][3][4], which is technically defensible and practically useless: it points at everything and therefore nothing. Verification helps here too. If a claim genuinely follows only from [2], drop the other three citations even though the model offered them, because a citation the reader can’t act on is noise.

The refusal that should have happened. The worst output is a fully cited answer to a question the sources don’t actually address. Every citation passes a loose overlap check because the topic words match, but the specific claim isn’t there. Tighten the entailment check and lean on the retrieval side: if retrieval quality is poor, no citation layer will save you, because the model is choosing the least-wrong of several wrong chunks.

Tradeoffs, and what I’d do differently

Verification costs latency and, if you use a judge, money. On Archi, I don’t verify every claim with an LLM on the hot path. The checks are tiered instead:

  • String overlap runs on everything as a fast fabrication tripwire.
  • NLI is the per-claim gate on live answers.
  • The expensive LLM judge is reserved for offline evaluation runs, where I’m measuring citation precision across a test set rather than gating a single live answer.

That tiered setup keeps the interactive path responsive while still catching the failures that matter most.

If I were starting fresh, I’d reach for a provider’s native citation API before building the parse-and-verify machinery by hand. Character-offset citations remove the misattribution class almost entirely, because the model isn’t picking a label; it’s pointing at a span. The hand-rolled version in this post is what you build when you’re running open models, mixing providers, or need the verification to be inspectable rather than a black box. That describes a lot of production systems, including the ones I work on.

The through-line back to Archi and the wider CMS operations tooling is trust. Operators don’t need the copilot to be right every time; they need to be able to check it every time. A grounded, verified citation is how a RAG answer earns the right to be acted on instead of double-checked from scratch, which is the entire reason the tool exists.


Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.