LLM-as-a-Judge: Scoring RAG Answer Quality

Retrieval metrics say the right docs came back, not that the answer is right. Build an LLM-as-a-judge to score RAG answers for faithfulness and quality.

A while back I wrote about measuring RAG retrieval quality with recall@k (did a right chunk land in the top k results) and MRR (mean reciprocal rank: how high the first right chunk ranked). That post ends with a deliberate cop-out. It grades whether the retriever pulled the right chunks and stops there, because “grading the generated text is a separate exercise.” This is the separate exercise.

The gap is real. Your retriever can score a perfect recall@5, and the answer built on top of it can still be wrong. The model might ignore the chunk that mattered, stitch two facts into a claim that neither supports, or hedge so much it says nothing. Retrieval metrics never see any of that, because they never read the answer. So you are back to reading outputs by hand, which does not scale past the first dozen and quietly stops happening the week you get busy.

This post is for engineers who have a working RAG or LLM app and need to score the answers automatically, without a human in the loop for every run. I’ll cover:

  • what an LLM-as-a-judge actually is;
  • how to write one that grades faithfulness instead of vibes;
  • the biases that make a naive judge useless;
  • how to check the judge against human labels before you trust a single number it produces.

What an LLM judge is, and why it works at all

An LLM-as-a-judge is exactly what it sounds like. You take a second model, hand it the question, the answer, and a rubric, and ask it to score the answer. It sounds circular: if a model can’t be trusted to write a correct answer, why trust another model to grade one?

It works because grading is an easier task than generating. Deciding whether a written answer is supported by a paragraph of context is closer to reading comprehension than to open-ended recall, and models are good at reading comprehension.

The MT-Bench paper by Zheng et al. (NeurIPS 2023) put a number on it. A strong judge model reached about 85% agreement with human raters on open-ended answers, which is higher than the 81% agreement humans reached with each other. That is the whole case for the technique. The judge is not perfect, but it is about as consistent as a second human reviewer, available in seconds, for a fraction of a cent.

The diagram below is the picture I keep in my head. The golden question (a test question from your evaluation set) feeds both halves of the pipeline. Retrieval metrics grade the left half, and the judge grades the text the model wrote from whatever the retriever handed it.

A diagram of a RAG evaluation loop. A golden question flows into a RAG system that retrieves chunks and generates an answer, emitting both the answer and the retrieved context. An LLM judge takes the answer, the context, and a rubric, then scores faithfulness, relevance, and correctness, producing a structured JSON score that is aggregated across the set. A callout notes that you should calibrate the judge against about fifty hand-labeled answers using Cohen's kappa before trusting it, and that retrieval is scored separately.

Grade faithfulness, not “goodness”

The first mistake is asking the judge “is this a good answer?” on a 1-to-10 scale. “Good” is undefined, so the model invents its own standard on every call, and you get noise. Break the vague question into specific ones the judge can actually decide.

For RAG, the single most useful thing to grade is faithfulness: is every claim in the answer supported by the retrieved context? This is the metric that catches hallucination, and it is the one your retrieval numbers physically cannot see.

RAGAS, a widely used RAG eval library, defines faithfulness cleanly:

  1. Decompose the answer into atomic claims.
  2. Check each claim against the context.
  3. Score the fraction of claims that the context supports.

An answer with four claims, three of them grounded, scores 0.75. The nice property is that this is reference-free. You do not need a human-written gold answer, only the context the model was given, which you already have.

Here is that decomposition as a judge prompt. Note that the model is told to extract claims first and judge each one, not to emit a single gut-feel number:

FAITHFULNESS_PROMPT = """You are grading whether an answer is grounded in the
provided context. Do not use outside knowledge.

Context:
{context}

Answer:
{answer}

Steps:
1. Break the answer into a list of standalone factual claims.
2. For each claim, decide if the context directly supports it: yes or no.
3. Report the claims, the per-claim verdicts, and the fraction supported.

Return JSON only:
{{"claims": [{{"text": "...", "supported": true}}],
  "faithfulness": 0.0}}"""

The per-claim supported verdict is doing the real work; averaging over claims is just arithmetic. Forcing the model to lay out its claims before scoring is not decoration either. The G-Eval paper (EMNLP 2023) tested a chain-of-thought, form-filling judge: one that reasons step by step and fills in a structured form. That judge correlates far better with humans than one asked for a bare score. The reasoning constrains the number, instead of the number being pulled from the air.

Get the score out as structured data

A judge that answers in prose is useless for aggregation. You want one JSON object per answer, so you can compute a mean, sort the worst offenders, and diff two pipeline versions.

I already covered the general problem of getting reliable JSON out of an LLM. The same rules apply to a judge, plus one more: define the scale in the schema, because an open integer field invites the score-clustering problem described below.

The schema below gives each answer a faithfulness fraction, a three-step correctness label, a relevance flag, and a short reasoning string:

from pydantic import BaseModel, Field
from enum import IntEnum

class Correctness(IntEnum):
    wrong = 1          # contradicts the context or the known answer
    partial = 2        # some right, some missing or unsupported
    correct = 3        # fully supported and complete

class Judgment(BaseModel):
    faithfulness: float = Field(ge=0.0, le=1.0)
    correctness: Correctness
    relevant: bool           # does it actually address the question?
    reasoning: str           # one or two sentences, for spot-checking

Three anchored fields beat one vague 1-to-10 score every time. A short label on each rung of the scale (“wrong”, “partial”, “correct”) gives the model something concrete to map onto. It also gives you something you can grep for when a score looks off.

The reasoning field is for spot-checking. When you scan the low scores by hand, you see why the judge marked them down, and you can catch the cases where the judge itself was wrong.

Direct scoring or pairwise: pick by what you’re doing

There are two ways to run a judge, and they answer different questions.

Direct scoring grades one answer against an absolute rubric, which is what the code above does. Use it when you want a number you can track over time. “Faithfulness went from 0.82 to 0.88 after I switched the reranker” is a direct-scoring statement, and it is the mode you want for regression testing a pipeline.

Pairwise comparison shows the judge two answers to the same question and asks which is better. It is more reliable when the thing you care about is relative, because “A is better than B” is an easier call than “B deserves a 7.” It is how Chatbot Arena ranks models. Reach for it when you are choosing between two prompts or two models and absolute scores are hard to calibrate. The cost is that it gives you no stable number to trend, and it opens the door to position bias, the first entry in the next section.

The judge lies in predictable ways

A naive judge is confidently biased. The biases are documented well enough that you can design around them instead of discovering them in production. These are measured effects from the papers above, not things I am guessing at. The figure below lists four of them with their fixes.

A table of four LLM-judge biases and their fixes. Position bias, in pairwise judging only, prefers whichever answer is shown first regardless of content; the fix is to run both orderings and keep only agreeing verdicts. Verbosity bias scores longer answers higher even when the extra words add nothing, and repetitive padding fools weak judges; the fix is to score conciseness in the rubric. Self-preference rates a judge's own model family higher than a human would, worst when the judge and generator are the same model; the fix is to judge with a different model than you generate with. Score clustering bunches scores into a narrow band on a wide scale so real differences vanish; the fix is a short, anchored scale.

Three of these deserve more than the one-liner in the figure.

Position bias is the sharp one for pairwise setups. Swap which answer comes first, and the verdict can flip on identical content. The fix is cheap and non-negotiable: judge every pair twice, once in each order, and count the comparison only if both orderings agree. Disagreement is not a tie to break; it is a signal that the judge can’t actually tell these two apart.

Verbosity bias is why a raw “which is better?” often just picks the longer answer. The MT-Bench authors demonstrated it with a “repetitive list” attack: they padded answers with restated points that added no information, and weaker judges preferred the padded version the large majority of the time. If length is not part of what you are grading, put conciseness in the rubric explicitly, so the judge has a reason to push back on padding.

Self-preference is the one people forget. If you generate answers with a model and grade them with the same model, the judge tends to favor its own style, an effect G-Eval flagged directly as a bias toward LLM-generated text. The practical rule I follow: the judge should be a different model from the generator, ideally a stronger one. Grading your own homework is a bias, not a convenience.

Calibrate the judge before you trust it

This is the step that separates an eval you can act on from a dashboard of numbers nobody believes: check the judge against humans, once, on a sample. Skipping it is how teams end up “improving” a metric that never tracked anything real.

  1. Hand-label 40 to 50 answers yourself.
  2. Run the judge on the same set.
  3. Measure agreement with Cohen’s kappa.

Kappa is the right statistic here because it corrects for the agreement you would get by chance. Raw percent-agreement, by contrast, looks great when most answers are “correct” anyway. The snippet below computes kappa with scikit-learn on a toy set of labels:

from sklearn.metrics import cohen_kappa_score

human  = [3, 1, 2, 3, 3, 1, 2]   # your labels
judge  = [3, 1, 3, 3, 2, 1, 2]   # the model's labels
print(cohen_kappa_score(human, judge))   # ~0.6 here

As a rough reading, kappa above 0.6 is substantial agreement, enough to automate with periodic spot-checks. Below 0.4, the judge and you are not measuring the same thing, and you should fix the rubric before trusting a single run.

When agreement is poor, the rubric is almost always the culprit, not the model. Look for vague criteria, an unanchored scale, or a “goodness” question you never turned into something the judge can decide. Tighten those and re-measure.

Tradeoffs, and what I’d watch

An LLM judge is not free, and it is not truth. Every graded answer is another model call, so evaluating a thousand-question set doubles your token bill for that run. The judge also has its own error bar of a few percent, which you carry into every comparison.

That error bar is why a jump from 0.82 to 0.83 means nothing, while a jump from 0.82 to 0.90 across a few hundred questions probably does. Treat judge scores like measurements with noise, not like ground truth, and never make a ship decision on a difference smaller than your calibration error.

What I would do differently on a first build is start with one metric: faithfulness, reference-free, direct-scored, and judged by a model from a different family than the generator. That single number catches the failure retrieval metrics are blind to (hallucination on top of good retrieval), and it needs no gold answers to compute. Add correctness and relevance once faithfulness is calibrated and stable. Resist the pull to grade five dimensions on day one; five noisy metrics are harder to trust than one you have checked against your own eyes.

Why this matters for Archi

This is the eval I care about for Archi, the retrieval copilot I worked on for CMS computing operations at CERN. Retrieval quality tells me the logbook search surfaced the right past incident. It says nothing about whether the answer the operator reads is faithful to that incident or quietly made up a detail. For an ops tool that someone acts on at 3 a.m., that second question is the one that matters. Retrieval metrics grade the plumbing; the judge grades the thing the human actually reads.


Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.