LLM-as-a-Judge: Evaluating LLM Output Quality

Human review does not scale for grading LLM answers. How to use an LLM as a judge: write a rubric, score with structured output, and control the biases.

You change a prompt, swap a model, or add a reranker, and now you need to know whether the answers got better. For retrieval, that is a solved problem: freeze a set of queries, score the ranked chunks, read the numbers. I wrote that up in measuring RAG retrieval quality.

The generated answer is fuzzier. There is no single correct string, and two good answers can look nothing alike. “Is this a good answer?” is exactly the kind of judgment that string-matching metrics like BLEU or ROUGE (which score word overlap with a reference text) were never built for.

The scalable option people reach for is another LLM. Give a model the question, the answer, and a rubric, and ask it to grade. This works better than it has any right to. It also fails in ways that quietly flatter whatever you are testing, if you do not know they are there.

This post is for engineers who have a working LLM or RAG feature and want to score its output on every change without reading a hundred transcripts by hand. It covers the two grading modes, a judge you can actually run, the three biases that make a naive judge lie, and how I use it on a production copilot.

Two ways to grade: scores and pairwise comparisons

There are two shapes of LLM judging, and picking the wrong one wastes a lot of tokens.

Single-answer scoring hands the judge one answer and a rubric, and asks for a score. It is cheap, it parallelizes, and it gives you an absolute number you can track over time. The catch is that absolute scores drift. A model’s idea of “7 out of 10” is not stable across runs or versions, so the number is only meaningful relative to itself under a fixed rubric.

Pairwise comparison shows the judge two answers to the same question and asks which is better. This is what powers Chatbot Arena and MT-Bench. It is more reliable, because “A is better than B” is an easier call than “B is a 7.” The cost is that it only gives you a relative signal. Comparing N systems means N-squared matchups unless you use a ranking scheme on top.

I default to pairwise when I am choosing between two candidate configurations. I use single-answer scoring when I want a dashboard number that tracks one pipeline over time.

The diagram below shows how the pieces fit for a RAG answer. There, the judge also gets the retrieved context, so it can check whether the answer is actually grounded in it.

Where an LLM judge sits: the question, the model's answer, and (for RAG) the retrieved context all feed a judge model driven by a rubric, which returns a rationale and a score; the two modes shown are single-answer scoring against a rubric and pairwise A-vs-B comparison

A judge you can run

The whole thing is a prompt plus a parser. Two things make it trustworthy: the score comes back in a fixed shape you can compute on, and you ask for the reasoning before the number.

Here is a single-answer scorer. It uses a typed schema so the output always parses, and its rubric describes what earns each point on the scale:

from pydantic import BaseModel, Field

class Judgment(BaseModel):
    reasoning: str = Field(description="Step-by-step assessment against the rubric")
    score: int = Field(ge=1, le=5, description="1 = unusable, 5 = fully correct and complete")

RUBRIC = """You are grading an answer to an operations question.
Score 1-5 on this rubric:
5 - correct, complete, and supported by the provided context
4 - correct but missing a minor detail
3 - partially correct or partially unsupported
2 - mostly wrong or largely unsupported by the context
1 - wrong, empty, or fabricated
Judge only against the context. If the answer states something the context
does not support, that is a fabrication and caps the score at 2."""

def judge(question, answer, context, model):
    prompt = (
        f"{RUBRIC}\n\n"
        f"Question:\n{question}\n\n"
        f"Context provided to the answerer:\n{context}\n\n"
        f"Answer to grade:\n{answer}\n\n"
        "Return your reasoning first, then the score."
    )
    return model.complete(prompt, schema=Judgment)  # constrained to the schema

Two details are doing most of the work here.

The schema asks for reasoning before score. That order is not decoration. When you force the model to write its assessment first, the score it lands on is conditioned on that reasoning, rather than blurted out and rationalized afterward. This is the core idea behind G-Eval, which showed that chain-of-thought judging (the model reasons step by step before scoring) correlates with human ratings noticeably better than asking for a bare number.

The score lives in a fixed field. Reading the number from a field, rather than parsing it out of prose, is what lets a scorer run unattended. A prose parser throws on the first answer the model wraps in a sentence.

For pairwise, the prompt changes to show two answers and ask for a winner, which the schema restricts to A, B, or tie:

class Comparison(BaseModel):
    reasoning: str
    winner: str = Field(pattern="^(A|B|tie)$")

def compare(question, answer_a, answer_b, context, model):
    prompt = (
        "Two answers to the same question are below. Judge which better "
        "satisfies the question using only the provided context.\n\n"
        f"Question:\n{question}\n\nContext:\n{context}\n\n"
        f"Answer A:\n{answer_a}\n\nAnswer B:\n{answer_b}\n\n"
        "Explain briefly, then output the winner as A, B, or tie."
    )
    return model.complete(prompt, schema=Comparison)

That looks complete. It is not, because the judge has opinions that have nothing to do with the answers.

The three biases that make a judge lie

The MT-Bench paper is worth reading in full. It reports that a strong judge like GPT-4 agrees with human preference over 80% of the time, and it names the failure modes precisely. Three of them will bite a naive setup.

Position bias. In pairwise mode, the judge favors whichever answer it saw first, regardless of content. Swap A and B, and a nontrivial fraction of verdicts flip. The fix is mechanical: run every comparison twice with the order reversed, and count a win only if the same answer wins both times. When the two orders disagree, record a tie, which is the honest reading.

Verbosity bias. Judges reward length. A longer, more detailed-looking answer scores higher even when the extra words add nothing correct, and sometimes when they add something wrong. This one is dangerous because it lines up with a real regression: a model change that makes answers wordier will look like an improvement to an unguarded judge. Watch answer length alongside the score, and if both climb together, be suspicious.

Self-enhancement bias. A judge tends to prefer answers from its own model family, so if you generate with GPT-4 and judge with GPT-4, the score tilts toward the home team. When it matters, judge with a different model than the one that produced the answer. At least check that your conclusion survives a second judge from another family.

The three LLM-judge biases and their mitigations: position bias, fixed by running both A-B and B-A orderings and keeping only order-consistent wins; verbosity bias, watched by tracking answer length next to the score; self-enhancement bias, reduced by judging with a different model family than the one that generated the answer

None of these are exotic; they are the default behavior. Every one of them pushes toward a false positive, which is the worst kind of measurement error when you are deciding whether to ship a change.

Judging RAG answers: faithfulness

For a retrieval system, the useful question is not just “is this a good answer?” It is “did the model stay inside the retrieved context, or did it make something up?” That property is called faithfulness, and it is what the Ragas framework measures for RAG pipelines.

The rubric above bakes faithfulness in with its fabrication cap, but you can also score it on its own:

  1. Break the answer into individual claims.
  2. For each claim, ask the judge whether the context supports it.
  3. Take the fraction of claims that are grounded. That is the faithfulness score.

This catches the failure that matters most in operations. An answer can be fluent, confident, and completely detached from the documents it was supposed to summarize. A retrieval metric will not see it, because retrieval did its job and surfaced the right chunks; the model then ignored them. Faithfulness judging is the layer that notices.

Pair it with an answer-relevance check (does the response actually address the question that was asked?). Together, the two cover the ways generation goes wrong independently of retrieval.

Where the judge lies to you

You never validated the judge. An LLM judge is a model, and a model you have not measured is a guess. Before trusting it, hand-label thirty to fifty examples yourself and check the judge’s Spearman correlation (how closely its ranking matches yours) against your labels. If it does not track your ratings on a small set, its scores on ten thousand examples are noise with a decimal point. This is the same discipline as building a golden set for retrieval.

The rubric is vague. “Rate the quality from 1 to 10” gives you a random number generator with good manners. Every point on the scale needs a concrete description of what earns it, and the more the rubric names specific, checkable properties, the more stable the scores. Ten-point scales are worse than five-point ones here, because the model cannot reliably tell a 6 from a 7, and neither can you.

You are grading against the answer you wanted. If you give the judge a reference answer and ask it to score similarity, you are back to measuring string overlap with extra steps. You will also penalize a correct answer that happens to be phrased differently. For open-ended output, reference-free rubric scoring (grading against the question and the context rather than a gold string) is usually the better call.

The judge and the generator are the same model. This is self-enhancement bias again, but it is common enough to repeat. Sharing one model between generation and judging quietly inflates the score, and the mistake is easy to make because it is the model you already have wired up.

What I would do differently

I trusted a single-answer score for too long before I validated it against my own labels, and the absolute number lulled me into reading drift that was not there. If I were starting over, I would build pairwise comparison first. “Did this change beat the old version?” is the question I actually have most of the time, and pairwise answers it more reliably than watching a 1-to-5 average wobble.

I would also log answer length next to every score from day one. It costs nothing, and it is the fastest way to catch verbosity bias dressing up a regression as a win.

The honest framing is that an LLM judge is a cheap, biased, tireless grader that agrees with a human most of the time. That is genuinely useful, and it is not the same as correct. Use it to catch big regressions and to rank candidates. Keep a small human-labeled set as the thing you actually trust, and re-check the judge against that set whenever you change the judge model.

Where this runs

I leaned on LLM judging to keep Archi, the RAG copilot I worked on for CMS computing operations at CERN, from regressing on the answers themselves, not just the retrieved chunks. Operators ask about specific error codes and procedures. An answer that sounds right but drifts off the logbook it cited is worse than a blank one, and faithfulness scoring is what flags that before it ships.

The retrieval side has its own metrics. This is the other half, and the two together are the difference between knowing a pipeline changed and knowing whether it got better. A generation quality you cannot score is a quality you are shipping on faith.

Image credit: diagrams by M. Hassan Ahmed, created for this post and released under CC0. Example rubric references the Archi project.