Functional AI is coming soon. Join the waitlist for early access.

// blog

7 min read

LLM as Judge: How It Works and When to Use It

LLM as judge uses one model to score another's outputs. Learn how it works, the three judging modes, common biases, and how to wire it into your CI/CD pipeline.

LLM as judge means using one language model to score, classify, or rank the outputs of another. Feed it your agent's response, a rubric, and the original input — and it returns a score plus reasoning. No annotation queue, no waiting for domain experts, no five-figure labeling contract. At production scale, it is the only realistic way to evaluate free-form AI outputs continuously.

What is LLM as judge?

LLM as judge is an evaluation method where a "judge" model receives an input, the AI-generated output, and a scoring rubric, then returns a numeric score or verdict with a natural-language explanation. The MT-Bench paper (Zheng et al., NeurIPS 2023) established the technique at scale: GPT-4 as judge agreed with expert human reviewers approximately 85% of the time — higher than the 81% agreement between two humans evaluating the same output. That number made continuous automated evaluation practical for production systems.

The approach fills a gap that rule-based tests cannot. Traditional metrics like BLEU and ROUGE measure surface overlap; they cannot tell you whether a customer-support response actually resolved the issue, whether a summarization was faithful to the source, or whether a multi-step agent followed a coherent reasoning path.

When does LLM as judge fit your stack?

Use it when the quality dimension is qualitative and context-dependent:

  • Correctness without a ground truth: your agent summarizes a meeting — there is no single right answer, but you know what a bad one looks like.
  • Safety and tone: responses that are technically accurate but inappropriate for your product context.
  • Faithfulness to context: RAG responses that drift from the retrieved documents.
  • Agent trajectory quality: did the agent take sensible intermediate steps, or did it loop and hallucinate tool calls?

Skip it for deterministic outputs. If you can write a passing/failing schema check or regex, do that first — it is faster, cheaper, and not subject to the biases covered below.

The three judging modes

Point-wise scoring — the judge reads one output and rates it on a rubric ("score 1–5 for faithfulness, where 1 means the response contradicts the retrieved context and 5 means every claim is grounded"). This is the default mode for production monitoring: run it on every live trace, track the score distribution over time, alert when the P10 drops below your threshold.

Pairwise comparison — the judge reads two outputs (e.g., response from prompt@v3 vs. prompt@v4) and picks the winner. Use this when deciding whether a prompt change is net-positive across a representative eval set before shipping. It answers the question "is this version better?" rather than "how good is this version?"

Trajectory evaluation — the judge reads the full ordered sequence of an agent's actions: tool calls, intermediate reasoning steps, final output. This is the right scope for multi-step agents where the final answer looks correct but the path was brittle or non-deterministic. Feed the judge the complete trace, not just the last message.

Why LLM judges go wrong

The same NeurIPS paper that validated LLM as judge also catalogued its failure modes — position bias, verbosity bias, and self-enhancement bias — and later research has added more. Each silently corrupts eval scores if left unaddressed.

Position bias: in pairwise mode, the judge tends to favor whichever response appears first in the prompt. Research across 22 tasks and 15 judge models confirms this is not random — it varies by judge and by the quality gap between candidates. Mitigation: run each pairwise comparison twice with candidates swapped, then aggregate. If the verdicts contradict, score it as a tie.

Verbosity bias: longer responses score higher even when the extra length adds no value. Research found that padding a response with a redundant summary paragraph can shift scores by 0.5–1.5 points on a 5-point scale. Mitigation: add explicit rubric language ("do not reward length; reward precision") and test your rubric by scoring a known-bad verbose response against a known-good concise one.

Self-preference bias: if the judge model and the generator share architecture or training lineage, the judge subtly favors the generator's style — not by recognizing its own output, but because it assigns lower perplexity to text that matches its own distribution. Mitigation: use a different model family for judging than for generation.

Style bias: recent systematic research found that style bias — favoring markdown-formatted responses over plain prose — can be stronger than position or verbosity effects in some configurations. Normalize output format before judging when possible, and validate your judge on pairs where format varies but content quality is equivalent.

Writing a reliable judge prompt

A good judge prompt has five components:

Role: You are a strict evaluator assessing AI responses for a B2B SaaS product.
Task: Score the following response for [CRITERION] on a scale from 1 to 5.
Input: {{user_query}}
Response: {{ai_response}}
[Optional] Context: {{retrieved_documents}}
Rubric:
  5 — [specific description of what excellence looks like]
  3 — [specific description of what adequacy looks like]
  1 — [specific description of what failure looks like]
  Do NOT reward length. Do NOT consider formatting unless explicitly asked.
Output: Return JSON: {"score": <int>, "reasoning": "<one sentence>"}

Concrete rubric definitions matter more than anything else. Vague criteria ("helpful", "relevant") let the judge's priors fill in the gaps — which reintroduces annotation subjectivity at the prompt level. Specify exact failure modes.

Always request structured output: a score field and a reasoning field. The reasoning is what your team reviews when a regression fires — it tells you why the score dropped, not just that it dropped. Without it, a failing eval is an alarm with no signal.

Running LLM judge as a CI/CD gate

The same evaluator should run at every stage:

  1. Development — run the judge against a frozen eval set (20–50 cases covering your known failure modes) on every git push. A failing run blocks the PR. This is the cheapest place to catch regressions.
  2. Pre-release — expand to a larger eval set before deploying a new agent version. Set a minimum pass rate and fail the deploy if it's missed.
  3. Production — run the judge on a sample of live traffic. Track score distributions over time. Alert on drift, not just on individual failures.

The critical discipline is keeping the eval set frozen and versioned. When you update the set, you reset your historical baseline — and lose the ability to compare against past deploys.

With Functional AI, a pre-release gate looks like this:

fn eval classify --threshold 0.90 --judge gpt-4o --criteria faithfulness

This evaluates the classify Agentic Function against its versioned eval set, using GPT-4o as judge for the faithfulness criterion, and returns a pass/fail with per-case reasoning. Ship only on green.

Where Functional AI fits

LLM as judge is a technique; you still need somewhere to run it reliably against every version of your agents. Functional AI is an Agentic Function Runtime that evaluates each new version of your Agentic Function before it is promoted — same judge, same rubric, same threshold — so the quality bar does not quietly drift as your team ships faster. Automatic model fallback keeps the judge itself available even when a provider has an outage. And because every Agentic Function version is pinned, you can reproduce the exact inputs and outputs of any past eval run.

Trust your agents in production — and that starts with knowing they passed the same bar every time.

Join the waitlist at fnai.dev to run LLM-as-judge evals as part of your deploy pipeline.

FAQ

What is LLM as judge? LLM as judge is an evaluation method where a language model scores or ranks the outputs of another AI system based on a written rubric, returning a score and reasoning without requiring human annotators.

How accurate is LLM as judge compared to human evaluation? The MT-Bench paper (Zheng et al., NeurIPS 2023) showed that GPT-4 as judge agreed with expert human reviewers approximately 85% of the time (excluding ties), which is higher than the 81% agreement between two humans evaluating the same output.

What biases affect LLM judge results? The main documented biases are position bias (favoring the first response in pairwise comparisons), verbosity bias (preferring longer responses regardless of quality), self-preference bias (favoring outputs from architecturally similar models), and style bias (favoring markdown-formatted text over plain prose).

When should I use LLM as judge instead of rule-based tests? Use LLM as judge for qualitative dimensions that cannot be checked deterministically — faithfulness to source, tone, reasoning quality, or agent trajectory coherence. For outputs with a deterministic ground truth (structured data, code correctness), deterministic tests are faster and cheaper.

How do I prevent an LLM judge from being biased toward longer outputs? Write explicit rubric language ("do not reward length; reward precision") and verify it by scoring a known-bad verbose response against a known-good concise one. If the verbose response wins, tighten the rubric language and re-test.

Can I use the same model as both generator and judge? It is possible but risks self-preference bias — the judge subtly favors outputs that match its own distribution. Using a different model family for generation and evaluation reduces this risk and produces more reliable eval scores.

// private beta

Ship agents you can trust.

Build your prompt, agent, or multi-agent workflow as an Agentic Function. Functional AI hosts it, certifies each version before it ships, switches to a backup model when the primary fails, and lets you release without a redeploy. You just call it from your code.

No spam. One email when Functional AI launches.

Shipping agents today? Become a design partner — design partners shape the roadmap and get early access.