Advanced 11 minOptimization

LLM-as-a-judge: have a model rate a model

Direct response

LLM-as-a-judge means having one model score a system's responses. The foundational study (Zheng et al., MT-Bench and Chatbot Arena, 2023) shows that a strong judge such as GPT-4 reaches over 80% agreement with human preferences, comparable to agreement between two humans—with documented biases involving position, verbosity, self-preference, and limited reasoning. Locally, a 13- to 14-billion-parameter judge using pairwise comparison produces usable results.

Having one model score another model's responses has become the most widespread evaluation method for a simple reason: it's the only one that scales. It is also biased, manipulable, and misleading if read as an absolute score. The study that established the reference framework, published by the LMSYS team (Chatbot Arena) and presented at NeurIPS 2023, says so itself: used properly, it compares two versions of a system; used improperly, it produces a reassuring number that measures nothing.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#The principle and why it matters

You give a model a question, the response produced by the system being evaluated, optionally the passages used to answer it and a reference answer, plus a scoring rubric. It returns a score and a justification. Repeated over one hundred cases, this produces a metric you can track over time, whenever you change the prompt, model, or retrieval parameter.

The alternative is human evaluation, which remains the standard but does not scale: no one will reread one hundred responses every time the prompt changes. An automated judge lets you answer “did this change improve or degrade the system?” in ten minutes, changing the way you work. Zheng et al.’s foundational paper was motivated precisely by this observation: existing benchmarks (MMLU and the like) do not measure what matters in an open-ended conversation, and large-scale human evaluation is too expensive to iterate quickly on a system that evolves every week.

#Write a grid that fits

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
  1. 01
    One criterion at a time
    Asking for an overall score is meaningless. Score source fidelity, relevance, and completeness separately.
  2. 02
    A short, documented scale
    Three or four levels, with what each one means. A ten-point scale produces scores that cannot be reproduced.
  3. 03
    Require justification before the score
    A judge asked to reason and then conclude is much more stable than one that simply produces a number. It follows the same logic as the “limited reasoning ability” identified by Zheng et al.: forcing explicit reasoning reduces this limitation.
  4. 04
    Prefer pairwise comparison
    “Which of these two answers is better?” is a question a model answers much better than “rate this one out of five”—that's the format Chatbot Arena uses to compare models.
i
Pairwise comparison is the most cost-effective trick
It eliminates the question of scale, greatly reduces variance, and directly answers what you care about: is the new version better than the old one? Just remember to alternate the order of the two answers, because judges have a documented position bias.

#Biases and how to contain them

The reference article on the subject identifies exactly three biases and one limitation: position bias, verbosity bias, self-preference bias (self-enhancement), and limited reasoning ability. This is not a blog hypothesis; it is the central result of the study that introduced MT-Bench and Chatbot Arena, two references still used today to rank models.

What skews an automated evaluation, according to Zheng et al. (2023)
BiasEffectCountermeasure
PositionThe answer presented first is favoredAlternate the order and average the two passes
VerbosityA longer answer is judged better, all else being equalSpecify it in the table, and check the length of both sides
Self-preferenceA judge favors responses from its own model familyDo not evaluate a model against itself or a close cousin
Limited reasoningThe judge gets calculations or deduction tasks wrongRequire a written justification before the score, never just a score
SycophancyEverything is well documented; the ratings are stacked at the top of the scaleA demanding rubric and examples of bad answers in the prompt

The good news, documented by the same study: a strong judge such as GPT-4 achieves more than 80% agreement with human preferences on MT-Bench and Chatbot Arena—the same level of agreement measured between two human evaluators. The rule of thumb that sums up this entire chapter: a judge is used for comparison, never to certify quality in absolute terms. A score of 4.2 out of 5 means nothing; a change from 3.1 to 3.8 on the same test set with the same judge indicates something real.

The verbosity bias is not a theoretical concern: the AlpacaEval benchmark, widely used to compare instruction-tuned models, is explicitly known to favor models that generate longer responses when quality is equal. Its length-controlled version raises the correlation with Chatbot Arena rankings from 0.94 to 0.98 — quantitative proof that neutralizing this single bias mechanically brings automated judgment closer to human judgment rather than farther away.

#Is a local judge serious?

Yes, under two conditions. The first is size: below roughly 14 billion parameters, judgments become unstable and the output format breaks, making the evaluation unusable—consistent with the site’s memory guideline (14B ≈ 9 GB in Q4). The second is context: the judge must contain the question, passages, and answer all at once—a saturated context causes it to score a truncated text, and no one notices because the model continues responding normally based on what it actually received.

The value of a local judge is obvious when the data being evaluated is confidential: sending every response and every passage to an external API for scoring defeats the purpose of keeping the system local, especially in a medical, legal, or financial use case where the source passages themselves contain sensitive data. And unlike production, an evaluation campaign tolerates slowness very well: launch it overnight on a GPU that serves another purpose during the day, and collect the results in the morning.

#Tools to avoid reinventing the pipeline

You do not need to write your own judge harness from scratch. Prometheus, an open-source research project, specifically trains a 13-billion-parameter model for this purpose: “We train Prometheus, a 13B evaluator LLM that can assess any given long-form text based on customized score rubric provided by the user” (we train Prometheus, a 13-billion-parameter evaluator LLM capable of evaluating any long-form text according to a customized scoring rubric provided by the user). The authors report a Pearson correlation of 0.897 with human evaluators across 45 customized rubrics, compared with 0.882 for GPT-4 and only 0.392 for ChatGPT under the same protocol—an gap that illustrates how badly a nonspecialized model can judge, even when it otherwise generates correct conversational responses.

Prometheus
13B, open, specifically designed for fine-grained evaluation on a custom rubric; a solid foundation for a dedicated local judge rather than a general-purpose model repurposed from its primary use.
A general-purpose model with 14B parameters or more
Qwen, Llama, or Mistral in this size range also work as judges, with a carefully written rubric rather than training dedicated specifically to this task.
An orchestration framework
Useful for launching the campaign, storing results, and tracking score changes over time instead of recoding everything each time a test needs to be rerun on a new version.

#What does a judge prompt that holds up look like?

Let's revisit the four rules from the previous section with a concrete example: an internal RAG system that answers procedural questions. The judge prompt must explicitly include the question asked, the retrieved source passages, the answer being evaluated, a rubric with separate criteria, and an instruction to provide justification before the score. It takes longer to write than a simple “rate this answer out of ten,” but that detail is what separates a usable evaluation campaign from a reassuring number that measures nothing.

Judge prompt skeleton (pairwise comparison)
Question de l'utilisateur : {question}
Passages source fournis au système : {passages}

Réponse A : {reponse_a}
Réponse B : {reponse_b}

Pour chaque réponse, évalue séparément :
1. Fidélité aux passages fournis (aucune affirmation absente des sources)
2. Pertinence par rapport à la question posée
3. Complétude (la question est traitée entièrement)

Justifie d'abord ton analyse pour chaque critère, puis conclus par :
Meilleure réponse : A ou B
Raison en une phrase.

Two details make all the difference in this scaffold. First, the source passages are sent to the judge, not just the answer: without them, it is impossible to verify faithfulness, only form. Second, the justification comes before the conclusion in the prompt order—a judge asked for the score first tends to justify it afterward instead of reasoning before deciding, which worsens the limited reasoning ability identified by the reference study. Finally, alternating the order of answers A and B from one case to the next, then averaging the two passes, neutralizes most of the position bias documented by Zheng et al.

#When not to use it

For validating a high-stakes system
Medical, legal, financial: an automated judge doesn't replace human review for cases involving liability.
To compare two very different systems
Style and length biases dominate as soon as response formats differ, as shown by the verbosity bias documented by Zheng et al.
When a deterministic test exists
If the right answer is a number or an exact value, a simple comparison is more accurate and infinitely cheaper than an LLM judge.
On too few cases
Across ten questions, generation variance exceeds the difference you are trying to measure.

#FAQ

Can one model really judge another?+
Enough to compare two versions of the same system on a fixed test set: the reference study measures over 80% agreement between a strong judge and human preferences, comparable to agreement between two humans. Not enough to issue an absolute score or independently validate a high-stakes use case.
What model size for the judge?+
At least 13 to 14 billion parameters in practice, and more if the answers are long: that is the size of the Prometheus model designed specifically for this role. Below that, judgments are unstable and the expected output format (score, justification) frequently breaks.
Do you need a judge different from the evaluated model?+
Yes. The study by Zheng et al. documents a self-preference bias (self-enhancement): a model tends to rate responses from its own family more highly, including a slightly different version of itself. Using the same model on both sides produces a flattering, useless figure for making any decisions, especially before a production version change.
Single rating or pairwise comparison?+
Pairwise comparison, almost always: less variance, no scaling problem, and it directly answers the question “is it better than before.” This is the format Chatbot Arena uses to rank models. Remember to alternate the presentation order to counter position bias.
How many cases are needed for a usable campaign?+
Thirty to fifty real cases are a reasonable foundation for starting to distinguish a genuine difference from generation noise, provided they cover the actual diversity of questions asked of the system. Below around ten questions, generation variability far exceeds the differences you are trying to measure between two versions of the same system.
Is Prometheus better than a general-purpose model as a judge?+
On the custom rubrics tested in the article, Prometheus (13B) reaches a 0.897 correlation with humans, versus 0.882 for GPT-4 and 0.392 for ChatGPT. That's a strong signal, but it was obtained on its own evaluation protocol: validating it on your rubric is still necessary before trusting it in production.
Is verbosity bias really measurable?+
Yes: the AlpacaEval benchmark, known for favoring long answers, measured the effect precisely by correcting for this bias. Its length-neutralized version raises the correlation with Chatbot Arena’s human rankings from 0.94 to 0.98, concretely quantifying the gain from eliminating this single bias in automated evaluation.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.