LLM-as-a-judge: have a model rate a model
LLM-as-a-judge means having one model score a system's responses. The foundational study (Zheng et al., MT-Bench and Chatbot Arena, 2023) shows that a strong judge such as GPT-4 reaches over 80% agreement with human preferences, comparable to agreement between two humans—with documented biases involving position, verbosity, self-preference, and limited reasoning. Locally, a 13- to 14-billion-parameter judge using pairwise comparison produces usable results.
Having one model score another model's responses has become the most widespread evaluation method for a simple reason: it's the only one that scales. It is also biased, manipulable, and misleading if read as an absolute score. The study that established the reference framework, published by the LMSYS team (Chatbot Arena) and presented at NeurIPS 2023, says so itself: used properly, it compares two versions of a system; used improperly, it produces a reassuring number that measures nothing.
#The principle and why it matters
You give a model a question, the response produced by the system being evaluated, optionally the passages used to answer it and a reference answer, plus a scoring rubric. It returns a score and a justification. Repeated over one hundred cases, this produces a metric you can track over time, whenever you change the prompt, model, or retrieval parameter.
The alternative is human evaluation, which remains the standard but does not scale: no one will reread one hundred responses every time the prompt changes. An automated judge lets you answer “did this change improve or degrade the system?” in ten minutes, changing the way you work. Zheng et al.’s foundational paper was motivated precisely by this observation: existing benchmarks (MMLU and the like) do not measure what matters in an open-ended conversation, and large-scale human evaluation is too expensive to iterate quickly on a system that evolves every week.
#Write a grid that fits
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
- 01One criterion at a timeAsking for an overall score is meaningless. Score source fidelity, relevance, and completeness separately.
- 02A short, documented scaleThree or four levels, with what each one means. A ten-point scale produces scores that cannot be reproduced.
- 03Require justification before the scoreA judge asked to reason and then conclude is much more stable than one that simply produces a number. It follows the same logic as the “limited reasoning ability” identified by Zheng et al.: forcing explicit reasoning reduces this limitation.
- 04Prefer pairwise comparison“Which of these two answers is better?” is a question a model answers much better than “rate this one out of five”—that's the format Chatbot Arena uses to compare models.
#Biases and how to contain them
The reference article on the subject identifies exactly three biases and one limitation: position bias, verbosity bias, self-preference bias (self-enhancement), and limited reasoning ability. This is not a blog hypothesis; it is the central result of the study that introduced MT-Bench and Chatbot Arena, two references still used today to rank models.
| Bias | Effect | Countermeasure |
|---|---|---|
| Position | The answer presented first is favored | Alternate the order and average the two passes |
| Verbosity | A longer answer is judged better, all else being equal | Specify it in the table, and check the length of both sides |
| Self-preference | A judge favors responses from its own model family | Do not evaluate a model against itself or a close cousin |
| Limited reasoning | The judge gets calculations or deduction tasks wrong | Require a written justification before the score, never just a score |
| Sycophancy | Everything is well documented; the ratings are stacked at the top of the scale | A demanding rubric and examples of bad answers in the prompt |
The good news, documented by the same study: a strong judge such as GPT-4 achieves more than 80% agreement with human preferences on MT-Bench and Chatbot Arena—the same level of agreement measured between two human evaluators. The rule of thumb that sums up this entire chapter: a judge is used for comparison, never to certify quality in absolute terms. A score of 4.2 out of 5 means nothing; a change from 3.1 to 3.8 on the same test set with the same judge indicates something real.
The verbosity bias is not a theoretical concern: the AlpacaEval benchmark, widely used to compare instruction-tuned models, is explicitly known to favor models that generate longer responses when quality is equal. Its length-controlled version raises the correlation with Chatbot Arena rankings from 0.94 to 0.98 — quantitative proof that neutralizing this single bias mechanically brings automated judgment closer to human judgment rather than farther away.
#Is a local judge serious?
Yes, under two conditions. The first is size: below roughly 14 billion parameters, judgments become unstable and the output format breaks, making the evaluation unusable—consistent with the site’s memory guideline (14B ≈ 9 GB in Q4). The second is context: the judge must contain the question, passages, and answer all at once—a saturated context causes it to score a truncated text, and no one notices because the model continues responding normally based on what it actually received.
The value of a local judge is obvious when the data being evaluated is confidential: sending every response and every passage to an external API for scoring defeats the purpose of keeping the system local, especially in a medical, legal, or financial use case where the source passages themselves contain sensitive data. And unlike production, an evaluation campaign tolerates slowness very well: launch it overnight on a GPU that serves another purpose during the day, and collect the results in the morning.
#Tools to avoid reinventing the pipeline
You do not need to write your own judge harness from scratch. Prometheus, an open-source research project, specifically trains a 13-billion-parameter model for this purpose: “We train Prometheus, a 13B evaluator LLM that can assess any given long-form text based on customized score rubric provided by the user” (we train Prometheus, a 13-billion-parameter evaluator LLM capable of evaluating any long-form text according to a customized scoring rubric provided by the user). The authors report a Pearson correlation of 0.897 with human evaluators across 45 customized rubrics, compared with 0.882 for GPT-4 and only 0.392 for ChatGPT under the same protocol—an gap that illustrates how badly a nonspecialized model can judge, even when it otherwise generates correct conversational responses.
- Prometheus
- 13B, open, specifically designed for fine-grained evaluation on a custom rubric; a solid foundation for a dedicated local judge rather than a general-purpose model repurposed from its primary use.
- A general-purpose model with 14B parameters or more
- Qwen, Llama, or Mistral in this size range also work as judges, with a carefully written rubric rather than training dedicated specifically to this task.
- An orchestration framework
- Useful for launching the campaign, storing results, and tracking score changes over time instead of recoding everything each time a test needs to be rerun on a new version.
#What does a judge prompt that holds up look like?
Let's revisit the four rules from the previous section with a concrete example: an internal RAG system that answers procedural questions. The judge prompt must explicitly include the question asked, the retrieved source passages, the answer being evaluated, a rubric with separate criteria, and an instruction to provide justification before the score. It takes longer to write than a simple “rate this answer out of ten,” but that detail is what separates a usable evaluation campaign from a reassuring number that measures nothing.
Two details make all the difference in this scaffold. First, the source passages are sent to the judge, not just the answer: without them, it is impossible to verify faithfulness, only form. Second, the justification comes before the conclusion in the prompt order—a judge asked for the score first tends to justify it afterward instead of reasoning before deciding, which worsens the limited reasoning ability identified by the reference study. Finally, alternating the order of answers A and B from one case to the next, then averaging the two passes, neutralizes most of the position bias documented by Zheng et al.
#When not to use it
- For validating a high-stakes system
- Medical, legal, financial: an automated judge doesn't replace human review for cases involving liability.
- To compare two very different systems
- Style and length biases dominate as soon as response formats differ, as shown by the verbosity bias documented by Zheng et al.
- When a deterministic test exists
- If the right answer is a number or an exact value, a simple comparison is more accurate and infinitely cheaper than an LLM judge.
- On too few cases
- Across ten questions, generation variance exceeds the difference you are trying to measure.
- Ragas: evaluate your local RAG with numbers
- Langfuse: observe your local LLMs (traces, prompts, evals)
- Read model rankings correctly
- Local RAG: introduction (to frame what we're evaluating)
- Source: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (Zheng et al., 2023)
- Source: Prometheus, a 13B open-source judge (Kim et al., 2023)
- Source: Prometheus official GitHub repository
#FAQ
Can one model really judge another?+
What model size for the judge?+
Do you need a judge different from the evaluated model?+
Single rating or pairwise comparison?+
How many cases are needed for a usable campaign?+
Is Prometheus better than a general-purpose model as a judge?+
Is verbosity bias really measurable?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.