AI 101 — What is LLM-as-a-judge?
If you read a week of AI news, you'll notice almost nobody grades models by hand any more. When a lab says its new model is better at writing, summarising or handling customers, the number usually came from another AI.
An LLM judge is a language model used to grade the work of another model — or of your own AI feature — when no piece of code can decide whether the answer was good. The industry term is LLM-as-a-judge, and the job is simple to state: read a question and an answer, then return a score or a ranking instead of a paragraph.
It exists because grading stopped being automatic. As we explained in What is an AI benchmark?, code and maths have right answers a script can check. The moment you ask whether a summary was useful, or whether a support reply kept the tone the brand wanted, there is no test to run. Until recently that meant a human read the output — accurate, slow, expensive, and impossible to run on thousands of examples a night. A judge model is the compromise the industry settled on: a fast, cheap, consistent reader that mostly agrees with a careful human and can be pointed at a million samples.
Why it matters right now
Nearly every model ranking you read is judge-mediated. Arena-style leaderboards, launch-day tables and the eval reports inside enterprise tooling all lean on a model as the scorer. Our own coverage has spent September on results where the grading is done by models or by the teams themselves — the CyberGym leaderboard where Alibaba's 27B security model took the top model-track slot is self-reported, meaning every row was submitted by the team that ran it.
The past week also produced a research wave on what that costs you. One paper, DIAL, collected more than 410,000 judgments from 21 different judge models in both display orders, specifically to separate a judge's position habit from genuine preference. A second argues that running more comparisons cannot remove documented biases such as position, verbosity, judge severity and self-enhancement — you have to model them. A third, CARGO, built a test suite where the right answer is known by construction and found a standard reference-based judge penalising 100% of correct answers once the entities involved were swapped, with its ability to tell good work from bad sitting near zero on the same set.
The tooling moved the same way. AWS's CloudWatch Omni went generally available on September 23 with an evaluation workflow for agent traces, plus third-party evaluators — Braintrust, DeepEval and Ragas — plugging into it. Judges are no longer a research technique; they are a feature in the observability stack.

The mental model
Judges come in two shapes and do two jobs.
The shapes: pairwise judging shows the model two answers and asks which is better — that is how arena-style rankings are built. Rubric judging shows one answer against a written checklist and asks for a score. Pairwise gets cleaner human agreement; rubric gets you a number you can track release over release.
The jobs: an instrument (measuring a system someone else built) and a training signal (grading candidate answers during post-training, standing in for the human labellers). The training role is the one people forget — What is RLHF? describes the reward model that scores answers so a model can be optimised against human preference, and an LLM judge is often what that scorer is made of.
How judges get validated is the part worth remembering: agreement with humans. The 2023 paper that established the practice, by Lianmin Zheng and colleagues at UC Berkeley, found that a strong judge model matched controlled and crowdsourced human preferences with over 80% agreement — the same level of agreement humans show with each other. That number is why the idea spread.
The hiring analogy
Picture 400 applications for one job and no time to read them. So you hire a recruiter to shortlist — fast, consistent, and usually right.
Now watch how the recruiter behaves. Longer CVs get the benefit of the doubt, so padding pays. Whoever was read first feels stronger, because the shortlist forms in order. And the recruiter quietly favours candidates who resemble the last great hire they made. None of that is dishonesty; it is a taste profile with a shape you can measure. Judges behave the same way: they reward length, they are swayed by position, and a model tends to be kinder to output from its own family.
The rubric matters just as much. Hand a recruiter a vague job ad and they optimise for whatever is easy to see — long CVs, confident tone, tidy formatting. Hand a judge a vague instruction and it rewards tidy Markdown and confident phrasing instead of the thing you actually wanted.
Common misconceptions
"Judges only measure; they don't train anything." Judge output is also training data. When a lab says it trained on AI feedback, a judge model is usually the thing doing the feedback.
"A judge is objective where humans aren't." The biases are measured, not hypothetical. The 2023 paper named position, verbosity and self-enhancement; a later framework catalogued a dozen. You can partly cancel them — randomise order, strip length cues, never let a model family judge itself — but knowing the shape of a bias isn't removing it.
"A judge score is comparable across papers." It isn't. The judge model, the rubric, the prompt, the display order and the effort setting all move the number. A leaderboard graded by the team that entered it is one lab's judge on one day, which is why the honest tables publish their evaluation traces alongside the score.
"Once a judge agrees with humans, it's calibrated for good." Agreement is measured against a particular generation of models. When the models being judged change, the target moves and the judge has to be re-checked.
"You need the biggest model to be the judge." Cost drives the choice, and judges are not interchangeable — the DIAL paper needed 21 of them to study the spread. Pick your judge the way you'd pick a human rater: with a test set you graded yourself, not by brand.
"Judge scores mean the feature works for users." A judge measures the rubric you wrote, and rubrics are written by teams. It will happily confirm that your outputs match your assumptions while users quietly disagree.
Where to learn more
The fastest way to understand judging is to read one. The judge prompts from the original MT-Bench work are public, and the prompt is the whole artifact: a question, an answer, a checklist, and a request for a verdict. For testing your own judge, CARGO's approach is the best template — build cases where the correct answer is known in advance, then check whether the judge can tell them apart from wrong ones. And if you're about to point a judge at your own product, start by hand-grading 30 examples yourself — our guide to building a small eval set for your own AI feature makes that the first step, because a judge you have not calibrated is just a second opinion.
Would you trust a model to grade a model — or does every decision that matters still need a human in the loop? Tell us in the comments.
Sources: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv) · DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration (arXiv) · Accounting for Bias Enables Sustainable LLM Evaluation (arXiv) · CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production (arXiv) · Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge (arXiv) · MT-bench judge prompts (GitHub)