•11 min
Your LLM Judge Prefers the Longer Answer
You wired an LLM up as the grader for your evals, and now every release looks green. The problem is the judge is not scoring quality, it is scoring length, order, and answers that sound like its own. Fix the three biases that matter and calibrate the judge against human labels, so the number your pipeline gates on actually means what you think it means.
Evals
LLM-as-judge