The evaluator receives a rubric and may compare outputs with references or with each other. Calibration against human judgments, order controls, and bias checks are needed because the judge can share model errors or favor superficial traits.
LLM as a judge uses a language model to score, rank, or critique other model outputs.
The evaluator receives a rubric and may compare outputs with references or with each other. Calibration against human judgments, order controls, and bias checks are needed because the judge can share model errors or favor superficial traits.