A protocol specifies tasks, rubrics, blind presentation, sampling, and how disagreements are handled. Rater training and inter-rater agreement help reveal whether the resulting scores are dependable.
Human evaluation asks people to judge model outputs against defined criteria or preferences.
A protocol specifies tasks, rubrics, blind presentation, sampling, and how disagreements are handled. Rater training and inter-rater agreement help reveal whether the resulting scores are dependable.