A judge model is given a rubric and the content to grade, then returns a score or structured verdict; it is used to evaluate chatbot responses, coding agent patches, retrieval quality, and even human-authored documents like academic papers. Strong implementations score multiple named dimensions rather than one number, emit structured qualitative feedback, and disclose what the judge could and could not actually read or verify — weak implementations return a single confident-sounding score with no way to audit it.