explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

  1. Home
  2. /
  3. Dictionary
  4. /
  5. LLM-as-Judge
Evaluation & Benchmarksaka AI refereeaka LLM judge

LLM-as-Judge

LLM-as-judge is the practice of using a language model call to score another system output or document against defined criteria, instead of relying only on exact-match rules or human review.

Ask Melo about this← all terms

A judge model is given a rubric and the content to grade, then returns a score or structured verdict; it is used to evaluate chatbot responses, coding agent patches, retrieval quality, and even human-authored documents like academic papers. Strong implementations score multiple named dimensions rather than one number, emit structured qualitative feedback, and disclose what the judge could and could not actually read or verify — weak implementations return a single confident-sounding score with no way to audit it.

Related terms

AI BenchmarkBenchmark ContaminationDouble-Blind EvaluationArtificial Analysis Intelligence IndexRecallPass at K