After each comparison, ratings update according to the expected outcome implied by the current rating difference. In model evaluation, votes or judge decisions supply the match results and ordering can be estimated with uncertainty. Ratings depend on the comparison pool and are not absolute measures of capability.