Two anonymous models answer the same prompt, and a user selects a winner or tie before model identities are revealed. Aggregated pairwise outcomes are converted into relative ratings with uncertainty and sampling considerations. The population of voters and prompts influences what the ranking represents.