Contributors defined tasks, examples, and metrics spanning reasoning, language, knowledge, and other behaviors. Results are normally examined by task or task group because one aggregate can hide large differences. Prompt format and model settings remain part of any comparison.