Domain experts wrote and validated the questions so answering them requires specialized reasoning rather than simple recall. Models choose among candidate answers, and evaluators report accuracy across the tested subjects. The result measures performance on this constrained expert-question format, not scientific ability in general.