The subset emphasizes tasks on which earlier language models performed poorly relative to human baselines. Evaluators commonly compare direct prompting with prompts that encourage intermediate reasoning. Scores should state the exact prompts, task set, and answer parsing used.