Questions span many academic disciplines and may include diagrams, charts, maps, or other images. A model combines visual perception with subject knowledge and reasoning to select or generate an answer. Scores reflect the benchmark's question distribution and input formatting rather than all multimodal ability.