Questions are constructed around relationships between concepts and paired with several plausible answer choices. A model must use ordinary world knowledge to select the best option rather than rely only on surface word overlap. Accuracy should be interpreted alongside prompt format and possible exposure to the public dataset.