Released in January 2025, HLE is intentionally adversarial: questions were solicited from subject-matter experts specifically to resist models that had memorized older benchmarks like MMLU. Frontier scores moved quickly after launch, and model cards increasingly report separate no-tools and with-tools figures, since web and code access change the result more than raw model capability does.