The term surfaced widely in Hacker News and X discussions of model launches where a vendor's self-reported benchmark table showed unusually large jumps on one specific eval while other metrics moved less, or where different harnesses and prompting setups made the comparison hard to reproduce independently — as commenters raised about Qwen3.8-27B's cited SWE-bench Pro score against Claude Opus in August 2026. It's a practical case of Goodhart's Law: once a benchmark becomes the target, it stops reliably measuring the underlying capability.