The term comes from systems and design theory, where problems are placed on a spectrum from tame (concise spec, mechanical verification, closed feedback loop — chess, math, passing a compiler) to wicked (no ground truth, subjective success metric, context that keeps shifting). It explains why LLMs improved quickly at math and code, which have verifiable right-and-wrong answers reinforcement learning can train against, while plateauing on prose quality: writing has no objective verifier, and success is measured by resonance inside another human mind rather than a checkable output. Some commentators go further and call writing an AI-complete problem, meaning fully solving it may require general human-level understanding rather than more scale on current architectures.