Reinforcement learning from verifiable rewards optimizes a policy against an automatically gradable signal — a unit test passing, a proof checker accepting, an exact-match answer — rather than a learned preference model as in RLHF. It scales cleanly for math, code, and other checkable domains, but a task with no checkable answer, such as knowing when to pause and ask a clarifying question, gets no reward signal at all. Community discussion of Claude Opus 5's reduced tendency to ask clarifying questions has pointed to heavy RLVR training as the likely cause, since committing to an answer scores higher than pausing to ask under a purely verifiable reward.