The original SWE-bench drew criticism for issues with ambiguous requirements or broken test harnesses, which understated model performance. SWE-bench Verified keeps only the subset professional software engineers confirmed was solvable and well-specified, and has become the default citation for coding-agent claims. Scaffold, tool permissions, and time limits are set by each vendor's own harness, so scores across labs are not directly comparable without matching those details.