An agent receives an issue and repository snapshot, produces a patch, and is scored by tests that check the intended behavior and regressions. Environment setup, tool access, and issue selection materially affect results.
SWE-bench evaluates systems by asking them to resolve software issues in real code repositories.
An agent receives an issue and repository snapshot, produces a patch, and is scored by tests that check the intended behavior and regressions. Environment setup, tool access, and issue selection materially affect results.