SWE-bench
SWE-bench asks a system to generate a repository patch for a real GitHub issue.
SWE-bench turns issue resolution into a repeatable evaluation. The input includes an issue description and a codebase. The output is a patch.
The evaluation applies that patch and runs tests in a containerized environment. This makes the central result concrete: did the patch resolve the benchmark instance under its tests? Reports should name the dataset variant because the full, Lite and Verified sets contain different instances.
The score has a boundary. OpenAI’s Verified report found original unit tests that were overly specific or unrelated to the issue. It also describes cases where environment setup could make tests fail regardless of the solution.
Short code completions do not represent repository-level software work.
Reproduce the repository, apply a candidate patch, and run the evaluation tests.
- 1 · prepareLoad the issue description and codebase.
- 2 · patchAsk the system to generate edits that address the issue.
- 3 · testApply the patch and run the benchmark's tests in its evaluation environment.
SWE-bench Verified is a 500-problem subset that software engineers confirmed as solvable.
| Who | What they ask | What it works with |
|---|---|---|
| Agent developer | “Does this system resolve repository-level issues?” | Resolved-instance rate and traces |
| Evaluation operator | “Can I reproduce this patch result?” | Dataset version, image and test output |
| Researcher | “How does performance change across issue subsets?” | Full, Lite and Verified results |
- Connects issue descriptions to repository-level code changes.
- Uses real GitHub issues and corresponding pull requests.
- Provides a Docker-based evaluation harness for reproducible runs.
- Offers a 500-problem expert-verified subset.
- Some original unit tests were overly specific or unrelated to the issue.
- The full, Lite and Verified datasets contain different numbers of instances.
- The same model can score differently when used with different scaffolds.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperSWE-bench Can Language Models Resolve Real-World GitHub Issues, Jimenez and colleagues · read 28 Sept 2026
- repoSWE-bench repository, SWE-bench · read 28 Sept 2026
- docsSWE-bench datasets guide, SWE-bench · read 28 Sept 2026
- officialIntroducing SWE-bench Verified, OpenAI · read 28 Sept 2026
- paperSWE-agent Agent-Computer Interfaces Enable Automated Software Engineering, Yang and colleagues · read 28 Sept 2026