Concepts

SWE-bench

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

SWE-bench asks a system to generate a repository patch for a real GitHub issue.

1 · What it is

SWE-bench turns issue resolution into a repeatable evaluation. The input includes an issue description and a codebase. The output is a patch.

The evaluation applies that patch and runs tests in a containerized environment. This makes the central result concrete: did the patch resolve the benchmark instance under its tests? Reports should name the dataset variant because the full, Lite and Verified sets contain different instances.

The score has a boundary. OpenAI’s Verified report found original unit tests that were overly specific or unrelated to the issue. It also describes cases where environment setup could make tests fail regardless of the solution.

2 · Why it exists

Short code completions do not represent repository-level software work.

ContextProblems are drawn from real issues and corresponding pull requests.
ChangeThe system must edit a codebase rather than return only a short function.
CheckA containerized harness evaluates the generated patch with tests.
3 · How it works

Reproduce the repository, apply a candidate patch, and run the evaluation tests.

SWE-bench measures whether a generated patch passes the benchmark's repository tests.
  1. 1 · prepareLoad the issue description and codebase.
  2. 2 · patchAsk the system to generate edits that address the issue.
  3. 3 · testApply the patch and run the benchmark's tests in its evaluation environment.

SWE-bench Verified is a 500-problem subset that software engineers confirmed as solvable.

4 · Where it's used
WhoWhat they askWhat it works with
Agent developer“Does this system resolve repository-level issues?”Resolved-instance rate and traces
Evaluation operator“Can I reproduce this patch result?”Dataset version, image and test output
Researcher“How does performance change across issue subsets?”Full, Lite and Verified results
5 · What it solves, and what it doesn't
solves
  • Connects issue descriptions to repository-level code changes.
  • Uses real GitHub issues and corresponding pull requests.
  • Provides a Docker-based evaluation harness for reproducible runs.
  • Offers a 500-problem expert-verified subset.
doesn't solve
  • Some original unit tests were overly specific or unrelated to the issue.
  • The full, Lite and Verified datasets contain different numbers of instances.
  • The same model can score differently when used with different scaffolds.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperSWE-bench Can Language Models Resolve Real-World GitHub Issues, Jimenez and colleagues · read 28 Sept 2026
  2. repoSWE-bench repository, SWE-bench · read 28 Sept 2026
  3. docsSWE-bench datasets guide, SWE-bench · read 28 Sept 2026
  4. officialIntroducing SWE-bench Verified, OpenAI · read 28 Sept 2026
  5. paperSWE-agent Agent-Computer Interfaces Enable Automated Software Engineering, Yang and colleagues · read 28 Sept 2026