Concepts

HumanEval

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

HumanEval measures functional correctness for programs synthesized from docstrings.

1 · What it is

HumanEval asks a model to synthesize Python programs from docstrings. The evaluator executes candidate code against tests to measure functional correctness.

For pass@k, a problem counts as solved if any of k generated samples passes its unit tests. Record the chosen k and sampling setup with each result.

Two cautions matter. EvalPlus shows that substantially more tests can expose failures missed by the original tests. The official HumanEval repository also warns that executing model-generated code requires a robust security sandbox.

2 · Why it exists

Similar-looking code can differ in whether it actually works.

ExecuteHumanEval evaluates functional correctness by running generated code.
SamplePass@k counts a problem as solved when any of k samples passes the tests.
CompareThe released harness provides a common dataset and evaluator.
3 · How it works

Complete the function, execute it against tests, and estimate pass@k.

HumanEval checks short Python function completions with tests; executing untrusted generations requires a sandbox.
  1. 1 · promptGive the model a Python programming prompt.
  2. 2 · sampleGenerate one or more candidate function bodies.
  3. 3 · executeRun candidates against tests and compute the chosen pass@k statistic.

The official repository warns that its program-execution call is not a security sandbox.

4 · Where it's used
WhoWhat they askWhat it works with
Code-model researcher“Do sampled functions satisfy the test cases?”pass@1 or another stated pass@k
Evaluation engineer“Was generated code executed safely?”Isolation and resource controls
Benchmark reader“Are two reported numbers directly comparable?”Sampling count, temperature and pass@k
5 · What it solves, and what it doesn't
solves
  • Evaluates functional correctness for handwritten Python problems.
  • Runs generated code against tests.
  • Defines pass@k around whether any of k samples passes the tests.
  • Provides a released evaluation harness and dataset.
doesn't solve
  • Existing benchmark test cases can be limited in quantity and quality.
  • HumanEval+ catches wrong code that the original tests did not detect.
  • The official evaluator is not a security sandbox for untrusted code.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperEvaluating Large Language Models Trained on Code, Chen and colleagues · read 28 Sept 2026
  2. repoHumanEval repository, OpenAI · read 28 Sept 2026
  3. paperEvalPlus Rigorous Evaluation of Neural Code Generation Models, Liu and colleagues · read 28 Sept 2026
  4. paperMultiPL-E A Scalable and Polyglot Approach to Benchmarking Neural Code Generation, Cassano and colleagues · read 28 Sept 2026
  5. repoBigCode Evaluation Harness, BigCode · read 28 Sept 2026