HumanEval
HumanEval measures functional correctness for programs synthesized from docstrings.
HumanEval asks a model to synthesize Python programs from docstrings. The evaluator executes candidate code against tests to measure functional correctness.
For pass@k, a problem counts as solved if any of k generated samples passes its unit tests. Record the chosen k and sampling setup with each result.
Two cautions matter. EvalPlus shows that substantially more tests can expose failures missed by the original tests. The official HumanEval repository also warns that executing model-generated code requires a robust security sandbox.
Similar-looking code can differ in whether it actually works.
Complete the function, execute it against tests, and estimate pass@k.
- 1 · promptGive the model a Python programming prompt.
- 2 · sampleGenerate one or more candidate function bodies.
- 3 · executeRun candidates against tests and compute the chosen pass@k statistic.
The official repository warns that its program-execution call is not a security sandbox.
| Who | What they ask | What it works with |
|---|---|---|
| Code-model researcher | “Do sampled functions satisfy the test cases?” | pass@1 or another stated pass@k |
| Evaluation engineer | “Was generated code executed safely?” | Isolation and resource controls |
| Benchmark reader | “Are two reported numbers directly comparable?” | Sampling count, temperature and pass@k |
- Evaluates functional correctness for handwritten Python problems.
- Runs generated code against tests.
- Defines pass@k around whether any of k samples passes the tests.
- Provides a released evaluation harness and dataset.
- Existing benchmark test cases can be limited in quantity and quality.
- HumanEval+ catches wrong code that the original tests did not detect.
- The official evaluator is not a security sandbox for untrusted code.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperEvaluating Large Language Models Trained on Code, Chen and colleagues · read 28 Sept 2026
- repoHumanEval repository, OpenAI · read 28 Sept 2026
- paperEvalPlus Rigorous Evaluation of Neural Code Generation Models, Liu and colleagues · read 28 Sept 2026
- paperMultiPL-E A Scalable and Polyglot Approach to Benchmarking Neural Code Generation, Cassano and colleagues · read 28 Sept 2026
- repoBigCode Evaluation Harness, BigCode · read 28 Sept 2026