Skip to content

evaluation

HumanEval

A benchmark for evaluating AI code generation consisting of 164 Python programming problems with unit tests. HumanEval measures a model's ability to generate functionally correct code from docstrings.

In practice

Claude Opus 4 scores above 80% on HumanEval, meaning it correctly solves most programming challenges on the first attempt.