Skip to content

evaluation

Benchmark

A standardized test or dataset used to measure and compare the performance of AI models on specific tasks. Benchmarks provide objective metrics for tracking progress and comparing different approaches.

In practice

Common LLM benchmarks include MMLU for knowledge, HumanEval for coding, and HellaSwag for commonsense reasoning.