evaluation
Benchmark
A standardized test or dataset used to measure and compare the performance of AI models on specific tasks. Benchmarks provide objective metrics for tracking progress and comparing different approaches.
In practice
Common LLM benchmarks include MMLU for knowledge, HumanEval for coding, and HellaSwag for commonsense reasoning.
In the index
Tools that mention Benchmark
- Papers with CodeTrust 87
- SWE-AgentTrust 78
Matched on each tool’s own description and feature list, highest trust score first.
Briefings