evaluation
Eval
Short for evaluation, the systematic process of measuring AI model performance across various dimensions including accuracy, safety, speed, and user satisfaction. Evals are critical for model development, comparison, and deployment decisions.
In practice
Before deploying a new model version, teams run comprehensive evals across safety, helpfulness, and task-specific benchmarks.