Skip to content

evaluation

Leaderboard

A ranking system that compares AI model performance on standardized benchmarks. Leaderboards like the Open LLM Leaderboard track how different models score across multiple evaluation metrics.

In practice

The Chatbot Arena Leaderboard uses human votes to rank models like Claude, GPT-4, and Gemini based on response quality.