evaluation
HellaSwag
A benchmark for evaluating commonsense reasoning by asking models to select the most plausible continuation of a scenario. HellaSwag questions are designed to be easy for humans but challenging for AI models.
In practice
Given 'A woman picks up a guitar. She...' the model must choose the most plausible next action from several options.