Skip to content

safety

Interpretability

The degree to which humans can understand and reason about an AI model's internal workings and decision-making process. Interpretability research develops tools for understanding what models learn, how they process information, and why they produce specific outputs.

In practice

Attention visualization shows which parts of the input a model focuses on when making predictions, aiding interpretability.