Skip to content

safety

Constitutional AI

An approach developed by Anthropic where AI models are trained to follow a set of explicit principles (a constitution) that guide their behavior. The model critiques and revises its own outputs based on these principles during training.

In practice

Constitutional AI trains a model to evaluate its own responses against principles like 'be helpful but avoid harm' and revise accordingly.