safety
Constitutional AI
An approach developed by Anthropic where AI models are trained to follow a set of explicit principles (a constitution) that guide their behavior. The model critiques and revises its own outputs based on these principles during training.
In practice
Constitutional AI trains a model to evaluate its own responses against principles like 'be helpful but avoid harm' and revise accordingly.