Skip to content

safety

AI Alignment

The challenge of ensuring AI systems' goals and behaviors are aligned with human values and intentions. Alignment research aims to build AI that reliably does what its users want while avoiding unintended negative consequences.

In practice

Anthropic's work on Constitutional AI is an alignment technique that trains models to follow explicit ethical principles.