safety
AI Alignment
The challenge of ensuring AI systems' goals and behaviors are aligned with human values and intentions. Alignment research aims to build AI that reliably does what its users want while avoiding unintended negative consequences.
In practice
Anthropic's work on Constitutional AI is an alignment technique that trains models to follow explicit ethical principles.