safety
AI Safety
The research field focused on ensuring AI systems behave as intended and do not cause unintended harm. AI safety encompasses technical research on alignment, robustness, and interpretability, as well as policy and governance frameworks.
In practice
AI safety research includes developing methods to prevent models from producing harmful content or being manipulated by adversarial inputs.