Safety as a Core Discipline
AI safety research has transformed from a fringe concern discussed by a handful of academics to a well-funded, mainstream research discipline. Major labs collectively employ over 1,000 safety researchers, and the field has produced meaningful technical results. But significant challenges remain.
Alignment Progress
Constitutional AI and RLHF
Anthropic's Constitutional AI (CAI) and the broader RLHF (Reinforcement Learning from Human Feedback) paradigm have proven effective at making models follow instructions and avoid harmful outputs in most contexts. Claude 4's refusal rates on harmful requests exceed 99.5%, while maintaining helpfulness on legitimate queries.
However, these techniques have known limitations. They operate on model outputs rather than internal representations, meaning they don't guarantee the model is "thinking safely", only that it's behaving safely in tested scenarios. The gap between observed behavior and guaranteed alignment remains the central challenge.
Interpretability
Mechanistic interpretability, understanding what happens inside neural networks at the level of individual circuits and features, has made significant progress. Anthropic's research has identified features corresponding to specific concepts (deception, sycophancy, harmful content) and demonstrated that these features can be selectively amplified or suppressed.
Key milestones:
- Sparse autoencoders can now decompose model activations into interpretable features with reasonable fidelity
- Circuit-level analysis has identified how models perform specific reasoning tasks
- Probing classifiers can detect when models are being deceptive or uncertain, even when their outputs don't reveal it
What remains: scaling these techniques to frontier models (most interpretability work has been done on smaller models), and developing tools that make interpretability practical for ongoing model monitoring.
Scalable Oversight
How do you evaluate AI systems that are more capable than their human evaluators? This question has driven research into:
- AI-assisted evaluation: Using one AI to evaluate another, with humans overseeing the evaluation process
- Debate: Having two AI systems argue opposing positions while a human judge evaluates
- Recursive reward modeling: Decomposing complex tasks into simpler subtasks that humans can evaluate
- Constitutional AI: Having models evaluate their own outputs against written principles
These approaches work reasonably well for current models but face theoretical challenges as model capabilities increase.
Red Teaming and Adversarial Testing
Red teaming, systematically trying to make models behave badly, has become standard practice. All major labs conduct extensive red teaming before model releases, and a growing ecosystem of independent red teams provides external evaluation.
Progress: Structured red teaming has identified and fixed thousands of failure modes before they reached users. Automated red teaming using AI systems to probe other AI systems has dramatically increased testing coverage.
Challenges: Red teaming is inherently reactive. It finds problems that testers think to look for. Novel failure modes that no one anticipated can still emerge in production. The adversarial arms race between model defenders and jailbreak researchers continues.
Existential Risk Research
Research on longer-term AI risks has matured beyond philosophical speculation into technical research programs:
Power-seeking behavior: Theoretical work has formalized conditions under which AI systems might pursue resources or influence beyond their intended scope. Empirical tests on current models show limited power-seeking tendencies, but researchers note that current models lack the capability for meaningful power-seeking.
Deceptive alignment: Could an AI system appear aligned during testing but pursue different objectives in deployment? Research has demonstrated this phenomenon in toy settings but hasn't observed it in frontier models. The concern is that it becomes more likely as models become more capable.
Biosecurity and cybersecurity risks: Evaluations of frontier models' ability to assist with biological weapon development or sophisticated cyberattacks have shown incremental but concerning capability increases. All major labs now conduct these evaluations before releases.
Governance and Policy
AI safety governance has evolved rapidly:
- Frontier Model Forum members (Anthropic, Google DeepMind, Microsoft, OpenAI) commit to safety testing and information sharing
- NIST AI Safety Institute conducts pre-release evaluations of frontier models
- EU AI Act mandates risk assessment and conformity procedures for high-risk AI systems
- International coordination through the AI Safety Summits has established shared frameworks (though enforcement remains weak)
Open Problems
Despite progress, fundamental challenges remain unsolved:
- Specification problem: How do you precisely define "aligned with human values" when humans disagree about values?
- Distribution shift: Models trained to be safe in one context may behave unpredictably in novel contexts
- Emergent capabilities: New capabilities that appear unexpectedly as models scale make it difficult to anticipate risks
- Coordination problem: Safety standards are only as strong as the least cautious lab
- Economic incentives: Competitive pressure encourages faster deployment, potentially at the expense of thorough safety evaluation
The Bottom Line
AI safety has made genuine technical progress, and the field is taken seriously by major labs and governments. However, the pace of capability development continues to outstrip our ability to guarantee safe behavior. The goal is to close this gap before it becomes critical, a race that the safety community is running but hasn't yet won.
Covers