Safety, Ethics & Governance · intermediate · concept 168 of 176
AI Alignment
The challenge of ensuring AI systems pursue goals beneficial to humans. The most important unsolved problem in AI. Includes technical approaches (RLHF, Constitutional AI) and governance frameworks.
Key terms
Value alignmentOuter alignmentInner alignmentReward hacking
Videos
▶ Alignment faking in large language models ↗
Anthropic · YouTube
▶ How difficult is AI alignment? | Anthropic Research Salon ↗
Anthropic · YouTube
Guides and articles
Courses, papers, and more
This unlocks