Encyclopedia · 187 concepts

Safety, Ethics & Governance · intermediate · concept 177 of 187

AI Alignment

The challenge of ensuring AI systems pursue goals beneficial to humans. The most important unsolved problem in AI. Includes technical approaches (RLHF, Constitutional AI) and governance frameworks.

Key terms

Value alignmentOuter alignmentInner alignmentReward hacking

Guides and articles