Safety, Ethics & Governance · advanced · concept 175 of 176
Constitutional AI & RLAIF
Alignment where the feedback comes from AI guided by an explicit list of principles, a constitution, instead of thousands of human raters. The model critiques and revises its own outputs against the principles, making the values legible and editable rather than buried in rater preferences.
Key terms
ConstitutionSelf-critiqueRLAIFPrinciple-based feedbackHarmlessness
Learn these first
Where you meet it in the real world
How Claude is trained, scalable alignment pipelines, auditable AI values
Videos
▶ Reinforcement Learning from Human Feedback (RLHF) Explained ↗
IBM Technology · YouTube
Guides and articles