Encyclopedia · 176 concepts

Safety, Ethics & Governance · advanced · concept 175 of 176

Constitutional AI & RLAIF

Alignment where the feedback comes from AI guided by an explicit list of principles, a constitution, instead of thousands of human raters. The model critiques and revises its own outputs against the principles, making the values legible and editable rather than buried in rater preferences.

Key terms

ConstitutionSelf-critiqueRLAIFPrinciple-based feedbackHarmlessness

Where you meet it in the real world

How Claude is trained, scalable alignment pipelines, auditable AI values