Encyclopedia · 176 concepts

NLP & Language · advanced · concept 97 of 176

RLHF (Reinforcement Learning from Human Feedback)

The technique that made ChatGPT helpful and safe. Trains a reward model on human preferences, then optimizes the LLM against it. The key ingredient between a base model and a useful assistant.

Key terms

Reward modelPPOPreference dataConstitutional AIDPO