NLP & Language · advanced · concept 97 of 176
RLHF (Reinforcement Learning from Human Feedback)
The technique that made ChatGPT helpful and safe. Trains a reward model on human preferences, then optimizes the LLM against it. The key ingredient between a base model and a useful assistant.
Key terms
Reward modelPPOPreference dataConstitutional AIDPO
Learn these first
Videos
▶ Reinforcement Learning with Human Feedback (RLHF), Clearly Explained!!! ↗
StatQuest with Josh Starmer · YouTube
▶ Reinforcement Learning from Human Feedback (RLHF) Explained ↗
IBM Technology · YouTube
Guides and articles
Courses, papers, and more
This unlocks