Reinforcement Learning · advanced · concept 130 of 176
Inverse Reinforcement Learning
Learning the reward function from observing expert behavior, inferring WHAT the expert is optimizing, not just imitating their actions. Key for building AI that understands human preferences.
Key terms
Reward inferenceExpert demonstrationsImitation learningIRL
Learn these first
Videos
▶ RL Course by David Silver - Lecture 8: Integrating Learning and Planning ↗
Google DeepMind · YouTube
Guides and articles
CS 185/285 ↗
Berkeley CS285
Courses, papers, and more