Reinforcement Learning · advanced · concept 126 of 176
Reward Design & Reward Hacking
Designing reward functions is RL's hardest problem. Reward hacking: agents find unintended shortcuts to maximize reward without actually solving the task. A core challenge in AI alignment.
Key terms
Sparse rewardsReward shapingGoodhart's lawSpecification gaming
Learn these first
Videos
▶ What is Al "reward hacking"—and why do we worry about it? ↗
Anthropic · YouTube
▶ Stanford CS221 I The AI Alignment Problem: Reward Hacking & Negative Side Effects I 2023 ↗
Stanford Online · YouTube
Guides and articles
Reward Hacking in Reinforcement Learning | Lil'Log ↗
Lil'Log (OpenAI researcher)
Courses, papers, and more