Encyclopedia · 187 concepts

Reinforcement Learning · advanced · concept 127 of 187

Reward Design & Reward Hacking

Designing reward functions is RL's hardest problem. Reward hacking: agents find unintended shortcuts to maximize reward without actually solving the task. A core challenge in AI alignment.

Key terms

Sparse rewardsReward shapingGoodhart's lawSpecification gaming

Guides and articles

Courses, papers, and more