Encyclopedia · 176 concepts

Reinforcement Learning · advanced · concept 126 of 176

Reward Design & Reward Hacking

Designing reward functions is RL's hardest problem. Reward hacking: agents find unintended shortcuts to maximize reward without actually solving the task. A core challenge in AI alignment.

Key terms

Sparse rewardsReward shapingGoodhart's lawSpecification gaming

Guides and articles

Courses, papers, and more