Encyclopedia · 176 concepts

Reinforcement Learning · advanced · concept 125 of 176

PPO (Proximal Policy Optimization)

For years the default RL algorithm, simpler and more stable than its predecessors, and the method used in the original RLHF pipelines. Newer group-relative variants have largely displaced it for language-model post-training, though PPO remains a standard baseline in robotics and control. Used in RLHF to fine-tune ChatGPT, in robotics, and in game AI. Balances exploration with stable training.

Key terms

ClippingTrust regionSurrogate objectiveKL penalty

Courses, papers, and more