Reinforcement Learning · advanced · concept 124 of 176
Policy Gradient & Actor-Critic
Policy gradient methods directly optimize the policy (action selection) rather than value estimates. Actor-Critic combines both: the actor selects actions, the critic evaluates them.
Key terms
REINFORCEAdvantage functionA2CA3CBaseline
Learn these first
Videos
▶ DeepMind x UCL RL Lecture Series - Policy-Gradient and Actor-Critic methods [9/13] ↗
Google DeepMind · YouTube
▶ Reinforcement Learning 6: Policy Gradients and Actor Critics ↗
Google DeepMind · YouTube
Guides and articles
Policy Gradient Algorithms | Lil'Log ↗
Lil'Log (OpenAI researcher)
Part 3: Intro to Policy Optimization — Spinning Up documentation ↗
OpenAI Spinning Up
Courses, papers, and more
This unlocks