Encyclopedia · 176 concepts

Reinforcement Learning · advanced · concept 124 of 176

Policy Gradient & Actor-Critic

Policy gradient methods directly optimize the policy (action selection) rather than value estimates. Actor-Critic combines both: the actor selects actions, the critic evaluates them.

Key terms

REINFORCEAdvantage functionA2CA3CBaseline