Encyclopedia · 187 concepts

Reinforcement Learning · advanced · concept 125 of 187

Policy Gradient & Actor-Critic

Policy gradient methods directly optimize the policy (action selection) rather than value estimates. Actor-Critic combines both: the actor selects actions, the critic evaluates them.

Key terms

REINFORCEAdvantage functionA2CA3CBaseline