Reinforcement Learning · intermediate · concept 123 of 176
Markov Decision Process (MDP)
The mathematical framework underlying RL. Defines states, actions, transition probabilities, and rewards. The Markov property: the future depends only on the current state, not history.
Key terms
State spaceAction spaceTransition functionDiscount factor
Videos
▶ Markov Decision Processes 1 - Value Iteration | Stanford CS221: AI (Autumn 2019) ↗
Stanford Online · YouTube
▶ RL Course by David Silver - Lecture 2: Markov Decision Process ↗
Google DeepMind · YouTube
Guides and articles
Part 1: Key Concepts in RL — Spinning Up documentation ↗
OpenAI Spinning Up
Courses, papers, and more