Math & Optimization · intermediate · concept 46 of 176
Information Theory: Entropy & KL Divergence
Shannon's mathematics of surprise. Entropy measures how unpredictable a distribution is, cross-entropy is the loss nearly every classifier and LLM trains on, and KL divergence measures how far one distribution drifts from another, which is how RLHF keeps a tuned model close to its base.
Key terms
EntropyCross-entropy lossKL divergenceBitsPerplexity
Learn these first
Where you meet it in the real world
The loss function of LLMs, compression, perplexity benchmarks, the KL penalty in RLHF
Videos
▶ Entropy (for data science) Clearly Explained!!! ↗
StatQuest with Josh Starmer · YouTube
Guides and articles
Visual Information Theory -- colah's blog ↗
Chris Olah (Anthropic co-founder)