Encyclopedia · 187 concepts

Math & Optimization · intermediate · concept 47 of 187

Information Theory: Entropy & KL Divergence

Shannon's mathematics of surprise. Entropy measures how unpredictable a distribution is, cross-entropy is the loss nearly every classifier and LLM trains on, and KL divergence measures how far one distribution drifts from another, which is how RLHF keeps a tuned model close to its base.

Key terms

EntropyCross-entropy lossKL divergenceBitsPerplexity

Where you meet it in the real world

The loss function of LLMs, compression, perplexity benchmarks, the KL penalty in RLHF

Videos

▶ Entropy (for data science) Clearly Explained!!! ↗

StatQuest with Josh Starmer · YouTube

Guides and articles

Visual Information Theory -- colah's blog ↗

Chris Olah (Anthropic co-founder)