Architectures · intermediate · concept 52 of 176
Transformer Architecture
THE architecture that changed everything (2017). Uses self-attention to process sequences in parallel, replacing RNNs. Foundation of GPT, BERT, and virtually every modern AI system.
Interactive · 3D
The transformer stack, explorable in 3D →
Attention beams, expert routing, the KV cache: click every stage.
Key terms
Self-attentionMulti-head attentionPositional encodingEncoder-decoder
Learn these first
Videos
▶ Transformers, the tech behind LLMs | Deep Learning Chapter 5 ↗
3Blue1Brown · YouTube
▶ Attention in transformers, step-by-step | Deep Learning Chapter 6 ↗
3Blue1Brown · YouTube
Guides and articles
Courses, papers, and more
This unlocks