Architectures · intermediate · concept 54 of 176
Self-Attention
A special case of attention where a sequence attends to itself, each token computes relationships with every other token. This is the core mechanism powering Transformers.
Key terms
QKV matricesAttention scoreMulti-headCausal masking
Learn these first
Videos
▶ Attention in transformers, step-by-step | Deep Learning Chapter 6 ↗
3Blue1Brown · YouTube
▶ Attention for Neural Networks, Clearly Explained!!! ↗
StatQuest with Josh Starmer · YouTube
Guides and articles
Courses, papers, and more