Encyclopedia · 187 concepts

Architectures · intermediate · concept 55 of 187

Self-Attention

A special case of attention where a sequence attends to itself, each token computes relationships with every other token. This is the core mechanism powering Transformers.

Key terms

QKV matricesAttention scoreMulti-headCausal masking

Learn these first

Guides and articles

Courses, papers, and more