Architectures · advanced · concept 64 of 176
Vision Transformer (ViT)
Applying the Transformer architecture to images by splitting them into patches and treating each patch as a token. Showed that attention can match or exceed CNNs for vision tasks at scale.
Key terms
Patch embeddingPosition embeddingCLS tokenImage patches
Learn these first
Videos
▶ Attention in transformers, step-by-step | Deep Learning Chapter 6 ↗
3Blue1Brown · YouTube
▶ An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (Paper Explained) ↗
Yannic Kilcher · YouTube
Guides and articles
Vision Transformer (ViT) · Hugging Face ↗
Hugging Face
11.8. Transformers for Vision — Dive into Deep Learning 1.0.3 documentation ↗
Dive into Deep Learning
Courses, papers, and more