Computer Vision · advanced · concept 108 of 176
Contrastive Learning & CLIP
Train two encoders so matching pairs land close in embedding space and mismatched pairs land far apart. CLIP did this for 400 million image-caption pairs and got zero-shot classification for free: name the classes in English and it picks one. The bridge most multimodal systems stand on.
Key terms
Contrastive lossShared embedding spaceZero-shot classificationImage-text pairsInfoNCE
Learn these first
Where you meet it in the real world
Visual search, image moderation, the text encoder guiding Stable Diffusion, dataset curation
Videos
▶ How AI 'Understands' Images (CLIP) - Computerphile ↗
Computerphile · YouTube
▶ Supervised Contrastive Learning ↗
Yannic Kilcher · YouTube
Guides and articles
Contrastive Representation Learning | Lil'Log ↗
Lil'Log (OpenAI researcher)
Courses, papers, and more