Architectures · advanced · concept 65 of 176
Multimodal Models
AI models that process multiple types of data, text, images, audio, video, simultaneously. Today's frontier models from OpenAI, Google, and Anthropic are all natively multimodal; text-only models are the exception, not the rule.
Key terms
Cross-modal attentionVision-languageContrastive learningCLIP
Learn these first
Where you meet it in the real world
Image captioning, visual Q&A, document understanding, video analysis
Videos
▶ Large Language Models explained briefly ↗
3Blue1Brown · YouTube
▶ Stanford CS25: V4 I From Large Language Models to Large Multimodal Models ↗
Stanford Online · YouTube
Guides and articles
Vision Language Models Explained ↗
Hugging Face
Generalized Visual Language Models | Lil'Log ↗
Lil'Log (OpenAI researcher)
Courses, papers, and more
This unlocks