Encyclopedia · 176 concepts

Architectures · advanced · concept 65 of 176

Multimodal Models

AI models that process multiple types of data, text, images, audio, video, simultaneously. Today's frontier models from OpenAI, Google, and Anthropic are all natively multimodal; text-only models are the exception, not the rule.

Key terms

Cross-modal attentionVision-languageContrastive learningCLIP

Where you meet it in the real world

Image captioning, visual Q&A, document understanding, video analysis

Guides and articles

Courses, papers, and more