AI Agents & Applications · advanced · concept 163 of 176
Vision-Language-Action Models
Models that take camera images plus a natural-language instruction and output robot actions directly, trained across large mixed datasets of demonstrations rather than one policy per task. They port the pretrain-then-adapt pattern from language into physical control. The property under test is generalisation: whether a policy handles objects, scenes, and phrasings it never saw during training.
Key terms
Action tokensFlow matching policyTeleoperation dataGeneralist policyEmbodied AI
Learn these first
Where you meet it in the real world
Warehouse manipulation, household robots, industrial pick and place, humanoid control
Videos
▶ What Are Vision Language Models? How AI Sees & Understands Images ↗
IBM Technology · YouTube
▶ Inside the World's Smartest Robot Brain [VLA] ↗
Welch Labs · YouTube
Guides and articles
Courses, papers, and more