Encyclopedia · 176 concepts

AI Agents & Applications · advanced · concept 163 of 176

Vision-Language-Action Models

Models that take camera images plus a natural-language instruction and output robot actions directly, trained across large mixed datasets of demonstrations rather than one policy per task. They port the pretrain-then-adapt pattern from language into physical control. The property under test is generalisation: whether a policy handles objects, scenes, and phrasings it never saw during training.

Key terms

Action tokensFlow matching policyTeleoperation dataGeneralist policyEmbodied AI

Where you meet it in the real world

Warehouse manipulation, household robots, industrial pick and place, humanoid control