Skip to content
Launchpad Library logo
← All glossary terms

How It Works

VLA (vision-language-action model)

A model that looks at camera images, reads an instruction in plain language, and outputs actions for a robot to take — such as arm movements or steps. It extends a vision-language model so that its output is motion rather than words.

Origin

The term was popularized by Google DeepMind’s RT-2 paper (2023), which described its robot model as a vision-language-action model; later open models such as OpenVLA adopted the label.

Read the source →