Definition
Vision-language-action (VLA) models use visual observations and language instructions to produce actions for robots.
RoboAtlas
Explore vision-language-action models for robotics. Compare published versions, checkpoints, robot compatibility, licenses and benchmark evidence.
Vision-language-action (VLA) models use visual observations and language instructions to produce actions for robots.
Includes published model families with a released version classified as VLA in the RoboAtlas model catalog. Other embodied AI categories are listed separately.
Compare model versions, action spaces, checkpoints, licenses, robot compatibility and benchmark evidence on the linked model profiles.
A VLA label does not establish general-purpose capability. Training data, hardware requirements and evaluation conditions may be undisclosed.
13 published VLA model families
Catalog last updated: Sep 28, 2026
Open the VLA catalog, select released versions and compare their documented capabilities and sources.
Browse VLA models