OpenVLA

OpenVLA Collaboration

OpenVLA 7B

OpenVLA 7B

Checkpoint

Open 7B vision-language-action model pretrained on Open X-Embodiment.

Specifications

Architecture components
Autoregressive action tokens
Architecture components
VLM backbone
Backbone
Prismatic VLM with DINOv2, SigLIP and Llama 2
Frameworks
Transformers
Frameworks
Transformers / PyTorch
Input modalities
Language instruction
Input modalities
RGB camera images
Languages
English
Model type
Vision-language-action model
Output action space
Normalized 7-DoF end-effector deltas
Parameter count
7B parameters
Pipeline task
Image-text-to-text
Pipeline task
Robotics
Training data scale
970K manipulation episodes
Training data sources
Open X-Embodiment
Training data sources
Open X-Embodiment

Lineage

Architecture & I/O

ee pose · continuous

7 DoF

Normalized 7-DoF end-effector deltas

Default checkpoint

openvla/openvla-7b

openMIT License
Checkpoint

Other releases

1

published checkpoints

Supported robots

No verified robot compatibility recorded.

Benchmark results

No published benchmark results.

Deployment

Training data

Resources

Submit info