# OpenDriveLab/UniVLA

[RSS 2025] Learning to Act Anywhere with Task-centric Latent Actions

Repository: https://github.com/OpenDriveLab/UniVLA
Canonical: https://ross.abutalabs.com/products/univla
Homepage: https://arxiv.org/abs/2505.06111
Language: Python
License: Apache-2.0
License Family: permissive
Topics: robot-learning, vla, vision-language-actions-models
Last push: 2025-11-19T03:31:25+00:00

## Health v2 (maintenance only)
Score: 43/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 53, release rhythm 35, longevity 35
- inputs: {"age_days": 497, "days_push": 287, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1124, forks 69 (observed 2026-08-28T04:03:40.890995+00:00)

## What it is
UniVLA is an open-source framework for training cross-embodiment vision-language-action (VLA) robot policies using task-centric latent actions extracted from videos. It provides the full training recipe—latent action model, generalist policy pretraining, and post-training/evaluation pipelines for benchmarks like LIBERO, CALVIN, and real robots.

## Use cases
- train a vision-language-action policy for robot manipulation
- learn robot policies from cross-embodiment and human videos
- evaluate VLA models on LIBERO and CALVIN benchmarks
- deploy a generalist policy on a real robot arm
- extract task-centric latent actions from videos
- pretrain a robot foundation model with less compute than OpenVLA

## When to choose
- you need a compute-efficient, state-of-the-art VLA training recipe
- you want to leverage heterogeneous video data including human demonstrations
- you need cross-embodiment transfer for manipulation or navigation tasks

## When to avoid
- you need a production-ready robot control stack rather than research code
- you lack GPU resources for large-scale model training
- your task requires a single-embodiment, plug-and-play controller with no training

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, deep-learning, llm-training, simulation
- domain: robotics, machine-learning, autonomous-vehicles
- platform: python
- tags: vla, vision-language-action, robot-learning, cross-embodiment, latent-actions, manipulation, research-code, linux, gpu

## Member repositories
- OpenDriveLab/UniVLA (main) score 43

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:40.890995+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:39:37.355272+00:00, confidence not recorded.
  - readme: https://github.com/OpenDriveLab/UniVLA (fetched 2026-08-28T04:03:40.890995+00:00, sha fefd6f4c5f38)
  - homepage: https://arxiv.org/abs/2505.06111 (fetched 2026-08-29T12:44:16.803355+00:00, sha c1b520e59394)
  - site_page: https://info.arxiv.org/about/donate.html (fetched 2026-08-29T12:44:16.812881+00:00, sha cca9c3a11c56)
  - site_page: https://info.arxiv.org/about/ourmembers.html (fetched 2026-08-29T12:44:16.816525+00:00, sha 47cbc55ff1de)
  - site_page: https://info.arxiv.org/about (fetched 2026-08-29T12:44:16.818635+00:00, sha a1f16f915a9a)
  - site_page: https://info.arxiv.org/labs/index.html (fetched 2026-08-29T12:44:16.814824+00:00, sha b14a8d05a0ec)
- Data as of 2026-08-30T08:39:29.467469+00:00.
