Pub-AI: AI in Science

Driving · 2026-09-30

Rho: A Foundation for Efficiently Adaptable VLA Models

Rho Team, Simran Bagaria, Daphne Chen, Dean Fortier, Jianlong Fu, Michael Harrison, Tess Hellebrekers, Neel Joshi, Andrey Kolobov, Dalton Moore, Galen Mullins, Michael Murray, Eduardo Salinas, Reuben Tan

arXiv:2609.38164PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

Rho is an open-weights family of 5B-parameter vision-language-action (VLA) models for bimanual robotic manipulation across three embodiments (YAM Box, UR AI Trainer, FR3 Duo). The authors design regularized progressive adaptation with embodiment midtraining, and show improved downstream adaptation in simulation and physical-robot experiments. They also add online latent adaptation that learns from corrective feedback using as few as 15 corrected episodes.

Why it matters

The paper targets efficient adaptation for physical AI that must work across robot embodiments while maintaining control performance. By separating embodiment adaptation from task adaptation (via embodiment midtraining) and providing lightweight online latent-policy adaptation, the authors present a practical route from a general-purpose VLA to reliable task execution, with open-weights releases to enable deployment and further research.

Method

  • Train an open-weights Rho family from a Phi-Phy vision-language backbone, coupled to a flow-matching continuous action expert that generates action chunks from noise conditioned on observation and language.
  • Use regularized progressive adaptation: Rho-base foundation followed by multi-task single-embodiment midtraining to produce Rho-YAM-Box, Rho-UR-AI-Trainer, and Rho-FR3-Duo before offline task finetuning.
  • For online adaptation, enable an optional internal latent policy that learns to select observation-conditioned noise inputs for a frozen flow-matching action expert using corrective feedback; only the latent policy is updated.
Abstract (from arXiv)

General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo. We systematically ablate Rho's action-expert architecture and training recipe, and show in controlled simulation and physical-robot experiments that embodiment midtraining improves downstream adaptation. The resulting Rho variants for YAM Box, UR AI Trainer, and FR3 Duo match or outperform existing open-weights VLAs and achieve the strongest overall performance across the tasks, embodiments, and baselines evaluated in this report. We further demonstrate the Rho model family's built-in capacity for online adaptation: an internal latent policy learns from corrective feedback to select observation-conditioned noise inputs for the frozen flow-matching action expert. With as few as 15 corrected episodes, adapting this lightweight module enables Rho to handle task situations at the fringe of its offline finetuning distribution. Together, these results position Rho as both a strong general-purpose robotic manipulation model and a practical foundation for adaptation. We release the base Rho model and the embodiment-specific checkpoints to facilitate Rho's deployment in research experiments and practical industrial use cases.

Related papers