Robotics · 2026-09-30
FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation
Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala, Catherine Weaver, Mouli Sivapurapu, Kai Yang, Jackson Lee, Thomas Wolf, Pragna Mannam
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
FineART is a densely annotated bimanual manipulation dataset (40,543 episodes, 1,718 hours, 533,913 subtasks across 151 tasks). The authors introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask and executes continuous actions. Dense subtask training improves success (32.0%→100.0% on spatial disambiguation; 16.0%→76.0% on a long-horizon task). Mid-training enables data-efficient cross-robot transfer, and the authors open-source dataset, model weights, and training code.
Why it matters
The authors target long-horizon, multi-step bimanual manipulation with persistent, dense subtask supervision inside a single end-to-end policy. They report that subtask training and knowledge insulation improve out-of-distribution generalization, and that mid-training reduces fine-tuning data needs and enables zero-shot transfer to unseen tasks on new hardware.
Method
- FineART dataset: bimanual teleoperated episodes with whole-episode demonstration text and contiguous, non-overlapping subtask labels (free-text) spanning 100% of hours/tasks.
- FineART-VLA: a single end-to-end System-2 subtask predictor paired with a System-1 continuous-action controller, using hierarchical inference rather than splitting into separate planner/executor models.
- Training recipe uses dense subtask training plus knowledge insulation (KI) to prevent continuous-action gradients from degrading pretrained VLM representations.
Limitation
The paper excerpt does not state a specific limitation; null.
Abstract (from arXiv)
Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode while the rare bimanual effort that does label subtasks annotates only a fraction of its hours. We present FineART, a densely annotated bimanual manipulation dataset of 40,543 episodes, 1,718 hours, and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask, and show that mid-training it this way yields substantial gains. Specifically, success on a spatial disambiguation task increases from 32.0% to 100.0%, and step-by-step human subtask guidance lifts success on an unseen long-horizon task from 16.0% to 76.0%. Furthermore, after minimal fine-tuning on a new robot, the policy requires one-tenth the data of baselines without mid-training and generalizes zero-shot to completely unseen tasks on the new hardware. We open-source the full dataset, model weights, and training code.
Related papers
- RoboChrono: A Real Robot Benchmark for Streaming Task Understanding
- WorldLine: Action-Driven Visual Simulation for Robotic Manipulation
- T$^2$Mem: Learning Test-Time Memory for Robotics
- Copper-Policy: Focus on the Representation for Robust Robot Manipulation
- Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks