Robotics · 2026-09-30
MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation
Wenbo Chen, Tianfu Li, Haoxuan Xu, Zhihao Cao, Zhenghan Chen, Zhengming Zhu, Zizhou Luo, Guosheng Yang, Yuan Liu, Lujia Wang, Wen Chen, Haoang Li
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
MVG-WAM is a world–action model for robotic manipulation that represents synchronized multi-view observations as geometrically related projections of one physical world. It combines an epipolar-constrained global state with view-indexed geometric states routed to corresponding video regions, and uses multi-horizon future metric-depth supervision for scale grounding without depth decoding at deployment. The authors report average success rates of 99.1% (LIBERO) and 92.07% (RoboTwin 2.0), plus 91.3% success on Cobot Magic.
Why it matters
The authors argue that common multi-view WAM interfaces (tiling images or concatenating tokens) leave calibrated cross-view geometry implicit, which makes it harder to connect global context with local interaction geometry. MVG-WAM explicitly structures the representation using calibrated epipolar constraints and camera-aware routing, and anchors geometry to metric scale via future-depth supervision, aiming to improve action prediction for manipulation.
Method
- Constructs a global state via epipolar-constrained cross-view retrieval in dense DINOv3 feature space to aggregate geometrically admissible correspondences.
- Encodes synchronized observations with a frozen geometry foundation model (VGGT-ΩΩ) to produce view-indexed register states; routes latent video tokens to camera regions using ν(i) and uses this in geometry-conditioned mixed attention.
- Uses multi-horizon future metric-depth supervision during training to metrically ground the geometry-aware representation, while disabling the depth decoder at inference.
Abstract (from arXiv)
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it harder to connect global scene context with the local geometry required for interaction. We introduce the Multi-View Geometry-Aware World-Action Model (MVG-WAM), which organizes these observations as related projections of one physical world rather than separate images on a canvas. Our model combines an epipolar-constrained global state with view-indexed geometric states jointly inferred from synchronized observations. Camera-aware routing supplies each video region with its corresponding geometric context and the shared global state, explicitly structuring the representation used for action prediction. We further ground the geometry-aware representation in metric scale through multi-horizon future-depth supervision, without requiring depth decoding during action rollout. MVG-WAM achieves average success rates of 99.1% on LIBERO and 92.07% on RoboTwin 2.0, demonstrating competitive performance across both benchmarks. Real-world experiments on Cobot Magic further demonstrate a 91.3% success rate across 150 trials spanning three manipulation tasks.
Related papers
- WorldLine: Action-Driven Visual Simulation for Robotic Manipulation
- Achieve What You Imagined: Learning to Align Actions with Visual Plans
- Copper-Policy: Focus on the Representation for Robust Robot Manipulation
- Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation
- EVO-WAM: Evolving World Action Models through Video-Action Verification