Driving · 2024-05-27
Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
Shenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta, Yihang Qiu, Andreas Geiger, Jun Zhang, Hongyang Li
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
Vista is a generalizable driving world model that predicts at high resolution, supports many action controls (command, goal point, trajectory, angle and speed), and can act as a reward function to evaluate actions without ground-truth actions.
Why it matters
Targets three weaknesses of prior driving world models: generalization to unseen scenes, fidelity of critical details, and action controllability. The authors report it beats the best previous driving world model by 55% in FID and 27% in FVD.
Method
- Two new losses: one for moving instances (dynamics-aware) and one for structural information.
- Latent replacement injects historical frames as priors for coherent long-horizon rollouts.
- A unified conditioning interface adds high-level and low-level action controls through an efficient learning strategy.
Limitation
The authors call it an early endeavor with limits in computation efficiency, quality maintenance and training scale, noting possible degradation in long-horizon rollouts or drastic view shifts.
Abstract (from arXiv)
World models can foresee the outcomes of different actions, which is of paramount importance for autonomous driving. Nevertheless, existing driving world models still have limitations in generalization to unseen environments, prediction fidelity of critical details, and action controllability for flexible application. In this paper, we present Vista, a generalizable driving world model with high fidelity and versatile controllability. Based on a systematic diagnosis of existing methods, we introduce several key ingredients to address these limitations. To accurately predict real-world dynamics at high resolution, we propose two novel losses to promote the learning of moving instances and structural information. We also devise an effective latent replacement approach to inject historical frames as priors for coherent long-horizon rollouts. For action controllability, we incorporate a versatile set of controls from high-level intentions (command, goal point) to low-level maneuvers (trajectory, angle, and speed) through an efficient learning strategy. After large-scale training, the capabilities of Vista can seamlessly generalize to different scenarios. Extensive experiments on multiple datasets show that Vista outperforms the most advanced general-purpose video generator in over 70% of comparisons and surpasses the best-performing driving world model by 55% in FID and 27% in FVD. Moreover, for the first time, we utilize the capacity of Vista itself to establish a generalizable reward for real-world action evaluation without accessing the ground truth actions.
Related papers
- DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving
- GAIA-1: A Generative World Model for Autonomous Driving
- V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving
- RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts
- PhysWAM: Physically Consistent World Action Model for Autonomous Driving