Driving · 2023-09-18
DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving
Xiaofeng Wang, Zheng Zhu, Guan Huang, Xinze Chen, Jiagang Zhu, Jiwen Lu
arXiv:2309.09777PDFProject page
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
DriveDreamer is a diffusion-based world model learned from real-world driving data (evaluated on nuScenes) that generates controllable driving videos and future driving policies, trained in two stages (structured traffic constraints first, then future-state prediction).
Why it matters
Claims to be the first world model established from real-world driving scenarios rather than games or simulators; its synthetic data improved 3D detection and it produced open-loop planning results on nuScenes.
Method
- Auto-DM: a diffusion model over structured traffic information as the environment representation.
- Two-stage training: first learn traffic-structure constraints, then anticipate future states.
- Can also output driving actions conditioned on history and Auto-DM features.
Abstract (from arXiv)
World models, especially in autonomous driving, are trending and drawing extensive attention due to their capacity for comprehending driving environments. The established world model holds immense potential for the generation of high-quality driving videos, and driving policies for safe maneuvering. However, a critical limitation in relevant research lies in its predominant focus on gaming environments or simulated settings, thereby lacking the representation of real-world driving scenarios. Therefore, we introduce DriveDreamer, a pioneering world model entirely derived from real-world driving scenarios. Regarding that modeling the world in intricate driving scenes entails an overwhelming search space, we propose harnessing the powerful diffusion model to construct a comprehensive representation of the complex environment. Furthermore, we introduce a two-stage training pipeline. In the initial phase, DriveDreamer acquires a deep understanding of structured traffic constraints, while the subsequent stage equips it with the ability to anticipate future states. The proposed DriveDreamer is the first world model established from real-world driving scenarios. We instantiate DriveDreamer on the challenging nuScenes benchmark, and extensive experiments verify that DriveDreamer empowers precise, controllable video generation that faithfully captures the structural constraints of real-world traffic scenarios. Additionally, DriveDreamer enables the generation of realistic and reasonable driving policies, opening avenues for interaction and practical applications.
Related papers
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
- GAIA-1: A Generative World Model for Autonomous Driving
- PhysWAM: Physically Consistent World Action Model for Autonomous Driving
- World4Scorer: Outcome-Grounded World Modeling for Autonomous Driving
- V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving