Pub-AI: AI in Science

Driving · 2023-09-29

GAIA-1: A Generative World Model for Autonomous Driving

Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, Gianluca Corrado

arXiv:2309.17080PDF

Editor pick

TL;DR

GAIA-1 is a generative world model for driving that takes video, text and action inputs, casts world modelling as next-token prediction over discrete tokens, and decodes with a video diffusion model to produce realistic driving scenes with control over ego-vehicle behaviour and scene features.

Why it matters

An early large generative world model for autonomous driving, pairing a token-based world model (6.5B parameters) with a diffusion decoder; the authors highlight emerging properties such as learning scene dynamics, contextual awareness and geometry.

Method

  • Video, text and actions are mapped to discrete tokens; the world model predicts the next token.
  • A separate video diffusion decoder turns the world model's predictions back into video.
  • Text and action conditioning give fine-grained control of the ego vehicle and scene.

Limitation

The authors note generation speed is slow, and suggest parallel sampling, fewer latent variables and better hardware as possible remedies.

Abstract (from arXiv)

Autonomous driving promises transformative improvements to transportation, but building systems capable of safely navigating the unstructured complexity of real-world scenarios remains challenging. A critical problem lies in effectively predicting the various potential outcomes that may emerge in response to the vehicle's actions as the world evolves. To address this challenge, we introduce GAIA-1 ('Generative AI for Autonomy'), a generative world model that leverages video, text, and action inputs to generate realistic driving scenarios while offering fine-grained control over ego-vehicle behavior and scene features. Our approach casts world modeling as an unsupervised sequence modeling problem by mapping the inputs to discrete tokens, and predicting the next token in the sequence. Emerging properties from our model include learning high-level structures and scene dynamics, contextual awareness, generalization, and understanding of geometry. The power of GAIA-1's learned representation that captures expectations of future events, combined with its ability to generate realistic samples, provides new possibilities for innovation in the field of autonomy, enabling enhanced and accelerated training of autonomous driving technology.

Related papers