Pub-AI: AI in Science

Latent Dynamics · 2024-05-20

Diffusion for World Modeling: Visual Details Matter in Atari

Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos Storkey, Tim Pearce, François Fleuret

arXiv:2405.12399PDFProject page

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

DIAMOND trains an RL agent inside a diffusion world model, arguing that discrete-latent compression can lose visual details that matter for control; it reaches a mean human-normalized score of 1.46 on Atari 100k.

Why it matters

Applies diffusion models to world modelling and reports the best Atari 100k score among agents trained entirely inside a world model. The diffusion model can also run as an interactive game engine (demonstrated on Counter-Strike: Global Offensive).

Method

  • Diffusion world model predicts the next observation from a short frame stack and the action.
  • Design choices (training objective, few denoising steps) make it stable over long horizons.
  • Uses only 3 network function evaluations per frame versus 16 for IRIS.

Limitation

The authors list: evaluation focuses on discrete control (continuous domains are untested), and conditioning on stacked frames is a minimal memory mechanism, so longer-term memory would need something like an autoregressive Transformer.

Lineage

Abstract (from arXiv)

World models constitute a promising approach for training reinforcement learning agents in a safe and sample-efficient manner. Recent world models predominantly operate on sequences of discrete latent variables to model environment dynamics. However, this compression into a compact discrete representation may ignore visual details that are important for reinforcement learning. Concurrently, diffusion models have become a dominant approach for image generation, challenging well-established methods modeling discrete latents. Motivated by this paradigm shift, we introduce DIAMOND (DIffusion As a Model Of eNvironment Dreams), a reinforcement learning agent trained in a diffusion world model. We analyze the key design choices that are required to make diffusion suitable for world modeling, and demonstrate how improved visual details can lead to improved agent performance. DIAMOND achieves a mean human normalized score of 1.46 on the competitive Atari 100k benchmark; a new best for agents trained entirely within a world model. We further demonstrate that DIAMOND's diffusion world model can stand alone as an interactive neural game engine by training on static Counter-Strike: Global Offensive gameplay. To foster future research on diffusion for world modeling, we release our code, agents, videos and playable world models at https://diamond-wm.github.io.

Related papers