Latent Dynamics · 2023-01-10
Mastering Diverse Domains through World Models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy Lillicrap
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
DreamerV3 is a general model-based RL algorithm that learns a world model and improves behavior by imagining future scenarios; with a single configuration it outperforms specialized methods across over 150 tasks, and is the first to collect diamonds in Minecraft from scratch without human data.
Why it matters
Aims to remove per-domain tuning: robustness techniques based on normalization, balancing and transformations let one configuration work across very different domains, from Atari to Minecraft.
Method
- Third generation of the Dreamer line; learns a world model and trains actor and critic in imagination.
- Robustness techniques (normalization, balancing, transformations) enable stable learning across domains with fixed hyperparameters.
- Applied out of the box to Minecraft, where rewards are sparse and horizons long.
Lineage
- builds on → Mastering Atari with Discrete World Models (Third generation of Dreamer, following DreamerV2.)
- ← contrasts by TD-MPC2: Scalable, Robust World Models for Continuous Control (DreamerV3 is one of the main model-based baselines it is compared against.)
Abstract (from arXiv)
Developing a general algorithm that learns to solve tasks across a wide range of applications has been a fundamental challenge in artificial intelligence. Although current reinforcement learning algorithms can be readily applied to tasks similar to what they have been developed for, configuring them for new application domains requires significant human expertise and experimentation. We present DreamerV3, a general algorithm that outperforms specialized methods across over 150 diverse tasks, with a single configuration. Dreamer learns a model of the environment and improves its behavior by imagining future scenarios. Robustness techniques based on normalization, balancing, and transformations enable stable learning across domains. Applied out of the box, Dreamer is the first algorithm to collect diamonds in Minecraft from scratch without human data or curricula. This achievement has been posed as a significant challenge in artificial intelligence that requires exploring farsighted strategies from pixels and sparse rewards in an open world. Our work allows solving challenging control problems without extensive experimentation, making reinforcement learning broadly applicable.