Latent Dynamics · 2022-09-01
Transformers are Sample-Efficient World Models
Vincent Micheli, Eloi Alonso, François Fleuret
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
IRIS is a data-efficient RL agent that learns inside a world model built from a discrete autoencoder plus an autoregressive Transformer, and reaches a mean human-normalized score of 1.046 on Atari 100k (about two hours of gameplay).
Why it matters
Brings the sequence-modelling recipe of Transformers to world models: frames become discrete tokens and dynamics are predicted autoregressively. The authors report a new state of the art for methods without lookahead search.
Method
- Discrete autoencoder turns each frame into a small set of tokens.
- Autoregressive Transformer predicts the next frame tokens, reward and termination given actions.
- Policy is trained entirely in imagination inside this model.
Lineage
- ← contrasts by Diffusion for World Modeling: Visual Details Matter in Atari (Compared directly with IRIS, which uses discrete tokens and 16 function evaluations per frame.)
Abstract (from arXiv)
Deep reinforcement learning agents are notoriously sample inefficient, which considerably limits their application to real-world problems. Recently, many model-based methods have been designed to address this issue, with learning in the imagination of a world model being one of the most prominent approaches. However, while virtually unlimited interaction with a simulated environment sounds appealing, the world model has to be accurate over extended periods of time. Motivated by the success of Transformers in sequence modeling tasks, we introduce IRIS, a data-efficient agent that learns in a world model composed of a discrete autoencoder and an autoregressive Transformer. With the equivalent of only two hours of gameplay in the Atari 100k benchmark, IRIS achieves a mean human normalized score of 1.046, and outperforms humans on 10 out of 26 games, setting a new state of the art for methods without lookahead search. To foster future research on Transformers and world models for sample-efficient reinforcement learning, we release our code and models at https://github.com/eloialonso/iris.