Latent Dynamics · 2019-03-01
Model-Based Reinforcement Learning for Atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, Afroz Mohiuddin, Ryan Sepassi, George Tucker, Henryk Michalewski
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
SimPLe (Simulated Policy Learning) trains a video prediction model of an Atari game and learns a policy inside it, using only 100k environment interactions, about two hours of real-time play.
Why it matters
Among the first model-based approaches shown to be competitive with tuned model-free RL on Atari in a low-data regime; the authors report it outperforms state-of-the-art model-free algorithms in most games, in some by over an order of magnitude.
Method
- Learns a stochastic video prediction model of the game (a novel architecture is compared with several alternatives).
- Trains a PPO policy inside the learned model, alternating with collecting new real data.
- Evaluated at 100k agent-environment interactions per game.
Lineage
- ← contrasts by Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (Compared against SimPLe, the previous best model-based approach on Atari.)
Abstract (from arXiv)
Model-free reinforcement learning (RL) can be used to learn effective policies for complex tasks, such as Atari games, even from image observations. However, this typically requires very large amounts of interaction -- substantially more, in fact, than a human would need to learn the same games. How can people learn so quickly? Part of the answer may be that people can learn how the game works and predict which actions will lead to desirable outcomes. In this paper, we explore how video prediction models can similarly enable agents to solve Atari games with fewer interactions than model-free methods. We describe Simulated Policy Learning (SimPLe), a complete model-based deep RL algorithm based on video prediction models and present a comparison of several model architectures, including a novel architecture that yields the best results in our setting. Our experiments evaluate SimPLe on a range of Atari games in low data regime of 100k interactions between the agent and the environment, which corresponds to two hours of real-time play. In most games SimPLe outperforms state-of-the-art model-free algorithms, in some games by over an order of magnitude.