Pub-AI: AI in Science

Latent Dynamics · 2019-11-19

Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model

Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, David Silver

arXiv:1911.08265PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

MuZero combines tree-based search with a learned model that predicts only the quantities needed for planning (reward, policy and value), reaching superhuman performance in Go, chess and shogi without being given the rules and a new state of the art on 57 Atari games.

Why it matters

Shows that a model trained only to predict planning-relevant quantities, not to reconstruct observations, is enough for strong lookahead search even when the environment dynamics are unknown.

Method

  • Learned model predicts reward, action-selection policy and value function; no observation reconstruction.
  • Monte Carlo tree search runs in the learned latent space.
  • Same algorithm applies to Go, chess, shogi and Atari with no knowledge of the dynamics or rules.

Limitation

On Atari the authors observe that planning helped much less than in Go, perhaps because of greater model inaccuracy, and performance plateaued at around 100 simulations.

Lineage

Abstract (from arXiv)

Constructing agents with planning capabilities has long been one of the main challenges in the pursuit of artificial intelligence. Tree-based planning methods have enjoyed huge success in challenging domains, such as chess and Go, where a perfect simulator is available. However, in real-world problems the dynamics governing the environment are often complex and unknown. In this work we present the MuZero algorithm which, by combining a tree-based search with a learned model, achieves superhuman performance in a range of challenging and visually complex domains, without any knowledge of their underlying dynamics. MuZero learns a model that, when applied iteratively, predicts the quantities most directly relevant to planning: the reward, the action-selection policy, and the value function. When evaluated on 57 different Atari games - the canonical video game environment for testing AI techniques, in which model-based planning approaches have historically struggled - our new algorithm achieved a new state of the art. When evaluated on Go, chess and shogi, without any knowledge of the game rules, MuZero matched the superhuman performance of the AlphaZero algorithm that was supplied with the game rules.

Related papers