Pub-AI: AI in Science

Latent Dynamics · 2019-12-03

Dream to Control: Learning Behaviors by Latent Imagination

Danijar Hafner, Timothy Lillicrap, Jimmy Ba, Mohammad Norouzi

arXiv:1912.01603PDFProject page

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

Dreamer learns behaviors purely by latent imagination: it trains an actor and a value function on trajectories imagined in a learned world model, backpropagating analytic value gradients through the learned dynamics.

Why it matters

Replaced PlaNet's online planning with a learned policy trained inside the world model, and reports better data efficiency, computation time and final performance across 20 visual control tasks.

Method

  • Uses the same latent world model as PlaNet, learned from past experience.
  • Actor-critic trained on imagined latent trajectories; value estimates are propagated with analytic gradients through the dynamics.
  • Predicts both actions and state values to handle long horizons beyond the imagination length.

Lineage

Abstract (from arXiv)

Learned world models summarize an agent's experience to facilitate learning complex behaviors. While learning world models from high-dimensional sensory inputs is becoming feasible through deep learning, there are many potential ways for deriving behaviors from them. We present Dreamer, a reinforcement learning agent that solves long-horizon tasks from images purely by latent imagination. We efficiently learn behaviors by propagating analytic gradients of learned state values back through trajectories imagined in the compact state space of a learned world model. On 20 challenging visual control tasks, Dreamer exceeds existing approaches in data-efficiency, computation time, and final performance.

Related papers