2020-10-05 · Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi et al. · arXiv:2010.02193
DreamerV2 learns behaviors purely from predictions in the latent space of a world model that uses discrete latent representations and is trained separately from the policy, and is the first such agent to reach human-level performance on the 55-game Atari benchmark.
2019-12-03 · Danijar Hafner, Timothy Lillicrap, Jimmy Ba et al. · arXiv:1912.01603
Dreamer learns behaviors purely by latent imagination: it trains an actor and a value function on trajectories imagined in a learned world model, backpropagating analytic value gradients through the learned dynamics.
2019-11-19 · Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert et al. · arXiv:1911.08265
MuZero combines tree-based search with a learned model that predicts only the quantities needed for planning (reward, policy and value), reaching superhuman performance in Go, chess and shogi without being given the rules and a new state of the art on 57 Atari games.
2019-03-01 · Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos et al. · arXiv:1903.00374
SimPLe (Simulated Policy Learning) trains a video prediction model of an Atari game and learns a policy inside it, using only 100k environment interactions, about two hours of real-time play.
2018-11-12 · Danijar Hafner, Timothy Lillicrap, Ian Fischer et al. · arXiv:1811.04551
PlaNet learns a latent dynamics model from images and chooses actions by fast online planning in latent space, using a model with both deterministic and stochastic transition parts and a multi-step 'latent overshooting' training objective.
2018-03-27 · David Ha, Jürgen Schmidhuber · arXiv:1803.10122
Trains a VAE to compress each frame and a recurrent mixture-density network to predict the next latent, then fits a tiny linear controller on their features; the agent can also be trained entirely inside the model's own generated 'dream' and transferred back to the real environment.