Latent Dynamics · 2023-10-25
TD-MPC2: Scalable, Robust World Models for Continuous Control
Nicklas Hansen, Hao Su, Xiaolong Wang
arXiv:2310.16828PDFProject page
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
TD-MPC2 improves TD-MPC, which does local trajectory optimization in the latent space of a decoder-free (implicit) world model; one set of hyperparameters works across 104 online RL tasks, and a single 317M-parameter agent performs 80 tasks across domains and embodiments.
Why it matters
Targets robustness and scaling in model-based control: the authors show capability grows with model and data size, and that a single agent can handle multiple embodiments and action spaces.
Method
- Implicit world model trained without a reconstruction (decoder) objective.
- Local trajectory optimization (MPC) in latent space to choose actions.
- Architectural changes let one world model handle multiple tasks, embodiments and action spaces.
Lineage
- contrasts → Mastering Diverse Domains through World Models (DreamerV3 is one of the main model-based baselines it is compared against.)
Abstract (from arXiv)
TD-MPC is a model-based reinforcement learning (RL) algorithm that performs local trajectory optimization in the latent space of a learned implicit (decoder-free) world model. In this work, we present TD-MPC2: a series of improvements upon the TD-MPC algorithm. We demonstrate that TD-MPC2 improves significantly over baselines across 104 online RL tasks spanning 4 diverse task domains, achieving consistently strong results with a single set of hyperparameters. We further show that agent capabilities increase with model and data size, and successfully train a single 317M parameter agent to perform 80 tasks across multiple task domains, embodiments, and action spaces. We conclude with an account of lessons, opportunities, and risks associated with large TD-MPC2 agents. Explore videos, models, data, code, and more at https://tdmpc2.com
Related papers
- Think Fast, Plan Selectively: Adaptive Deliberation for Efficient Data-Driven MPC
- QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
- Beyond a single latent space: a dual-latent world model for long-horizon planning
- FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales
- Beyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation Learning