Driving · 2026-09-30
Think Fast, Plan Selectively: Adaptive Deliberation for Efficient Data-Driven MPC
Yi Xian Goh, Sze Jue Yang, Hao Luan
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
Fast-TD-MPC reduces per-step computation in data-driven MPC by adaptively switching between a fast amortized policy (System 1) and the original MPPI planner (System 2) using an OOD gate in TD-MPC2’s latent space. Across 103 continuous control tasks, it achieves up to ~4× faster inference, with robustness under external disturbances comparable to TD-MPC2 when planning is triggered selectively.
Why it matters
Online trajectory optimization can be too slow for real-time control because planners must sample and evaluate many candidate trajectories each step. Fast-TD-MPC aims to keep the robustness of planning while lowering latency by routing “routine” states to a cheap policy and reserving planning for out-of-distribution states detected online in a learned latent space.
Method
- Builds on TD-MPC2: uses the original MPPI planner as System 2 (slow/robust fallback), and a compact amortized MLP policy trained by planner imitation as System 1 (fast).
- Adds an OOD gating mechanism in the latent space: computes Mahalanobis distance from an ID latent distribution and selects System 1 vs. System 2 based on threshold τ.
- Uses an adaptive runtime profile where effective per-step latency scales with the fraction of timesteps routed to System 2; on System 1→System 2 transitions it zeroes the MPPI warm-start buffer to avoid stale plans.
Abstract (from arXiv)
Data-driven model predictive control (MPC) combines learned world models with online trajectory optimization, achieving strong performance in continuous control. However, the per-step cost of sampling and evaluating hundreds of candidate trajectories restricts deployment to control frequencies well below what real-time robotics demands. Motivated by the dual-process theory of human cognition, which distinguishes between fast, intuitive processing (System 1) and slower, deliberative reasoning (System 2), we ask whether every decision requires the same degree of computational deliberation. We propose Fast-TD-MPC, a lightweight framework that adaptively routes between fast policy execution and test-time planning, reserving costly deliberation for states where it is most needed. Fast-TD-MPC delivers competitive task performance across 103 continuous control tasks while achieving up to ~4x faster inference. Under external disturbances, Fast-TD-MPC selectively falls back to planning, maintaining robustness comparable to the original planner.
Related papers
- TD-MPC2: Scalable, Robust World Models for Continuous Control
- Don't Throw Away the Tail: Action Upcycling for Policy Acceleration
- Learning Latent Dynamics for Planning from Pixels
- What Must a World Model Distinguish for Planning?
- Beyond a single latent space: a dual-latent world model for long-horizon planning