Pub-AI: AI in Science

Driving · 2026-09-30

Dynamic Manipulation with World-Action Models via Counterfactual Planning

Sunwoo Park, Wonbin Lee, Seonghyun Jin, Youngmin Kim, Jangho Park, Jong Chul Ye

arXiv:2609.33172PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

World–Action models trained on static demonstrations can fail on moving targets due to “target-response collapse” during reactive replanning. The authors propose Dynamic Predictive Planning (DPP), which predicts interaction timing using WAM rollouts, forecasts the target’s interaction position, synthesizes a counterfactual observation in a familiar robot context, and connects the resulting canonical plan to the robot’s live execution. DPP improves dynamic manipulation in simulation and on a real robot without additional training on dynamic data.

Why it matters

The paper addresses a deployment gap: reusable pretrained manipulation skills often break when targets move, even if the skill exists for static contexts. DPP’s key intervention is decoupling plan-generation context from execution state via counterfactual canonical planning and an activation bridge. The authors report consistent improvements across diverse target motions, including outperforming baselines trained with dynamic data in simulation, while running real-time planning on a single consumer GPU.

Method

  • Identify target-response collapse and decouple planning context from execution state in DPP.
  • Predict interaction timing from a WAM predictive rollout; combine timing with observed target motion to forecast target’s future interaction position.
  • Generate a counterfactual canonical plan by editing the target position into a familiar robot context; connect it to the robot’s current execution state via an activation bridge, then replan from fresh observations.
Abstract (from arXiv)

World-Action models (WAMs) trained on static demonstrations often fail to manipulate moving targets even when they possess the required manipulation skills. We attribute this failure to target-response collapse: as execution advances, the policy becomes increasingly biased toward the learned continuation of its ongoing behavior and less responsive to target relocation. To bridge the gap between what the model has learned and what it can generate from the current context, we formulate dynamic manipulation as counterfactual planning by decoupling the context used for plan generation from the physical state used for execution. Our framework, Dynamic Predictive Planning (DPP), first uses the WAM's predictive rollout to estimate when an interaction is expected to occur, and combines this timing estimate with observed target motion to predict the target's future interaction position. DPP then constructs a counterfactual observation that places this predicted target position in a familiar robot context, allowing the model to invoke an existing manipulation skill rather than generate a recovery behavior from an unfamiliar robot-target configuration. The resulting plan is connected to the robot's actual state during execution. DPP enables real-time dynamic manipulation on a single consumer GPU without additional training on dynamic data. Experiments in simulation and on a real robot demonstrate consistent improvements across diverse target motions, with simulation performance surpassing all evaluated baselines, including methods additionally trained on dynamic data. Project page: https://methoder00.github.io/DPP/

Related papers