Pub-AI: AI in Science

Driving · 2026-09-30

Don't Throw Away the Tail: Action Upcycling for Policy Acceleration

Taesung Kwon, Jangho Park, Sunwoo Park, Youngmin Kim, Seonghyun Jin, Youngjun Jun, Kyumin Choi, Jong Chul Ye

arXiv:2609.34911PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

Action Upcycling is a training-free way to extend the execution horizon of chunked robot policies by reusing “tail” actions the policy would otherwise discard. Using a velocity-fluctuation signal computed from a single sampled action chunk (no model internals, no extra samples), it executes tail actions until the predicted motion’s velocity starts to fluctuate. It reduces policy calls by 1.2–1.7× with no loss in success rate across several VLA models and a World Action Model, and works on real robots.

Why it matters

Chunked policies trade reactivity for efficiency by executing only a prefix of predicted actions and replanning frequently. The authors aim to reduce the number of policy calls without increasing inference cost per call and without relying on architecture-specific internal signals or extra sampling. Action Upcycling provides an orthogonal policy-acceleration axis: it preserves success rate while accelerating execution, and it can combine with other acceleration methods.

Method

  • Reuses a trustworthy part of each predicted action chunk’s tail: instead of executing exactly h actions then replanning, it executes h_exec = max_{c_k ≤ τ} k, choosing how many extra actions to run.
  • Defines a velocity fluctuation signal c_k by accumulating L2 changes between consecutive induced velocities beyond the default horizon; the tail is trusted while this signal stays small.
  • Selects the threshold τ by building a pool of c signals from chunks predicted by the policy itself and searching for the smallest τ that achieves a target mean execution length (upcycling ratio r).

Limitation

The authors report that the policy’s success rate drops beyond an upcycling ratio range, which they say is likely because a larger r admits tail actions with larger velocity fluctuation.

Abstract (from arXiv)

Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A short horizon keeps the policy reactive to the environment, but requires frequent policy calls. Recent test-time methods adaptively select the horizon for each chunk, but they either read model internals, where the signal must be chosen for each architecture, or draw extra samples, which adds cost. We propose Action Upcycling, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples. We find that discarded actions stay close to their replanned versions as long as the action velocity remains smooth. Action Upcycling therefore extends the execution horizon up to the point where the velocity begins to fluctuate. Extensive experiments on simulated and real-world manipulation tasks show that Action Upcycling reduces policy calls by 1.2-1.7x with no loss in success rate, across multiple Vision-Language-Action Models (VLAs) and even a World Action Model (WAM). It applies to any chunked policy at negligible cost and is orthogonal to other policy acceleration methods such as few-step sampling and streaming action decoding, opening a new axis for policy acceleration.

Related papers