Pub-AI: AI in Science

Driving · 2026-09-30

Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks

Guoheng Sun, Chen Chen, Jin Wang, Ang Li, Teresa Lv

arXiv:2609.36471PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

Staircase Policy (S-WAM) is a streaming inference/training framework for World-Action Models with large action chunks. It turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages, executing near-term actions early and continuously refining unexecuted actions by refreshing a lightweight future-latent predictor at each sub-chunk boundary.

Why it matters

WAMs add future-prediction overhead to expensive iterative action generation. While action chunking can amortize inference, effective throughput drops when later actions are conditioned on stale observations. S-WAM updates future latent conditioning during execution, enabling long-horizon execution without repeated full policy inference, and reports higher throughput and faster time-to-first-action while maintaining task performance.

Method

  • Turns a flow-matching VLA into a JEPA-style WAM by adding a lightweight future latent predictor; expert is conditioned on (h, f̂, e).
  • Streaming “denoising staircase”: partition chunk into K sub-chunks of size G; emit/read out the j-th sub-chunk after j+1 expert passes while denoising continues for the rest.
  • At each sub-chunk boundary, refresh the future latent from the latest observation (keeping backbone context fixed) and apply the updated condition to all unexecuted actions; optionally stop chunk when future-prediction error δ exceeds a threshold.

Limitation

When executing farther into a long chunk, the reliability of later actions decreases because they remain conditioned on the observation from which the chunk was conditioned; the paper notes that increasing executed horizon reduces success for prior vanilla large-chunk execution (e.g., by 12.8 points for LaWAM and 26.8 points for π0.5).

Abstract (from arXiv)

World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages. Near-term actions are executed as soon as they become available, while later actions continue to be refined. At each sub-chunk boundary, the future latent is re-predicted from the latest observation and used to update all unexecuted actions, enabling long-horizon execution without repeated full policy inference. The resulting future-prediction error can further serve as a signal for adaptive chunking. S-WAM achieves 97.7% on LIBERO and 87.9% on LIBERO-Plus, and improves performance across multiple policy backbones and real-robot tasks. It reaches 292.7 executed actions per second, $3.62\times$ the throughput of conventional execution at comparable accuracy, while reducing time-to-first-action from 123.6 to 73.3 ms. With additional inference optimizations, throughput further increases to 642.9 actions per second.

Related papers