Pub-AI: AI in Science

JEPA · 2026-09-30

FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales

Shidu Ren, Qilin Gu, Zhenghao Ni, Junhan Sun, Jiaqi Wang, Damien Scieur, Yunze Liu

arXiv:2609.35138PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

FlexiWorld is a JEPA-based latent world model for goal-directed planning that learns from mixed-span goal supervision while using variable-length action chunks across multiple time scales. It jointly trains a flexible (variable-length) causal action encoder, latent predictor, and an autoregressive goal-conditioned actor with Student Forcing. For planning, it proposes ARCEM (actor-residual CEM), combining action-residual search with within-chunk feedback and chunk-boundary latent prediction.

Why it matters

The authors aim to improve long-horizon control by addressing two issues in prior JEPA world-model planners: (1) fixed-length action chunks that fix temporal granularity, and (2) limited goal supervision that does not directly train action generation for distant goals. FlexiWorld’s variable-length chunks and mixed-span goal supervision are used together, and ARCEM adapts action-residual CEM to autoregressive chunks.

Method

  • Train on windows with varying goal spans (35/55/75 primitive steps) and randomly partition actions into variable-length chunks; jointly supervise latent prediction and goal-conditioned action generation.
  • Use a variable-length causal Transformer action encoder plus an autoregressive actor that generates primitive actions sequentially; apply Student Forcing to reduce exposure bias.
  • Plan with ARCEM: action-residual CEM where perturbed actions condition subsequent outputs within a chunk, and latent predictions are made only at chunk boundaries; also supports Direct planning without search.

Limitation

The authors report a controlled search failure in TwoRoom at high temperature (T=0.8): selected plans can fail despite low predicted cost, and failures are localized to candidate ranking rather than absence of successful proposals.

Abstract (from arXiv)

Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world model that combines mixed-span goal supervision with variable-length action chunks to improve long-horizon control. During training, we sample varying goal spans and randomly partition the actions into variable-length chunks. We jointly train the world model with a causal action encoder that embeds variable-length chunks and an autoregressive actor that generates primitive actions sequentially. Student Forcing reduces exposure bias by training on generated action prefixes. For planning, Actor-Residual Cross-Entropy Method (ARCEM) combines action-residual search with within-chunk autoregressive feedback and chunk-boundary latent prediction. Across four benchmarks and goal distances, FlexiWorld with ARCEM achieves 89.29% mean success, compared with 83.98% for the strongest baseline. PushT ablations show improved direct control from mixed-span supervision, variable-length chunks, and Student Forcing. Without retraining, FlexiWorld supports different planning chunk lengths: longer chunks accelerate ARCEM by approximately $1.3\times$ on average while maintaining comparable average success.

Related papers