Pub-AI: AI in Science

Driving · 2026-09-30

Copper-Policy: Focus on the Representation for Robust Robot Manipulation

Zexin Feng, Yixu Feng, Lingyu Xiao, Shang Su, Kexin Zheng, Chang Xu, Mengkai Shi, Shuo Feng, Xintao Yan

arXiv:2609.32779PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

Copper-Policy learns a compact, task-conditioned world representation jointly with the policy using temporal joint-embedding prediction. It predicts future observation embeddings (not pixels) and decodes actions from current-frame details plus the learned compact representation; no future is generated at test time. The authors report strong control and efficient training, including a 2B model trained in 9.67 hours on 8×RTX 5090 and faster training than Fast-WAM on matched A100 GPUs.

Why it matters

The paper argues that, for world action models, improved control may come less from generating detailed futures and more from shaping action-relevant representations. By learning the predictive target space jointly with control (and using asymmetric prediction/control inputs), Copper-Policy aims to produce a compact world representation that better separates task-driven change from perturbations while preserving current-frame spatial detail for execution.

Method

  • Learns a shared compact World representation z_t via an online–EMA compactor conditioned on task intention embeddings; targets come from an EMA compactor with stop-gradient.
  • Uses temporal joint-embedding prediction with conditional flow matching to predict future compact-view targets in embedding space (prediction conditioned on task intention).
  • Employs an asymmetric interface: Future Expert predicts only compact World tokens, while Action Expert also attends to current-frame visual tokens and robot state; no future tokens are instantiated at test time.

Limitation

The authors state: “These values are operational cost references rather than a controlled systems benchmark because the methods differ in model size, sequence length, distributed strategy, and implementation.”

Abstract (from arXiv)

World Action Models (WAMs) acquire behavioral priors by modeling future scene evolution, but predicting detailed futures in pixel or latent space incurs substantial cost. Recent evidence that co-training gains persist without test-time generation raises a question: what must a WAM learn to improve control? We introduce Copper-Policy, which learns a compact World representation with the policy rather than relying on a predefined target space. Through temporal joint-embedding prediction, it predicts future observation embeddings conditioned on task intention without reconstructing pixels. This prediction and action decoding shape the representation jointly, while the policy retains access to current-frame spatial detail for execution. Representation analyses show that the learned features better separate task-driven change from perturbations and provide complementary information for control. Compact prediction targets reduce training tokens per sample, enabling a 2B-parameter model trained in 9.67 hours on 8$\times$ RTX 5090 GPUs and 6$\times$ faster than Fast-WAM on matched A100 GPUs. Copper-Policy outperforms every compared method without embodied pretraining on RoboTwin and several embodied-pretrained VLAs on LIBERO-Plus (80.85%). On three challenging real-robot tasks, it performs comparably to $\pi_{0.5}$ and attains a higher average score. Together, these results show that Copper-Policy combines strong control performance with efficient training.

Related papers