Pub-AI: AI in Science

Robotics · 2026-09-30

CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces

Sen Wang, Liu Liu, Xinjiang Wang, Zequn Chen, Haoyi Jiang, Taojun Ding, Tingyang Xiao, Zhizhong Su, Jie Wang, Sanping Zhou

arXiv:2609.37721PDFCode

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

CogWAM is a cognition-guided world-action model for language-conditioned long-horizon robot manipulation. It maintains a persistent Semantic State that stores completed task events and the active subtask, updating only on semantic transitions. Progress-conditioned WORLD and ACTION queries use this state to condition future-world prediction (training only) and action generation (inference), aiming to keep predictions and actions aligned with task progress.

Why it matters

The authors target a mismatch where combining semantic reasoning and future-world prediction can produce locally plausible actions that correspond to an already-completed (or not-yet-reached) subtask. CogWAM’s event-driven Semantic State and progress-conditioned WORLD/ACTION interface explicitly align task progress with both predictive world modeling and action generation, including a training scheme to address sparse transitions and train–inference mismatch.

Method

  • Maintains a persistent Semantic State S_t=(m_t,s_t) with completed task events and active subtask; it is updated only when observations indicate semantic transitions (KEEP vs UPDATE).
  • Uses progress-conditioned WORLD and ACTION query tokens conditioned on the same Semantic State: WORLD supports future multi-view latent prediction (training-time only) and ACTION generates continuous action chunks.
  • Introduces semantic training strategies: boundary-aware semantic sampling for sparse transitions and progressive semantic conditioning to reduce exposure bias by mixing annotated vs predicted semantic contexts during training.
Abstract (from arXiv)

Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM, a cognition-guided world-action model that establishes an explicit semantic interface between task reasoning and world-action learning through a persistent Semantic State, which stores completed task events and the active subtask. CogWAM updates this state only when observations indicate semantic transitions, allowing task-level context to persist across multiple action chunks. To bridge semantic context with physical prediction and control, CogWAM employs progress-conditioned WORLD and ACTION queries that selectively extract task-relevant information for future-world prediction and action generation. During training, the Semantic State provides shared task-progress context for both branches, while inference removes the future-prediction branch and directly generates actions from observations and the maintained state. We further introduce semantic training strategies to improve transition learning and closed-loop conditioning. Without additional robot-action pretraining, CogWAM achieves 15.56 / 11.70 % Score/SR on RoboDojo and state-of-the-art performance on BiCoord, while real-world experiments demonstrate closed-loop dual-arm manipulation with 16.4 fewer Semantic State regenerations than step-wise updating.

Related papers