Pub-AI: AI in Science

Driving · 2026-09-30

V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving

Junwei You, Weizhe Tang, Can Wang, Yan Zhao, Jun Hua, Haotian Shi, Wei Zhang, Lin Wang, Bin Ran

arXiv:2609.37098PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

V2X-WAM is a cooperative world action model for end-to-end autonomous driving. It builds a reliability-aware spatiotemporal scene representation from ego and infrastructure observations, compressing infrastructure features into a compact quantized message. A multimodal planner generates prospective trajectories; the selected actions condition future occupancy and dynamic-flow prediction, and the predicted consequences are fed back to refine the final trajectory.

Why it matters

The authors aim to go beyond using V2X mainly to improve the current scene representation. V2X-WAM explicitly couples prospective action generation with future-world reasoning and then uses the predicted world consequences to refine the planned trajectory, enabling consequence-aware decision making in a cooperative driving setting.

Method

  • Reliability-aware cooperative scene encoding: shared scene representation from temporally aligned ego/infrastructure observations, with learned spatial reliability and temporal confidence regulating the contribution of a quantized infrastructure message.
  • Coupled world–action modeling: a multimodal planner produces prospective trajectories and an action sequence; predicted future traffic occupancy and dynamic flow are conditioned on the selected prospective action.
  • Feedback loop: predicted future-world consequences are fed back to refine the preliminary planned trajectory, forming an explicit action-to-world-to-action interaction.
Abstract (from arXiv)

Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providing valuable support for end-to-end autonomous driving. However, existing cooperative driving methods mainly exploit roadside information to enhance the representation of the current scene, while the future consequences of prospective driving actions are rarely modeled explicitly. This limits the ability of the planner to anticipate how its decisions may interact with the evolving traffic environment. To address this issue, we propose V2X-WAM, a cooperative world action model that tightly couples cooperative scene understanding, action generation, and future-world reasoning. V2X-WAM constructs a reliability-aware spatiotemporal representation from vehicle- and infrastructure-side observations, while compressing infrastructure information into a compact quantized message for efficient communication. Based on the resulting cooperative representation, a multimodal planner generates prospective trajectories, which explicitly condition future occupancy and dynamic-flow prediction. The predicted world consequences are then fed back to refine the planned trajectory, forming a closed interaction between action and future-world evolution. Experiments on a large-scale real-world cooperative driving dataset demonstrate that V2X-WAM consistently improves planning accuracy and safety over representative end-to-end cooperative driving methods, while achieving stronger future-world prediction and substantially lower communication overhead. Ablation studies further validate the effectiveness of the proposed design.

Related papers