Pub-AI: AI in Science

Driving · 2026-09-30

PhysWAM: Physically Consistent World Action Model for Autonomous Driving

Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye, Maulik Bhatt, Thang Vu, Charles Ochoa, Rowan McAllister, Igor Vasiljevic, Rajgopal Kannan, Viktor Prasanna, Vitor Guizilini, Yue Wang

arXiv:2609.37970PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

PhysWAM is a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. It introduces Coupled Point Projection (CPP) to unproject generated depth into 3D, transform it using generated SE(3) ego motion, and penalize the distance to LiDAR-transformed points using recorded ego motion, promoting physical consistency. Trajectory selection uses label-free consensus.

Why it matters

The authors address that joint world-action generation may not impose shared geometric constraints between predicted depth and ego motion. CPP couples these predictions by supervising the depth–motion pair against measured scene geometry through a single geometric loss. This enables physically consistent scene-action generation and is evaluated for driving planning/transfer and for future video and metric-depth prediction.

Method

  • Unified flow-matching transformer jointly denoises multiview video, metric depth, and ego motion, with bidirectional attention across modalities.
  • Coupled Point Projection (CPP) unprojects generated metric depth into 3D points, transforms them using generated SE(3) ego motion, and compares to LiDAR points transformed using recorded ego motion via a shared geometric loss.
  • Inference selects among multiple sampled trajectories using a parameter-free, label-free consensus rule based only on pairwise distances (no learned scorer or simulator feedback).

Limitation

CPP requires LiDAR sensor calibration to the cameras and an accurate ego trajectory, plus dense metric depth labels obtained from a reconstruction model anchored to the same LiDAR (and obstacle/drivable-area labels). Scaling to larger driving video without LiDAR would need a self-supervised form of coupling that the authors have not developed. CPP also enforces agreement only where LiDAR supports the recorded trajectory; consistency away from that trajectory is only measured indirectly.

Abstract (from arXiv)

World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated $\mathrm{SE}(3)$ ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.

Related papers