Driving · 2026-09-30
RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts
Hongbin Lin, Chaoda Zheng, Yiming Yang, Xiangyu Li, Shijia Chen, Jinhao Deng, Kangjie Chen, Dongbin Zhang, Jie Feng, Yu Zhang, Xianming Liu, Shuguang Cui, Boyang Wang, Zhen Li
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
RoXDrive is a closed-loop RL post-training framework for end-to-end autonomous driving using an action-conditioned video world model. It introduces an Action-Vision Faithfulness Evaluator (AVFE) trained with geometry-aware auxiliary trajectory supervision to filter action-faithful long-horizon rollouts, then performs dense safety-aware scoring and scene-level closed-loop GRPO using only faithful episodes.
Why it matters
The authors target action-vision mismatch in action-conditioned video world models, where multi-step rollouts may not faithfully reflect the conditioning ego actions, causing unreliable policy optimization. By explicitly evaluating and retaining action-faithful rollouts, RoXDrive aims to enable reliable closed-loop reinforcement learning with video world models while addressing causal confusion from open-loop imitation training.
Method
- Pre-train a driving policy via imitation learning and train an Action-Vision Faithfulness Evaluator (AVFE) for inverse dynamics estimation; use geometry-aware auxiliary trajectory supervision to mitigate cumulative relative-motion errors in long-horizon inverse dynamics.
- During RL post-training, interact with a frozen action-conditioned video diffusion world model to generate long-horizon scene rollouts; filter rollouts using AVFE to retain only action-faithful episodes.
- Use Dense Safety-Aware Scoring (collision/lane clearance, ego progress, comfort, lane centering) and Scene-level Closed-Loop GRPO computed from groups of faithful rollouts for policy optimization.
Limitation
The authors note that ego-action faithfulness alone may not always guarantee correctness of the entire simulated environment, and that physically consistent scene evolution and realistic responses of surrounding agents to novel ego actions remain important challenges, particularly beyond the logged data distribution.
Abstract (from arXiv)
End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their own actions, leading to causal confusion in closed-loop real-world deployment. To address this issue, reinforcement learning (RL) post-training offers a promising alternative by leveraging world models as interactive training environments to enable future scene generation for policy improvement. Nevertheless, existing approaches either rely on reconstruction-based simulators, offering limited counterfactual interaction, or adopt synthetic simulators to enable long-horizon closed-loop interaction at the cost of a substantial sim-to-real gap. Recently, video world models have exhibited the ability to generate realistic multi-step future rollouts but may not faithfully reflect action conditions, resulting in action-vision mismatch. In this paper, we introduce RoXDrive, a plug-and-play closed-loop RL framework that enables reliable policy optimization by identifying action-faithful world-model rollouts, consisting of two stages: 1) Model pre-training: In addition to imitation-based policy pre-training, we devise an Action-Vision Faithfulness Evaluator for inverse dynamics estimation with our geometry-aware auxiliary trajectory supervision, enabling long-horizon assessment of whether visual dynamics faithfully reflect the conditioning ego actions. 2) Action-faithful RL post-training: Agents iteratively interact with world models to form long-horizon scene rollouts, retaining only action-faithful ones for dense safety-aware scoring and scene-level closed-loop RL post-training. Extensive experiments on nuScenes and an in-house dataset with over 130K training scenarios demonstrate consistent gains across planners, reducing safety violations by 27.6% with DiffusionDrive on nuScenes and 33.7% with Qwen3-VL on the internal data.
Related papers
- AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
- DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving
- PhysWAM: Physically Consistent World Action Model for Autonomous Driving
- V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving