Pub-AI: AI in Science

Driving · 2026-09-30

World4Scorer: Outcome-Grounded World Modeling for Autonomous Driving

Jieyuan Pei, Meiyi Lu, Sining Ang, Yubo Zhao, Zhangyi Hu, Mingwei Xu, Haokai Ding, Wei Li, Zihan You, Jianwei Zheng, Li Yu, Yifeng Pan, Ji Tao, Rongjunchen Zhang, Yan Wang

arXiv:2609.36438PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

World4Scorer is an outcome-grounded, trajectory-conditioned JEPA-style world modeling framework for generate-and-select autonomous driving planners. A shared predictor maps each candidate plan to a latent state; simulator outcome labels supervise all candidates, while the observed future of the executed trajectory anchors shared parameters. Inertial re-ranking keeps consecutive selections consistent. The authors report state-of-the-art NAVSIM-v2 results and improved closed-loop Bench2Drive and manipulation on OGBench-Cube with a frozen LeWM world model.

Why it matters

The authors address a supervision gap in driving logs: only the executed trajectory has an observed future, while alternative candidates lack visual futures even though their predicted outcomes determine which plan is selected. World4Scorer combines simulator outcome labels for all candidates with realized-future anchoring for the executed one, using one shared predictor so the anchor can constrain parameters used to score unexecuted plans.

Method

  • Build the candidate scorer as a trajectory-conditioned predictor (JEPA-style): predict a state for each candidate, then score from candidate state.
  • Use simulator outcome labels to supervise predicted states for all generated and bank candidates; use the observed visual future of the executed trajectory to anchor the shared predictor.
  • Apply inertial re-ranking using a benchmark extended-comfort (EC) compatibility check and log-domain penalty to enforce continuity across frames.
Abstract (from arXiv)

Autonomous driving requires choosing a safe and efficient plan as surrounding traffic evolves. Generate-and-select planners propose multiple trajectories and score them for execution, and they have outperformed representative direct-prediction baselines on NAVSIM. Their scorer must compare plans that were never executed. Driving logs record the future of only the executed trajectory, so matching the logged future can leave predictions for the alternatives unconstrained; a simulator, in contrast, can label the outcome of every candidate. We introduce World4Scorer, which builds the scorer as a trajectory-conditioned JEPA-style predictor: it predicts a state for each candidate and reads the candidate's scores from that state. Simulator outcome labels supervise the states of all candidates, and the observed future of the executed trajectory anchors the predictor to real scene evolution. Because one predictor produces every candidate's state, the anchor can constrain shared parameters used to score unexecuted plans, while the future itself is needed only during training. Generated candidates mostly score well, so a scene-matched bank adds low-scoring plans to the outcome supervision; framewise choices can conflict, so inertial re-ranking keeps consecutive selections consistent. World4Scorer achieves state-of-the-art NAVSIM-v2 performance and a strong adapted-system result on closed-loop Bench2Drive. With the LeWM world model and planning budget fixed, outcome-based scoring also improves manipulation planning on the OGBench-Cube benchmark.

Related papers