Driving
World models for autonomous driving. (41 papers)
2026-09-30 · Ryozo Masukawa, Sanggeon Yun, Raheeb Hassan et al. · arXiv:2609.31893
CyberWorld is a Dreamer-style world model for autonomous cyber defense that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of a defended network. Using CyberWheel, the graph-based variant reaches the deploy_then_stop control after 3.6k–15.8k environment steps (vs model-free PPO needing 2.3M–3.1M steps or failing within a 3.2M budget).
Drivingauto-summary
2026-09-30 · Yi Xian Goh, Sze Jue Yang, Hao Luan · arXiv:2609.32591
Fast-TD-MPC reduces per-step computation in data-driven MPC by adaptively switching between a fast amortized policy (System 1) and the original MPPI planner (System 2) using an OOD gate in TD-MPC2’s latent space. Across 103 continuous control tasks, it achieves up to ~4× faster inference, with robustness under external disturbances comparable to TD-MPC2 when planning is triggered selectively.
Drivingauto-summary
2026-09-30 · Dongsheng Liu, Chao Jin, Wenkui Yang et al. · arXiv:2609.32679
GUI world models often predict future interfaces using only the current GUI observation and action, which can cause state aliasing: identical visible conditions correspond to different valid futures due to hidden transition-relevant state. The authors introduce StateAliasBench (strict-pair diagnostic) and a predictive-state recovery method that infers structured state from history and condition frozen GUI world models to restore state-sensitive prediction and improve AndroidWorld agent performance.
Drivingauto-summary
2026-09-30 · Zexin Feng, Yixu Feng, Lingyu Xiao et al. · arXiv:2609.32779
Copper-Policy learns a compact, task-conditioned world representation jointly with the policy using temporal joint-embedding prediction. It predicts future observation embeddings (not pixels) and decodes actions from current-frame details plus the learned compact representation; no future is generated at test time. The authors report strong control and efficient training, including a 2B model trained in 9.67 hours on 8×RTX 5090 and faster training than Fast-WAM on matched A100 GPUs.
Drivingauto-summary
2026-09-30 · Rongzhe Wei, Hans Hao-Hsun Hsu, Peizhi Niu et al. · arXiv:2609.33030
The authors formalize when a world model must preserve physical distinctions for planning, via mechanism/response/decision sufficiency and query-dependent resolution requirements. Experiments on collision, learned nonlinear dynamics, and robotic planning show that needed information depends on the query, candidate set, and planning stage. They also compare query placement: a modular design (query-aware proposal + query-independent action-conditioned prediction) reuses physical predictions across objectives and achieves lower held-out regret than action-conditioned prediction methods.
Drivingauto-summary
2026-09-30 · Sunwoo Park, Wonbin Lee, Seonghyun Jin et al. · arXiv:2609.33172
World–Action models trained on static demonstrations can fail on moving targets due to “target-response collapse” during reactive replanning. The authors propose Dynamic Predictive Planning (DPP), which predicts interaction timing using WAM rollouts, forecasts the target’s interaction position, synthesizes a counterfactual observation in a familiar robot context, and connects the resulting canonical plan to the robot’s live execution. DPP improves dynamic manipulation in simulation and on a real robot without additional training on dynamic data.
Drivingauto-summary
2026-09-30 · Xuanlin Chen, Ziyue Wang, Xunlan Zhou et al. · arXiv:2609.33336
RECON improves model-based imitation learning by separating conservative policy learning from active real-environment data collection. It trains a main (conservative) policy for task execution and an explorer optimized using epistemic uncertainty conditioned on recoverability estimated from multi-step imagination with the main policy, prioritizing uncertain yet recoverable dynamics. Experiments across locomotion, navigation, and manipulation show consistent gains in interaction efficiency, imitation performance, and robustness.
Drivingauto-summary
2026-09-30 · Teng Cao, Yu Deng, Quentin Delfosse et al. · arXiv:2609.33728
ALDER discovers and revises explicit, equation-based world models by acting in the environment. It proposes parametric equation structures, fits coefficients, and uses an independent verifier on held-out data. When multiple hypotheses fit, an experiment selector queries cost- and safety-aware interventions where predictions disagree. Counterexamples update evidence and drive structural revisions; validated equations are used for prediction and control via an inverse problem.
Drivingauto-summary
2026-09-30 · Yuheng Qiao, Ziran Wei, Xiaohan Wang et al. · arXiv:2609.33832
The authors treat a world-action model’s visual predictions as a goal-conditioned visual proposal and use a frozen action-conditioned world model to predict action consequences. They compute dense feedback from cross-model prediction consistency plus terminal goal alignment, then use Flow Policy Optimization to update only the action head. Across four real-world UR5 tasks, they report mean success increasing from 43.4% to 75.1%, without online robot interaction or task-specific reward models.
Drivingauto-summary
2026-09-30 · Weiqi Wang, Yuxin Zhou, Mouxiang Chen et al. · arXiv:2609.33848
QwenGyre is an end-to-end framework for x Long-horizon online RL with black-box LLM agent harnesses. It uses an elastic scheduler to elastically reallocate GPUs between rollout and training without interrupting live executions, and a trajectory processor that reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training cost. On NL2RepoBench with Qwen3.8 2.4T, it improves score in 48 steps and reports end-to-end speedups over Colocate and Async.
Drivingauto-summary
2026-09-30 · Haoran Zhu, Wancong Zhang, Yann LeCun et al. · arXiv:2609.34085
AD-E2E-JEPA is an action-conditioned JEPA world model for end-to-end autonomous driving that enables goal-conditioned zero-shot planning without training a driving policy. It adds a SIGReg-regularized learnable projector that compresses patch embeddings, reducing planning latency (0.8s for an 8-frame rollout over 256 trajectories) while retaining planning performance on NAVSIMv2.
Drivingcodeauto-summary
2026-09-30 · Haojie Yang, Ran Su · arXiv:2609.34391
P2P predicts unpaired single-cell perturbation response at the population level by learning shared condition effects from two stochastic “views” of the same control/perturbed condition. It encodes control sets with a permutation-invariant set encoder, perturbations with a structured encoder, and blends an empirical memory with a neural residual to output population mean and gene-wise variance. Trained with cross-view supervision (no cell correspondences).
Drivingauto-summary
2026-09-30 · Taesung Kwon, Jangho Park, Sunwoo Park et al. · arXiv:2609.34911
Action Upcycling is a training-free way to extend the execution horizon of chunked robot policies by reusing “tail” actions the policy would otherwise discard. Using a velocity-fluctuation signal computed from a single sampled action chunk (no model internals, no extra samples), it executes tail actions until the predicted motion’s velocity starts to fluctuate. It reduces policy calls by 1.2–1.7× with no loss in success rate across several VLA models and a World Action Model, and works on real robots.
Drivingauto-summary
2026-09-30 · Yichao Liang, Amber Li, Dat Nguyen et al. · arXiv:2609.35047
EMPIRIC learns residual world models for robot planning by extending a base physics simulator with Python programs for missing physical mechanisms (forces, constraints, recurrent hidden state). It uses Bayesian inference to estimate program parameters and hidden state from noisy observations, chooses informative experiments, revises the programs when predictions fail, and plans under uncertainty. It reports solving more tasks with fewer environment interactions than three baselines across five simulated domains, and demonstrates on a physical robot.
Drivingauto-summary
2026-09-30 · Yiqi Su, Rashed Shelim, Lingyi Wang et al. · arXiv:2609.35545
EpiMind is a graph world-model framework for multi-region epidemic policy planning under shared-resource constraints. It uses a graph-factored recurrent state-space model (GF-RSSM) to generate joint policy-conditioned rollouts, and a graph-temporal ADMM (GT-ADMM) planner to coordinate region interventions, enforce per-period feasibility via projection, and evaluate temporal specifications on projected actions.
Drivingauto-summary
2026-09-30 · Zongze Wu, Yani Guo, Runnan Li · arXiv:2609.35811
Lookahead-R reframes tool retrieval as a resource-constrained sequential decision problem. It trains an execution-aware surrogate world model to predict tool execution success, latency cost, and semantic utility without invoking real APIs, then uses a cost-sensitive, uncertainty-guided Monte Carlo Tree Search to select tools under latency budgets. On ToolBench I3, the authors report NDCG@5 = 91.40%, outperforming ToolGen (90.16%) by 1.24%.
Drivingauto-summary
2026-09-30 · Jie Yang, Jiajun Chen, Jiazheng Zhou et al. · arXiv:2609.35916
VehicleArena is a high-fidelity 3D urban-driving benchmark for independently operating multi-agent driving with passenger requests and cockpit/cabin actions in a single closed loop. It provides 112 held-out evaluation tasks (80 single-agent, 32 multi-agent). The authors report that the highest arrival rates across nine evaluated models are only 65.0% (single-agent) and 65.6% (multi-agent), and driving policies can reduce surrounding vehicles’ arrival rates versus SUMO’s native traffic controller.
Drivingauto-summary
2026-09-30 · Yizheng Huang, Wensheng Lin, Lixin Li et al. · arXiv:2609.35936
This tutorial proposes embodied semantic communication (ESC) for collective autonomous agents: an explicit wireless link carries action-oriented embodied semantics (multimodal perceptual states, intrinsic kinematic/body capabilities, and collaborative intents) so heterogeneous receivers can parse, align, and ground them in local motor control within a closed-loop perception–action cycle. The paper also outlines system characteristics, environment-constrained pathways, supporting mathematical tools, and open challenges.
Drivingauto-summary
2026-09-30 · Anna Pavlenko, Bogdan Crivat, Brandon Haynes et al. · arXiv:2609.36323
The authors propose an AI Software Factory for Data Systems that automates SDLC stages (Targeting, Coding, Reviewing, Ops) using a World Model and an evolutionary coding system (Darwin) wrapped in a self-improvement loop. They report scaled Microsoft deployments (tens of repositories) yielding 3× engineering efficiency vs agentic coding and up to 22× token efficiency, plus OSS and production evidence.
Drivingauto-summary
2026-09-30 · Jieyuan Pei, Meiyi Lu, Sining Ang et al. · arXiv:2609.36438
World4Scorer is an outcome-grounded, trajectory-conditioned JEPA-style world modeling framework for generate-and-select autonomous driving planners. A shared predictor maps each candidate plan to a latent state; simulator outcome labels supervise all candidates, while the observed future of the executed trajectory anchors shared parameters. Inertial re-ranking keeps consecutive selections consistent. The authors report state-of-the-art NAVSIM-v2 results and improved closed-loop Bench2Drive and manipulation on OGBench-Cube with a frozen LeWM world model.
Drivingauto-summary