Pub-AI: AI in Science

Driving

World models for autonomous driving. (41 papers)

CyberWorld: World Models for Sample-Efficient Autonomous Cyber Defense

2026-09-30 · Ryozo Masukawa, Sanggeon Yun, Raheeb Hassan et al. · arXiv:2609.31893

CyberWorld is a Dreamer-style world model for autonomous cyber defense that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of a defended network. Using CyberWheel, the graph-based variant reaches the deploy_then_stop control after 3.6k–15.8k environment steps (vs model-free PPO needing 2.3M–3.1M steps or failing within a 3.2M budget).

Drivingauto-summary

Think Fast, Plan Selectively: Adaptive Deliberation for Efficient Data-Driven MPC

2026-09-30 · Yi Xian Goh, Sze Jue Yang, Hao Luan · arXiv:2609.32591

Fast-TD-MPC reduces per-step computation in data-driven MPC by adaptively switching between a fast amortized policy (System 1) and the original MPPI planner (System 2) using an OOD gate in TD-MPC2’s latent space. Across 103 continuous control tasks, it achieves up to ~4× faster inference, with robustness under external disturbances comparable to TD-MPC2 when planning is triggered selectively.

Drivingauto-summary

The GUI Is Not the State: Diagnosing State Aliasing in GUI World Models

2026-09-30 · Dongsheng Liu, Chao Jin, Wenkui Yang et al. · arXiv:2609.32679

GUI world models often predict future interfaces using only the current GUI observation and action, which can cause state aliasing: identical visible conditions correspond to different valid futures due to hidden transition-relevant state. The authors introduce StateAliasBench (strict-pair diagnostic) and a predictive-state recovery method that infers structured state from history and condition frozen GUI world models to restore state-sensitive prediction and improve AndroidWorld agent performance.

Drivingauto-summary

Copper-Policy: Focus on the Representation for Robust Robot Manipulation

2026-09-30 · Zexin Feng, Yixu Feng, Lingyu Xiao et al. · arXiv:2609.32779

Copper-Policy learns a compact, task-conditioned world representation jointly with the policy using temporal joint-embedding prediction. It predicts future observation embeddings (not pixels) and decodes actions from current-frame details plus the learned compact representation; no future is generated at test time. The authors report strong control and efficient training, including a 2B model trained in 9.67 hours on 8×RTX 5090 and faster training than Fast-WAM on matched A100 GPUs.

Drivingauto-summary

What Must a World Model Distinguish for Planning?

2026-09-30 · Rongzhe Wei, Hans Hao-Hsun Hsu, Peizhi Niu et al. · arXiv:2609.33030

The authors formalize when a world model must preserve physical distinctions for planning, via mechanism/response/decision sufficiency and query-dependent resolution requirements. Experiments on collision, learned nonlinear dynamics, and robotic planning show that needed information depends on the query, candidate set, and planning stage. They also compare query placement: a modular design (query-aware proposal + query-independent action-conditioned prediction) reuses physical predictions across objectives and achieves lower held-out regret than action-conditioned prediction methods.

Drivingauto-summary

Dynamic Manipulation with World-Action Models via Counterfactual Planning

2026-09-30 · Sunwoo Park, Wonbin Lee, Seonghyun Jin et al. · arXiv:2609.33172

World–Action models trained on static demonstrations can fail on moving targets due to “target-response collapse” during reactive replanning. The authors propose Dynamic Predictive Planning (DPP), which predicts interaction timing using WAM rollouts, forecasts the target’s interaction position, synthesizes a counterfactual observation in a familiar robot context, and connects the resulting canonical plan to the robot’s live execution. DPP improves dynamic manipulation in simulation and on a real robot without additional training on dynamic data.

Drivingauto-summary

Beyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation Learning

2026-09-30 · Xuanlin Chen, Ziyue Wang, Xunlan Zhou et al. · arXiv:2609.33336

RECON improves model-based imitation learning by separating conservative policy learning from active real-environment data collection. It trains a main (conservative) policy for task execution and an explorer optimized using epistemic uncertainty conditioned on recoverability estimated from multi-step imagination with the main policy, prioritizing uncertain yet recoverable dynamics. Experiments across locomotion, navigation, and manipulation show consistent gains in interaction efficiency, imitation performance, and robustness.

Drivingauto-summary

ALDER: Discovering the Laws of a World by Acting in It

2026-09-30 · Teng Cao, Yu Deng, Quentin Delfosse et al. · arXiv:2609.33728

ALDER discovers and revises explicit, equation-based world models by acting in the environment. It proposes parametric equation structures, fits coefficients, and uses an independent verifier on held-out data. When multiple hypotheses fit, an experiment selector queries cost- and safety-aware interventions where predictions disagree. Counterexamples update evidence and drive structural revisions; validated equations are used for prediction and control via an inverse problem.

Drivingauto-summary

Achieve What You Imagined: Learning to Align Actions with Visual Plans

2026-09-30 · Yuheng Qiao, Ziran Wei, Xiaohan Wang et al. · arXiv:2609.33832

The authors treat a world-action model’s visual predictions as a goal-conditioned visual proposal and use a frozen action-conditioned world model to predict action consequences. They compute dense feedback from cross-model prediction consistency plus terminal goal alignment, then use Flow Policy Optimization to update only the action head. Across four real-world UR5 tasks, they report mean success increasing from 43.4% to 75.1%, without online robot interaction or task-specific reward models.

Drivingauto-summary

QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

2026-09-30 · Weiqi Wang, Yuxin Zhou, Mouxiang Chen et al. · arXiv:2609.33848

QwenGyre is an end-to-end framework for x Long-horizon online RL with black-box LLM agent harnesses. It uses an elastic scheduler to elastically reallocate GPUs between rollout and training without interrupting live executions, and a trajectory processor that reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training cost. On NL2RepoBench with Qwen3.8 2.4T, it improves score in 48 steps and reports end-to-end speedups over Colocate and Async.

Drivingauto-summary

AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving

2026-09-30 · Haoran Zhu, Wancong Zhang, Yann LeCun et al. · arXiv:2609.34085

AD-E2E-JEPA is an action-conditioned JEPA world model for end-to-end autonomous driving that enables goal-conditioned zero-shot planning without training a driving policy. It adds a SIGReg-regularized learnable projector that compresses patch embeddings, reducing planning latency (0.8s for an 8-frame rollout over 256 trajectories) while retaining planning performance on NAVSIMv2.

Drivingcodeauto-summary

P2P: Cross-View Population Denoising for Unpaired Single-Cell Perturbation Response Prediction

2026-09-30 · Haojie Yang, Ran Su · arXiv:2609.34391

P2P predicts unpaired single-cell perturbation response at the population level by learning shared condition effects from two stochastic “views” of the same control/perturbed condition. It encodes control sets with a permutation-invariant set encoder, perturbations with a structured encoder, and blends an empirical memory with a neural residual to output population mean and gene-wise variance. Trained with cross-view supervision (no cell correspondences).

Drivingauto-summary

Don't Throw Away the Tail: Action Upcycling for Policy Acceleration

2026-09-30 · Taesung Kwon, Jangho Park, Sunwoo Park et al. · arXiv:2609.34911

Action Upcycling is a training-free way to extend the execution horizon of chunked robot policies by reusing “tail” actions the policy would otherwise discard. Using a velocity-fluctuation signal computed from a single sampled action chunk (no model internals, no extra samples), it executes tail actions until the predicted motion’s velocity starts to fluctuate. It reduces policy calls by 1.2–1.7× with no loss in success rate across several VLA models and a World Action Model, and works on real robots.

Drivingauto-summary

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

2026-09-30 · Yichao Liang, Amber Li, Dat Nguyen et al. · arXiv:2609.35047

EMPIRIC learns residual world models for robot planning by extending a base physics simulator with Python programs for missing physical mechanisms (forces, constraints, recurrent hidden state). It uses Bayesian inference to estimate program parameters and hidden state from noisy observations, chooses informative experiments, revises the programs when predictions fail, and plans under uncertainty. It reports solving more tasks with fewer environment interactions than three baselines across five simulated domains, and demonstrates on a physical robot.

Drivingauto-summary

Graph World Models for Constrained Epidemic Policy Planning

2026-09-30 · Yiqi Su, Rashed Shelim, Lingyi Wang et al. · arXiv:2609.35545

EpiMind is a graph world-model framework for multi-region epidemic policy planning under shared-resource constraints. It uses a graph-factored recurrent state-space model (GF-RSSM) to generate joint policy-conditioned rollouts, and a graph-temporal ADMM (GT-ADMM) planner to coordinate region interventions, enforce per-period feasibility via projection, and evaluate temporal specifications on projected actions.

Drivingauto-summary

Lookahead-R: Budget-Aware Tool Retrieval via Execution-Centric Planning

2026-09-30 · Zongze Wu, Yani Guo, Runnan Li · arXiv:2609.35811

Lookahead-R reframes tool retrieval as a resource-constrained sequential decision problem. It trains an execution-aware surrogate world model to predict tool execution success, latency cost, and semantic utility without invoking real APIs, then uses a cost-sensitive, uncertainty-guided Monte Carlo Tree Search to select tools under latency budgets. On ToolBench I3, the authors report NDCG@5 = 91.40%, outperforming ToolGen (90.16%) by 1.24%.

Drivingauto-summary

VehicleArena: A Realistic Urban Environment for Multi-Agent Driving

2026-09-30 · Jie Yang, Jiajun Chen, Jiazheng Zhou et al. · arXiv:2609.35916

VehicleArena is a high-fidelity 3D urban-driving benchmark for independently operating multi-agent driving with passenger requests and cockpit/cabin actions in a single closed loop. It provides 112 held-out evaluation tasks (80 single-agent, 32 multi-agent). The authors report that the highest arrival rates across nine evaluated models are only 65.0% (single-agent) and 65.6% (multi-agent), and driving policies can reduce surrounding vehicles’ arrival rates versus SUMO’s native traffic controller.

Drivingauto-summary

Embodied Semantic Communication for Collective Autonomous Agents: A Tutorial on Representation, Wireless Delivery, and Closed-Loop Coordination

2026-09-30 · Yizheng Huang, Wensheng Lin, Lixin Li et al. · arXiv:2609.35936

This tutorial proposes embodied semantic communication (ESC) for collective autonomous agents: an explicit wireless link carries action-oriented embodied semantics (multimodal perceptual states, intrinsic kinematic/body capabilities, and collaborative intents) so heterogeneous receivers can parse, align, and ground them in local motor control within a closed-loop perception–action cycle. The paper also outlines system characteristics, environment-constrained pathways, supporting mathematical tools, and open challenges.

Drivingauto-summary

Towards an AI Software Factory for Data Systems

2026-09-30 · Anna Pavlenko, Bogdan Crivat, Brandon Haynes et al. · arXiv:2609.36323

The authors propose an AI Software Factory for Data Systems that automates SDLC stages (Targeting, Coding, Reviewing, Ops) using a World Model and an evolutionary coding system (Darwin) wrapped in a self-improvement loop. They report scaled Microsoft deployments (tens of repositories) yielding 3× engineering efficiency vs agentic coding and up to 22× token efficiency, plus OSS and production evidence.

Drivingauto-summary

World4Scorer: Outcome-Grounded World Modeling for Autonomous Driving

2026-09-30 · Jieyuan Pei, Meiyi Lu, Sining Ang et al. · arXiv:2609.36438

World4Scorer is an outcome-grounded, trajectory-conditioned JEPA-style world modeling framework for generate-and-select autonomous driving planners. A shared predictor maps each candidate plan to a latent state; simulator outcome labels supervise all candidates, while the observed future of the executed trajectory anchors shared parameters. Inertial re-ranking keeps consecutive selections consistent. The authors report state-of-the-art NAVSIM-v2 results and improved closed-loop Bench2Drive and manipulation on OGBench-Cube with a frozen LeWM world model.

Drivingauto-summary