2026-09-30 · Bang Du, Yichen Xie, Shuqi Zhao et al. · arXiv:2609.36413
The authors propose RoboActualizer to turn a pretrained video world model into a robot policy by “actualizing” a single task-conditioned future. A frozen V-JEPA 2.1 encoder provides a prior over plausible future latents; a tiny flow-matching head with two DiT experts learns task-conditioned selection and action realization. The approach is trained on cached latents and deployed with low latency for real-time control.
FineART is a densely annotated bimanual manipulation dataset (40,543 episodes, 1,718 hours, 533,913 subtasks across 151 tasks). The authors introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask and executes continuous actions. Dense subtask training improves success (32.0%→100.0% on spatial disambiguation; 16.0%→76.0% on a long-horizon task). Mid-training enables data-efficient cross-robot transfer, and the authors open-source dataset, model weights, and training code.
2026-09-30 · Jieyuan Pei, Meiyi Lu, Sining Ang et al. · arXiv:2609.36438
World4Scorer is an outcome-grounded, trajectory-conditioned JEPA-style world modeling framework for generate-and-select autonomous driving planners. A shared predictor maps each candidate plan to a latent state; simulator outcome labels supervise all candidates, while the observed future of the executed trajectory anchors shared parameters. Inertial re-ranking keeps consecutive selections consistent. The authors report state-of-the-art NAVSIM-v2 results and improved closed-loop Bench2Drive and manipulation on OGBench-Cube with a frozen LeWM world model.
2026-09-30 · Wenliang Guo, Zhanbo Huang, Yu Kong · arXiv:2609.36454
DynamicHOI reconstructs hand-object interaction (HOI) trajectories from monocular RGB video using geometry-grounded diffusion refinement plus coupled hand-object physics. It derives hand generalized forces via articulated inverse dynamics and object wrenches via Newton–Euler dynamics, couples them through contact-force transfer, and recovers active hand actuation regularized by a probabilistic prior to suppress mechanically implausible motion.
2026-09-30 · Guoheng Sun, Chen Chen, Jin Wang et al. · arXiv:2609.36471
Staircase Policy (S-WAM) is a streaming inference/training framework for World-Action Models with large action chunks. It turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages, executing near-term actions early and continuously refining unexecuted actions by refreshing a lightweight future-latent predictor at each sub-chunk boundary.
2026-09-30 · Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico) et al. · arXiv:2609.36531
The authors evaluate physical prediction in video world models at “event boundaries,” where release/impact is shown but its consequence is withheld. Using 62 controlled real free-fall recordings (124 event-anchored clips) and human trajectory prediction, they find Runway and Veo often produce release and impact events above 93% but with late onset, while Cosmos-Predict-2.5 and MAGI-1 frequently suppress measurable consequences. They also find that plausible timing does not guarantee physically consistent motion.
2026-09-30 · Dawei Xu, Zhijie Wang · arXiv:2609.36582
The authors test whether dry, cohesionless soil internal friction angle ϕ can be identified from proprioceptive foot–ground force histories. They generate a 2D plane-strain continuum forward model (MPM) for a rotating leg, benchmark it against measured rotating-leg force histories, fit two Gaussian-process surrogates (one per force component), and perform Bayesian inversion over the full history. In matched-model recovery, 14 off-grid friction angles are recovered with median absolute error ~0.1° (max ~0.7°).
2026-09-30 · Yuzhou Wu, Longteng Fan, Zimeng Li et al. · arXiv:2609.36605
RoboChrono is a real-robot benchmark for streaming task understanding with 39 scenarios and 34,713 evaluation instances, using a causal observation boundary (only visual observations up to the query time). Across 18 multimodal models under zero-shot evaluation, the authors report large task-dependent capability differences, especially on temporal ordering, view matching, and action time localization, plus input-ablations showing different dependencies on visual evidence.
2026-09-30 · Haoyang Wu, Shoudong Han, Chaoyue Li et al. · arXiv:2609.36677
ReWorld-Track is a recursive event world model for language-guided multi-camera tracking that carries association uncertainty forward through handoffs with a persistent recurrent belief. It predicts the next camera, arrival time, and entry region, then updates this belief using posterior probabilities over candidate matches and a temporary-null (no-match) alternative. The authors report improved identity continuity and HOTA on CityFlowV2 and MTMMC.
2026-09-30 · Yize Liu, Huang Huang, Yining Hong et al. · arXiv:2609.36720
T²Mem is a test-time training framework that turns a memory-free pretrained vision-language-action policy into a robotics controller with fast-weight test-time memory. It uses an observation-grounded interface to encode vision-language history into compact fast weights via online self-supervised updates, and action supervision plus alternating memory–policy learning shapes what is stored and how it is used. On RoboMME (16 tasks), it improves average success from 17.93% to 56.83% over the memory-free base policy.
2026-09-30 · Renlong Wu, Guanqiao Wang, Xuan Shang et al. · arXiv:2609.36810
MeteoVerse is a unified weather-controllable video world model for camera-controlled future prediction from either sunny or adverse-weather inputs. It explicitly estimates observed and target weather states (rain/snow/fog), derives a weather transition, and uses a transition-aware mixture of weather experts (MeteoMoE) to inject residual weather features into a frozen video backbone. It also introduces a dataset of over 50K real-world weather clips with sunny pseudo-pairs and weather-intensity annotations.
DSWM is a decomposed spatio-temporal world model for demand-driven UAV base station repositioning. It uses a recurrent state-space model trained with an EMA-based latent predictive objective plus variance regularization, a differentiable service simulator head (association, probabilistic LoS, Shannon rate, capped fulfillment), and observation-anchored CEM planning that executes only the first action of the best imagined rollout.
2026-09-30 · Hongbin Lin, Chaoda Zheng, Yiming Yang et al. · arXiv:2609.36851
RoXDrive is a closed-loop RL post-training framework for end-to-end autonomous driving using an action-conditioned video world model. It introduces an Action-Vision Faithfulness Evaluator (AVFE) trained with geometry-aware auxiliary trajectory supervision to filter action-faithful long-horizon rollouts, then performs dense safety-aware scoring and scene-level closed-loop GRPO using only faithful episodes.
2026-09-30 · Jingnan Pu, Zi-En Fan, Feng Lian · arXiv:2609.36952
ER-JEPA adds an episodic replay pathway to LLM-JEPA. It stores training token pairs in a memory and, at each step, retrieves relevant stored examples to provide additional supervision for both token prediction and representation alignment. The replay path is removed after training, so inference matches LLM-JEPA.
The authors propose Abductive World Modeling (AWM), which learns structured causal representations by “predict forward, then abduce backward.” Using the Hierarchical Abductive State Pyramid (HASP), AWM infers a hierarchical latent state with Entity, Dynamic, and Relation components by jointly reasoning over the current observation and its predicted future. Experiments on physical prediction, causal reasoning, and action understanding show consistent gains over a V-JEPA 2 backbone-only baseline.
2026-09-30 · Tu Fangyuan, Xiangyue Zhang, Yiyi Cai et al. · arXiv:2609.37004
World2Motion adapts Cosmos 3 into a single-stage generator that produces scene-aware 3D human motion and corresponding video from a single image and a text prompt. The authors build a hybrid training dataset (synthetic video–motion pairs plus real videos paired with estimated 3D motion) and introduce a shift-decoupled noise schedule to reduce motion jitter. On a multi-source interaction benchmark, it improves motion–text alignment and scene interaction and matches two-stage interaction success with ~3.3× faster inference.
2026-09-30 · Kerui Ren, Yingxiang Xu, Kaiwen Song et al. · arXiv:2609.37089
Real2Gym is an agentic Real2Sim2Real framework that builds editable, physics-validated simulation gyms from human and robot demonstration videos. It reconstructs aligned Blender+MuJoCo scenes, validates/retargets actions via native physics, augments feasible task variations, and distills successful executions into reusable, object-relative skills for simulation and real robots without model weight updates.
2026-09-30 · Junwei You, Weizhe Tang, Can Wang et al. · arXiv:2609.37098
V2X-WAM is a cooperative world action model for end-to-end autonomous driving. It builds a reliability-aware spatiotemporal scene representation from ego and infrastructure observations, compressing infrastructure features into a compact quantized message. A multimodal planner generates prospective trajectories; the selected actions condition future occupancy and dynamic-flow prediction, and the predicted consequences are fed back to refine the final trajectory.
Waypoint-1.5 is a real-time diffusion video world model for interactive generation on consumer GPUs. It uses a single-stream causal Diffusion Transformer over TAEHV 1.5 latents and is conditioned on full keyboard+mouse inputs. Trained on 100,000 hours of control-aligned gameplay, it produces playable rollouts with a KV-cache runtime and distillation that reduces denoising to 4 steps per frame. It is benchmarked via latent FPS throughput.
2026-09-30 · Ziqi Wen, Ting Xu, Lianyu Wang et al. · arXiv:2609.37156
LucidWM is a world model that learns “doubt” about imagined transitions from experience using Subjective Logic, then turns doubt into “trust” that compounds along imagination. Trust reweights λ-returns for learning and guides action choice. It needs no extra parameters or forward passes for uncertainty estimation, and the authors report earlier alarms for drift and fewer steps to reach a goal in a navigation study.