Pub-AI: AI in Science

Digest

Reset

126 papers

One from Infinity: Actualizing Futures from Pretrained World Models into Robot Actions

2026-09-30 · Bang Du, Yichen Xie, Shuqi Zhao et al. · arXiv:2609.36413

The authors propose RoboActualizer to turn a pretrained video world model into a robot policy by “actualizing” a single task-conditioned future. A frozen V-JEPA 2.1 encoder provides a prior over plausible future latents; a tiny flow-matching head with two DiT experts learns task-conditioned selection and action realization. The approach is trained on cached latents and deployed with low latency for real-time control.

Roboticsauto-summary

FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation

2026-09-30 · Jade Choghari, Pepijn Kooijmans, Mansi Agarwal et al. · arXiv:2609.36416

FineART is a densely annotated bimanual manipulation dataset (40,543 episodes, 1,718 hours, 533,913 subtasks across 151 tasks). The authors introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask and executes continuous actions. Dense subtask training improves success (32.0%→100.0% on spatial disambiguation; 16.0%→76.0% on a long-horizon task). Mid-training enables data-efficient cross-robot transfer, and the authors open-source dataset, model weights, and training code.

Roboticsauto-summary

World4Scorer: Outcome-Grounded World Modeling for Autonomous Driving

2026-09-30 · Jieyuan Pei, Meiyi Lu, Sining Ang et al. · arXiv:2609.36438

World4Scorer is an outcome-grounded, trajectory-conditioned JEPA-style world modeling framework for generate-and-select autonomous driving planners. A shared predictor maps each candidate plan to a latent state; simulator outcome labels supervise all candidates, while the observed future of the executed trajectory anchors shared parameters. Inertial re-ranking keeps consecutive selections consistent. The authors report state-of-the-art NAVSIM-v2 results and improved closed-loop Bench2Drive and manipulation on OGBench-Cube with a frozen LeWM world model.

Drivingauto-summary

DynamicHOI: Coupled Dynamics for Physics-aware HOI Reconstruction

2026-09-30 · Wenliang Guo, Zhanbo Huang, Yu Kong · arXiv:2609.36454

DynamicHOI reconstructs hand-object interaction (HOI) trajectories from monocular RGB video using geometry-grounded diffusion refinement plus coupled hand-object physics. It derives hand generalized forces via articulated inverse dynamics and object wrenches via Newton–Euler dynamics, couples them through contact-force transfer, and recovers active hand actuation regularized by a probabilistic prior to suppress mechanically implausible motion.

Roboticsauto-summary

Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks

2026-09-30 · Guoheng Sun, Chen Chen, Jin Wang et al. · arXiv:2609.36471

Staircase Policy (S-WAM) is a streaming inference/training framework for World-Action Models with large action chunks. It turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages, executing near-term actions early and continuously refining unexecuted actions by refreshing a lightweight future-latent predictor at each sub-chunk boundary.

Drivingauto-summary

Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

2026-09-30 · Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico) et al. · arXiv:2609.36531

The authors evaluate physical prediction in video world models at “event boundaries,” where release/impact is shown but its consequence is withheld. Using 62 controlled real free-fall recordings (124 event-anchored clips) and human trajectory prediction, they find Runway and Veo often produce release and impact events above 93% but with late onset, while Cosmos-Predict-2.5 and MAGI-1 frequently suppress measurable consequences. They also find that plausible timing does not guarantee physically consistent motion.

Video World Modelsauto-summary

Inferring Soil Friction Angle from Robot Foot-Ground Force Histories: A Bayesian Inverse Approach to Proprioceptive Soil Sensing

2026-09-30 · Dawei Xu, Zhijie Wang · arXiv:2609.36582

The authors test whether dry, cohesionless soil internal friction angle ϕ can be identified from proprioceptive foot–ground force histories. They generate a 2D plane-strain continuum forward model (MPM) for a rotating leg, benchmark it against measured rotating-leg force histories, fit two Gaussian-process surrogates (one per force component), and perform Bayesian inversion over the full history. In matched-model recovery, 14 off-grid friction angles are recovered with median absolute error ~0.1° (max ~0.7°).

auto-summary

RoboChrono: A Real Robot Benchmark for Streaming Task Understanding

2026-09-30 · Yuzhou Wu, Longteng Fan, Zimeng Li et al. · arXiv:2609.36605

RoboChrono is a real-robot benchmark for streaming task understanding with 39 scenarios and 34,713 evaluation instances, using a causal observation boundary (only visual observations up to the query time). Across 18 multimodal models under zero-shot evaluation, the authors report large task-dependent capability differences, especially on temporal ordering, view matching, and action time localization, plus input-ablations showing different dependencies on visual evidence.

Drivingauto-summary

ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking

2026-09-30 · Haoyang Wu, Shoudong Han, Chaoyue Li et al. · arXiv:2609.36677

ReWorld-Track is a recursive event world model for language-guided multi-camera tracking that carries association uncertainty forward through handoffs with a persistent recurrent belief. It predicts the next camera, arrival time, and entry region, then updates this belief using posterior probabilities over candidate matches and a temporary-null (no-match) alternative. The authors report improved identity continuity and HOTA on CityFlowV2 and MTMMC.

auto-summary

T$^2$Mem: Learning Test-Time Memory for Robotics

2026-09-30 · Yize Liu, Huang Huang, Yining Hong et al. · arXiv:2609.36720

T²Mem is a test-time training framework that turns a memory-free pretrained vision-language-action policy into a robotics controller with fast-weight test-time memory. It uses an observation-grounded interface to encode vision-language history into compact fast weights via online self-supervised updates, and action supervision plus alternating memory–policy learning shapes what is stored and how it is used. On RoboMME (16 tasks), it improves average success from 17.93% to 56.83% over the memory-free base policy.

Drivingauto-summary

MeteoVerse: Unified Weather-Controllable Video World Model

2026-09-30 · Renlong Wu, Guanqiao Wang, Xuan Shang et al. · arXiv:2609.36810

MeteoVerse is a unified weather-controllable video world model for camera-controlled future prediction from either sunny or adverse-weather inputs. It explicitly estimates observed and target weather states (rain/snow/fog), derives a weather transition, and uses a transition-aware mixture of weather experts (MeteoMoE) to inject residual weather features into a frozen video backbone. It also introduces a dataset of over 50K real-world weather clips with sunny pseudo-pairs and weather-intensity annotations.

Video World Modelsauto-summary

DSWM: Decomposed Spatio-Temporal World Model for Demand-Driven UAV Base Station Repositioning

2026-09-30 · Shengjie Zhong, Zhongliang Zhao, Jingxuan Chen et al. · arXiv:2609.36845

DSWM is a decomposed spatio-temporal world model for demand-driven UAV base station repositioning. It uses a recurrent state-space model trained with an EMA-based latent predictive objective plus variance regularization, a differentiable service simulator head (association, probabilistic LoS, Shannon rate, capped fulfillment), and observation-anchored CEM planning that executes only the first action of the best imagined rollout.

Latent Dynamicsauto-summary

RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts

2026-09-30 · Hongbin Lin, Chaoda Zheng, Yiming Yang et al. · arXiv:2609.36851

RoXDrive is a closed-loop RL post-training framework for end-to-end autonomous driving using an action-conditioned video world model. It introduces an Action-Vision Faithfulness Evaluator (AVFE) trained with geometry-aware auxiliary trajectory supervision to filter action-faithful long-horizon rollouts, then performs dense safety-aware scoring and scene-level closed-loop GRPO using only faithful episodes.

Drivingauto-summary

ER-JEPA: Experience Replay Improves Joint-Embedding Predictive Learning in Language Models

2026-09-30 · Jingnan Pu, Zi-En Fan, Feng Lian · arXiv:2609.36952

ER-JEPA adds an episodic replay pathway to LLM-JEPA. It stores training token pairs in a memory and, at each step, retrieves relevant stored examples to provide additional supervision for both token prediction and representation alignment. The replay path is removed after training, so inference matches LLM-JEPA.

JEPAauto-summary

Abductive World Modeling via Causal Representation Learning

2026-09-30 · Ziqi Liu, Songhan Yang, Linfan Zhou et al. · arXiv:2609.36985

The authors propose Abductive World Modeling (AWM), which learns structured causal representations by “predict forward, then abduce backward.” Using the Hierarchical Abductive State Pyramid (HASP), AWM infers a hierarchical latent state with Entity, Dynamic, and Relation components by jointly reasoning over the current observation and its predicted future. Experiments on physical prediction, causal reasoning, and action understanding show consistent gains over a V-JEPA 2 backbone-only baseline.

Latent Dynamicscodeauto-summary

World2Motion: Turning Video World Models into 3D Human Motion Generators

2026-09-30 · Tu Fangyuan, Xiangyue Zhang, Yiyi Cai et al. · arXiv:2609.37004

World2Motion adapts Cosmos 3 into a single-stage generator that produces scene-aware 3D human motion and corresponding video from a single image and a text prompt. The authors build a hybrid training dataset (synthetic video–motion pairs plus real videos paired with estimated 3D motion) and introduce a shift-decoupled noise schedule to reduce motion jitter. On a multi-source interaction benchmark, it improves motion–text alignment and scene interaction and matches two-stage interaction success with ~3.3× faster inference.

Video World Modelsauto-summary

Real2Gym: Building Gyms from Videos, Bringing Skills to Robots

2026-09-30 · Kerui Ren, Yingxiang Xu, Kaiwen Song et al. · arXiv:2609.37089

Real2Gym is an agentic Real2Sim2Real framework that builds editable, physics-validated simulation gyms from human and robot demonstration videos. It reconstructs aligned Blender+MuJoCo scenes, validates/retargets actions via native physics, augments feasible task variations, and distills successful executions into reusable, object-relative skills for simulation and real robots without model weight updates.

Drivingauto-summary

V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving

2026-09-30 · Junwei You, Weizhe Tang, Can Wang et al. · arXiv:2609.37098

V2X-WAM is a cooperative world action model for end-to-end autonomous driving. It builds a reliability-aware spatiotemporal scene representation from ego and infrastructure observations, compressing infrastructure features into a compact quantized message. A multimodal planner generates prospective trajectories; the selected actions condition future occupancy and dynamic-flow prediction, and the predicted consequences are fed back to refine the final trajectory.

Drivingauto-summary

Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware

2026-09-30 · Rajit Rajpal, Shahbuland Matiana, Liew Wei Pyn et al. · arXiv:2609.37107

Waypoint-1.5 is a real-time diffusion video world model for interactive generation on consumer GPUs. It uses a single-stream causal Diffusion Transformer over TAEHV 1.5 latents and is conditioned on full keyboard+mouse inputs. Trained on 100,000 hours of control-aligned gameplay, it produces playable rollouts with a KV-cache runtime and distillation that reduces denoising to 4 steps per frame. It is benchmarked via latent FPS throughput.

Video World Modelsauto-summary

Lucid Dreaming for World Models: Learning to Doubt Imagination and Decide by Trust

2026-09-30 · Ziqi Wen, Ting Xu, Lianyu Wang et al. · arXiv:2609.37156

LucidWM is a world model that learns “doubt” about imagined transitions from experience using Subjective Logic, then turns doubt into “trust” that compounds along imagination. Trust reweights λ-returns for learning and guides action choice. It needs no extra parameters or forward passes for uncertainty estimation, and the authors report earlier alarms for drift and fewer steps to reach a goal in a navigation study.

Latent Dynamicsauto-summary