InsightMap is a framework that uses top-down maps as explicit spatial memory and action-conditioned prediction targets. It links historical views to labeled map locations and trains a shared multimodal backbone (based on BAGEL) to jointly predict navigation actions and post-action map generation, with QA and 3D grounding sharing the same spatial interface. The authors report improved navigation and competitive static spatial reasoning and grounding results.
2026-09-30 · Yang Zhang, Jiangyuan Zhao, Chenyou Fan et al. · arXiv:2609.37250
V-JEPA Policy trains a world-action model on the latent space of a frozen V-JEPA 2.1 encoder, without inheriting a pretrained visual generative model. The authors jointly learn an instruction-conditioned future-latent predictor and a flow-matching action expert from scratch in one downstream stage, using the predictor’s layer-wise context key–value states to condition action generation.
2026-09-30 · Hossein Resani, Javen Qinfeng Shi · arXiv:2609.37378
Do-JEPA trains latent world models using paired simulator rollouts from the same saved state under an action a and a reference action a∅, supervising the latent difference Δz. The objective decomposes into effect, support, propagation, and invariance losses. In pixels and learned slots, the authors report improved action-causal (responsive) effect prediction versus masking-style baselines, and in fine-tuning reduces the factual-cost of training from scratch.
DEWO (Direct Experience World-Model Optimization) is a post-deployment learning paradigm for World-Action Models that refines world representations using visual futures collected around interaction turning points. It trains from matched successful and failed continuations for outcome-conditioned visual prediction, and uses a progress-estimating value head to activate classifier-free guidance when progress stalls. Across DexJoCo tasks and real-world dexterous-hand grids, the authors report improved control and predictive losses.
2026-09-30 · Mingu Kang, Yoori Oh, Sookyung Kim et al. · arXiv:2609.37441
The authors show that joint training of JEPA-style latent world models with isotropic Gaussian regularization can learn a representation geometry whose Euclidean latent planning cost ranks feasible outcomes differently from the task cost, even with accurate prediction. They propose AnisoWM with ΛReg: replacing the fixed isotropic target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. Across four visual goal-planning environments, it improves planning success over LeWorldModel in all four.
Video2STL converts observation-only videos into parametric Signal Temporal Logic (STL) specifications for robot learning. A VLM extracts an embodiment-independent semantic event trace, then generates a bank of parametric STL formulas; robot trajectories ground predicate thresholds and temporal bounds. The method uses two timescales: rolling-window robustness for dense local rewards and a causal monitor over a retained long-horizon formula for progress rewards, enabling cross-embodiment transfer.
2026-09-30 · Shifeng Bao, Fanding Huang, Yihan Lin et al. · arXiv:2609.37583
RoboHarn-Evo evolves Hierarchical Physical Knowledge (HPK) from physical interaction experience to improve long-horizon robotic manipulation without updating the base vision-language model or low-level executor. It uses dual loops: an inner execution loop that retrieves Task Knowledge and Action Knowledge, and an outer knowledge-update loop that revises entries via task-completion and physical-effect checks. Experiments on RMBench and RoboDojo report large held-out and transfer gains.
The authors report Dual-WM, a dual-latent world model that separates low-level execution from high-level long-range planning using distinct state representations/dynamics and learned macro-actions. They introduce LoRe to supervise self-generated recursive rollouts at both levels with horizon-weighted losses. On five goal-conditioned visual control tasks, the authors report improved long-horizon goal success versus task-wise strongest non-actor-guided baselines.
2026-09-30 · Jack Wei Lun Shi, Kaichen Zhou, Haoyu Chen et al. · arXiv:2609.37690
Honeycomb is a video world model that keeps persistent scene memory in a fixed-size HexMemory. HexMemory stores scene features as six fixed-dimension low-rank planes (three spatial, three spatiotemporal). A feed-forward writer updates only new video chunks by warping prior planes to expanded bounds and fusing with confidence-weighted pooling plus a learned residual. A reader reconstructs latent features from HexMemory to condition generation; the method avoids per-scene optimization and full-history reprocessing.
BRAID (Bilevel Representations for Agent Interaction Dynamics) is a hierarchical sequential latent-variable model for generative multi-person social motion. It uses a group latent for shared interaction dynamics and person latents for individual behavior conditioned on the evolving group context, enabling coherent generation and compact social-state vectors for full, sparse, or partial observations.
2026-09-30 · Sen Wang, Liu Liu, Xinjiang Wang et al. · arXiv:2609.37721
CogWAM is a cognition-guided world-action model for language-conditioned long-horizon robot manipulation. It maintains a persistent Semantic State that stores completed task events and the active subtask, updating only on semantic transitions. Progress-conditioned WORLD and ACTION queries use this state to condition future-world prediction (training only) and action generation (inference), aiming to keep predictions and actions aligned with task progress.
MVG-WAM is a world–action model for robotic manipulation that represents synchronized multi-view observations as geometrically related projections of one physical world. It combines an epipolar-constrained global state with view-indexed geometric states routed to corresponding video regions, and uses multi-horizon future metric-depth supervision for scale grounding without depth decoding at deployment. The authors report average success rates of 99.1% (LIBERO) and 92.07% (RoboTwin 2.0), plus 91.3% success on Cobot Magic.
GLaS-JEPA is a speech SSL framework that predicts the current encoder’s continuous representations at masked positions, without contrastive learning, discrete targets, or a separate EMA target encoder. It prevents representation collapse using SIGReg representation-space regularization. Pretrained on 960h LibriSpeech, the 57M model achieves 6.89% WER on frozen-encoder SUPERB ASR and 25.87% CER on slot filling, outperforming listed non-distilled sub-90M baselines.
2026-09-30 · Sicheng Xie, Yitong Chen, Haidong Cao et al. · arXiv:2609.37810
The authors introduce RoboSkill, a skill acquisition-and-reuse framework organized as an Explore–Execute–Evolve loop for embodied agents. The agent explores to gather missing task information, executes tasks with visual and tactile feedback while running reusable code, then evolves a skill library from session records. On LIBERO-10, they report higher first-episode success and lower runtime, and similar gains on real robots.
PhysWAM is a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. It introduces Coupled Point Projection (CPP) to unproject generated depth into 3D, transform it using generated SE(3) ego motion, and penalize the distance to LiDAR-transformed points using recorded ego motion, promoting physical consistency. Trajectory selection uses label-free consensus.
2026-09-30 · Parthib Roy, Yash Tandon, Marcus Blennemann et al. · arXiv:2609.38028
doPlan is a publicly available real-world autonomous driving dataset with 5,154 human-written passenger instructions aligned to variable-duration nuPlan contexts (30.0–508.8 s). The authors report that language-conditioned models may change predictions without reliably improving directional instruction following, and that the maneuver referenced by passenger instructions often occurs far beyond common 5 s prediction horizons.
2026-09-30 · Shiyang Zhou, Xionghao Wu, Wenbo Li et al. · arXiv:2609.38057
EVO-WAM adapts world action models to unseen tasks using self-generated video-action trajectories without external action execution. It enables complete autoregressive rollouts via state prediction and anchored multi-frame context, then selects reliable prefixes using a vision-language model for task-completing visuals and an inverse dynamics model to verify video-action consistency. Iteratively training on verified prefixes improves success on unseen RoboTwin 2.0 tasks and real-world long-horizon composites.
WorldLine is an action-driven visual simulator for robotic manipulation. It learns transferable robot–object dynamics from 10,000+ hours of action-free robot videos, then grounds heterogeneous actions using 2,000+ hours of action trajectories across 10+ embodiments via an image-space action interface and multi-view/failure-enriched relational training. A robot-focused few-step distillation enables efficient causal rollout for policy evaluation and planning.
2026-09-30 · I. Samuel Akinwande, Mykel J. Kochenderfer, Clark Barrett · arXiv:2609.38120
The authors train stochastic, physically grounded perception surrogates (world models) for verifying vision-based neural feedback systems. Their verification procedure combines falsification, adaptive refinement, symbolic, and backward analyses. On an emergency braking benchmark, they report resolving the entire state space (38% previously unresolved) with a GAN surrogate, and resolving over 80% of the state space on the RGB benchmark with a world-model surrogate.
2026-09-30 · Lei Ke, Jiahao Pan, Zeyue Tian et al. · arXiv:2609.38123
HelixWorld is a real-time interactive audio-visual world model where visual scenes and camera-grounded stereo sound co-evolve under user interaction. It trains a bidirectional teacher on a curated high-fidelity stereo spatial audio-visual dataset with metric camera poses, then distills it into a few-step causal streaming student using online trajectory distillation and streaming fine-tuning for 24 FPS on a single NVIDIA H800 GPU.