Pub-AI: AI in Science

Digest

Reset

126 papers

InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning

2026-09-30 · Hongpei Zheng, Hujun Yin · arXiv:2609.37187

InsightMap is a framework that uses top-down maps as explicit spatial memory and action-conditioned prediction targets. It links historical views to labeled map locations and trains a shared multimodal backbone (based on BAGEL) to jointly predict navigation actions and post-action map generation, with QA and 3D grounding sharing the same spatial interface. The authors report improved navigation and competitive static spatial reasoning and grounding results.

Drivingauto-summary

V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents

2026-09-30 · Yang Zhang, Jiangyuan Zhao, Chenyou Fan et al. · arXiv:2609.37250

V-JEPA Policy trains a world-action model on the latent space of a frozen V-JEPA 2.1 encoder, without inheriting a pretrained visual generative model. The authors jointly learn an instruction-conditioned future-latent predictor and a flow-matching action expert from scratch in one downstream stage, using the predictor’s layer-wise context key–value states to condition action generation.

Latent Dynamicsauto-summary

Do-JEPA: From Masking to Intervention in Latent World Models

2026-09-30 · Hossein Resani, Javen Qinfeng Shi · arXiv:2609.37378

Do-JEPA trains latent world models using paired simulator rollouts from the same saved state under an action a and a reference action a∅, supervising the latent difference Δz. The objective decomposes into effect, support, propagation, and invariance losses. In pixels and learned slots, the authors report improved action-causal (responsive) effect prediction versus masking-style baselines, and in fine-tuning reduces the factual-cost of training from scratch.

Latent Dynamicsauto-summary

Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation

2026-09-30 · Xiangcheng Zhan, Zirui Chen, Yicheng Zhao et al. · arXiv:2609.37398

DEWO (Direct Experience World-Model Optimization) is a post-deployment learning paradigm for World-Action Models that refines world representations using visual futures collected around interaction turning points. It trains from matched successful and failed continuations for outcome-conditioned visual prediction, and uses a progress-estimating value head to activate classifier-free guidance when progress stalls. Across DexJoCo tasks and real-world dexterous-hand grids, the authors report improved control and predictive losses.

Drivingauto-summary

Anisotropic Representations Improve Planning in JEPA World Models

2026-09-30 · Mingu Kang, Yoori Oh, Sookyung Kim et al. · arXiv:2609.37441

The authors show that joint training of JEPA-style latent world models with isotropic Gaussian regularization can learn a representation geometry whose Euclidean latent planning cost ranks feasible outcomes differently from the task cost, even with accurate prediction. They propose AnisoWM with ΛReg: replacing the fixed isotropic target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. Across four visual goal-planning environments, it improves planning success over LeWorldModel in all four.

JEPAauto-summary

Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning

2026-09-30 · Merve Atasever, Keyan Azbijari, Cagan Bakirci et al. · arXiv:2609.37519

Video2STL converts observation-only videos into parametric Signal Temporal Logic (STL) specifications for robot learning. A VLM extracts an embodiment-independent semantic event trace, then generates a bank of parametric STL formulas; robot trajectories ground predicate thresholds and temporal bounds. The method uses two timescales: rolling-window robustness for dense local rewards and a causal monitor over a retained long-horizon formula for progress rewards, enabling cross-embodiment transfer.

Latent Dynamicsauto-summary

RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

2026-09-30 · Shifeng Bao, Fanding Huang, Yihan Lin et al. · arXiv:2609.37583

RoboHarn-Evo evolves Hierarchical Physical Knowledge (HPK) from physical interaction experience to improve long-horizon robotic manipulation without updating the base vision-language model or low-level executor. It uses dual loops: an inner execution loop that retrieves Task Knowledge and Action Knowledge, and an outer knowledge-update loop that revises entries via task-completion and physical-effect checks. Experiments on RMBench and RoboDojo report large held-out and transfer gains.

Drivingauto-summary

Beyond a single latent space: a dual-latent world model for long-horizon planning

2026-09-30 · Delin Zhao, Zhengrong Yue, Shaobin Zhuang et al. · arXiv:2609.37644

The authors report Dual-WM, a dual-latent world model that separates low-level execution from high-level long-range planning using distinct state representations/dynamics and learned macro-actions. They introduce LoRe to supervise self-generated recursive rollouts at both levels with horizon-weighted losses. On five goal-conditioned visual control tasks, the authors report improved long-horizon goal success versus task-wise strongest non-actor-guided baselines.

Latent Dynamicsauto-summary

Honeycomb: Constant-Size Scene Memory Representation for Video World Models

2026-09-30 · Jack Wei Lun Shi, Kaichen Zhou, Haoyu Chen et al. · arXiv:2609.37690

Honeycomb is a video world model that keeps persistent scene memory in a fixed-size HexMemory. HexMemory stores scene features as six fixed-dimension low-rank planes (three spatial, three spatiotemporal). A feed-forward writer updates only new video chunks by warping prior planes to expanded bounds and fusing with confidence-weighted pooling plus a learned residual. A reader reconstructs latent features from HexMemory to condition generation; the method avoids per-scene optimization and full-history reprocessing.

Video World Modelsauto-summary

Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics

2026-09-30 · Ojas Shirekar, Yash Surange, Agustinas Ju\v{c}as et al. · arXiv:2609.37708

BRAID (Bilevel Representations for Agent Interaction Dynamics) is a hierarchical sequential latent-variable model for generative multi-person social motion. It uses a group latent for shared interaction dynamics and person latents for individual behavior conditioned on the evolving group context, enabling coherent generation and compact social-state vectors for full, sparse, or partial observations.

Video World Modelsauto-summary

CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces

2026-09-30 · Sen Wang, Liu Liu, Xinjiang Wang et al. · arXiv:2609.37721

CogWAM is a cognition-guided world-action model for language-conditioned long-horizon robot manipulation. It maintains a persistent Semantic State that stores completed task events and the active subtask, updating only on semantic transitions. Progress-conditioned WORLD and ACTION queries use this state to condition future-world prediction (training only) and action generation (inference), aiming to keep predictions and actions aligned with task progress.

Roboticscodeauto-summary

MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

2026-09-30 · Wenbo Chen, Tianfu Li, Haoxuan Xu et al. · arXiv:2609.37793

MVG-WAM is a world–action model for robotic manipulation that represents synchronized multi-view observations as geometrically related projections of one physical world. It combines an epipolar-constrained global state with view-indexed geometric states routed to corresponding video regions, and uses multi-horizon future metric-depth supervision for scale grounding without depth decoding at deployment. The authors report average success rates of 99.1% (LIBERO) and 92.07% (RoboTwin 2.0), plus 91.3% success on Cobot Magic.

Roboticsauto-summary

GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets

2026-09-30 · Gaspard Bott\'e, S\'everin Baroudi, Samir Sadok et al. · arXiv:2609.37798

GLaS-JEPA is a speech SSL framework that predicts the current encoder’s continuous representations at masked positions, without contrastive learning, discrete targets, or a separate EMA target encoder. It prevents representation collapse using SIGReg representation-space regularization. Pretrained on 960h LibriSpeech, the 57M model achieves 6.89% WER on frozen-encoder SUPERB ASR and 25.87% CER on slot filling, outperforming listed non-distilled sub-90M baselines.

JEPAauto-summary

Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents

2026-09-30 · Sicheng Xie, Yitong Chen, Haidong Cao et al. · arXiv:2609.37810

The authors introduce RoboSkill, a skill acquisition-and-reuse framework organized as an Explore–Execute–Evolve loop for embodied agents. The agent explores to gather missing task information, executes tasks with visual and tactile feedback while running reusable code, then evolves a skill library from session records. On LIBERO-10, they report higher first-episode success and lower runtime, and similar gains on real robots.

Roboticscodeauto-summary

PhysWAM: Physically Consistent World Action Model for Autonomous Driving

2026-09-30 · Dhruv Parikh, Fengcheng Yu, Quankai Gao et al. · arXiv:2609.37970

PhysWAM is a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. It introduces Coupled Point Projection (CPP) to unproject generated depth into 3D, transform it using generated SE(3) ego motion, and penalize the distance to LiDAR-transformed points using recorded ego motion, promoting physical consistency. Trajectory selection uses label-free consensus.

Drivingauto-summary

doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving

2026-09-30 · Parthib Roy, Yash Tandon, Marcus Blennemann et al. · arXiv:2609.38028

doPlan is a publicly available real-world autonomous driving dataset with 5,154 human-written passenger instructions aligned to variable-duration nuPlan contexts (30.0–508.8 s). The authors report that language-conditioned models may change predictions without reliably improving directional instruction following, and that the maneuver referenced by passenger instructions often occurs far beyond common 5 s prediction horizons.

Drivingcodeauto-summary

EVO-WAM: Evolving World Action Models through Video-Action Verification

2026-09-30 · Shiyang Zhou, Xionghao Wu, Wenbo Li et al. · arXiv:2609.38057

EVO-WAM adapts world action models to unseen tasks using self-generated video-action trajectories without external action execution. It enables complete autoregressive rollouts via state prediction and anchored multi-frame context, then selects reliable prefixes using a vision-language model for task-completing visuals and an inverse dynamics model to verify video-action consistency. Iteratively training on verified prefixes improves success on unseen RoboTwin 2.0 tasks and real-world long-horizon composites.

Drivingauto-summary

WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

2026-09-30 · Shenghe Zheng, Wenbo Li, Jiyao Zhang et al. · arXiv:2609.38059

WorldLine is an action-driven visual simulator for robotic manipulation. It learns transferable robot–object dynamics from 10,000+ hours of action-free robot videos, then grounds heterogeneous actions using 2,000+ hours of action trajectories across 10+ embodiments via an image-space action interface and multi-view/failure-enriched relational training. A robot-focused few-step distillation enables efficient causal rollout for policy evaluation and planning.

Roboticsauto-summary

Stochastic World Models for Verifying Vision-Based Neural Feedback Systems

2026-09-30 · I. Samuel Akinwande, Mykel J. Kochenderfer, Clark Barrett · arXiv:2609.38120

The authors train stochastic, physically grounded perception surrogates (world models) for verifying vision-based neural feedback systems. Their verification procedure combines falsification, adaptive refinement, symbolic, and backward analyses. On an emergency braking benchmark, they report resolving the entire state space (38% previously unresolved) with a GAN surrogate, and resolving over 80% of the state space on the RGB benchmark with a world-model surrogate.

Drivingauto-summary

HelixWorld: A Real-time Interactive Audio-Visual World Model

2026-09-30 · Lei Ke, Jiahao Pan, Zeyue Tian et al. · arXiv:2609.38123

HelixWorld is a real-time interactive audio-visual world model where visual scenes and camera-grounded stereo sound co-evolve under user interaction. It trains a bidirectional teacher on a curated high-fidelity stereo spatial audio-visual dataset with metric camera poses, then distills it into a few-step causal streaming student using online trajectory distillation and streaming fine-tuning for 24 FPS on a single NVIDIA H800 GPU.

Drivingcodeauto-summary