2026-09-30 · Ji Dai, Quan Fang, Junyu Gao et al. · arXiv:2609.34604
SPRII (Shaping Persistent Representations from Independent Interactions) is a relation-supervised training principle for world models that learns persistent context from related independent interactions without numerical property labels. It uses two components: Align (related contexts agree) and Cross (use one interaction’s context to predict another’s future), while keeping the learner’s native objective. The paper analyzes Formation, Use, and Value links from accessible persistent information to task error reduction.
2026-09-30 · Beomsu Kim, Chieh-Hsin Lai, Bac Nguyen et al. · arXiv:2609.34677
The authors propose Future-Aware Recall (FAR), a framework for episodic memory access in world models. FAR learns which past memories to retrieve and which retrieval cues (time, pose, vision, audio) to trust by supervising a future-blind retriever using future-aware predictive utility during training (approximated by negative diffusion prediction loss).
2026-09-30 · Taesung Kwon, Jangho Park, Sunwoo Park et al. · arXiv:2609.34911
Action Upcycling is a training-free way to extend the execution horizon of chunked robot policies by reusing “tail” actions the policy would otherwise discard. Using a velocity-fluctuation signal computed from a single sampled action chunk (no model internals, no extra samples), it executes tail actions until the predicted motion’s velocity starts to fluctuate. It reduces policy calls by 1.2–1.7× with no loss in success rate across several VLA models and a World Action Model, and works on real robots.
2026-09-30 · Yichao Liang, Amber Li, Dat Nguyen et al. · arXiv:2609.35047
EMPIRIC learns residual world models for robot planning by extending a base physics simulator with Python programs for missing physical mechanisms (forces, constraints, recurrent hidden state). It uses Bayesian inference to estimate program parameters and hidden state from noisy observations, chooses informative experiments, revises the programs when predictions fail, and plans under uncertainty. It reports solving more tasks with fewer environment interactions than three baselines across five simulated domains, and demonstrates on a physical robot.
2026-09-30 · Shidu Ren, Qilin Gu, Zhenghao Ni et al. · arXiv:2609.35138
FlexiWorld is a JEPA-based latent world model for goal-directed planning that learns from mixed-span goal supervision while using variable-length action chunks across multiple time scales. It jointly trains a flexible (variable-length) causal action encoder, latent predictor, and an autoregressive goal-conditioned actor with Student Forcing. For planning, it proposes ARCEM (actor-residual CEM), combining action-residual search with within-chunk feedback and chunk-boundary latent prediction.
The authors introduce SACReg, a spectral anti-collapse regularizer meant to prevent dimensional collapse in the backbone representation (not just the projected space) by using a log-determinant penalty on representation covariance (and a mean penalty). They derive the idea from a λ-balance analysis in a two-layer linear network, extend it to nonlinear encoders, and apply it to JEPA as λ-JEPA with SACReg on both backbone and projector. They report improved classification/transfer and video SSL results, with released code.
2026-09-30 · Haowei Xu, Wanyi Fu, Hongbin Han et al. · arXiv:2609.35338
NeuronDiscover is an Agent-in-Twin framework for mechanistic discovery in neuronal microenvironments under “twin confounding” (mechanism change vs computational twin error). A shared mechanism-grounded World Action Model (WAM) drives prediction, intervention proposals, and observation design; independently adjudicated outcomes revise a scoped Mechanism–Intervention–Observation–Outcome (MIOY) graph, compiling supported relations into executable programs with discrepancy-adjusted acceptance bounds. Evaluated in simulated transport worlds and donor-disjoint current-clamp recordings.
2026-09-30 · Carlos Garrido-Munoz, Jorge Calvo-Zaragoza · arXiv:2609.35473
For Handwritten Text Recognition, the authors report discriminative signal concentrated in high-variance pixel directions. They test six self-supervised learning (SSL) methods across six handwriting benchmarks and find pixel-grounded masked image modeling (MAE, SimMIM) achieves the lowest CER in frozen and fine-tuned settings, and benefits from real-handwriting pretraining; JEPA-style and contrastive methods do not.
2026-09-30 · Yiqi Su, Rashed Shelim, Lingyi Wang et al. · arXiv:2609.35545
EpiMind is a graph world-model framework for multi-region epidemic policy planning under shared-resource constraints. It uses a graph-factored recurrent state-space model (GF-RSSM) to generate joint policy-conditioned rollouts, and a graph-temporal ADMM (GT-ADMM) planner to coordinate region interventions, enforce per-period feasibility via projection, and evaluate temporal specifications on projected actions.
2026-09-30 · Ziang Fu, Ning Ning · arXiv:2609.35603
The authors introduce Control-Geometry Straightening (CGS), a single auxiliary loss for joint-embedding latent world models. CGS uses only local pixel–action transitions to match pairwise cosine similarities among actions to those among corresponding latent differences, shaping “control geometry” for sampling-efficient latent planning. Under linear-dynamics, they analyze how this relates to temporal straightening, finite-budget guarantees for MPPI, and convergence results for gradient descent; experiments show large success-rate gains over LeWM and LeWM+TS.
2026-09-30 · Zongze Wu, Yani Guo, Runnan Li · arXiv:2609.35811
Lookahead-R reframes tool retrieval as a resource-constrained sequential decision problem. It trains an execution-aware surrogate world model to predict tool execution success, latency cost, and semantic utility without invoking real APIs, then uses a cost-sensitive, uncertainty-guided Monte Carlo Tree Search to select tools under latency budgets. On ToolBench I3, the authors report NDCG@5 = 91.40%, outperforming ToolGen (90.16%) by 1.24%.
2026-09-30 · Jie Yang, Jiajun Chen, Jiazheng Zhou et al. · arXiv:2609.35916
VehicleArena is a high-fidelity 3D urban-driving benchmark for independently operating multi-agent driving with passenger requests and cockpit/cabin actions in a single closed loop. It provides 112 held-out evaluation tasks (80 single-agent, 32 multi-agent). The authors report that the highest arrival rates across nine evaluated models are only 65.0% (single-agent) and 65.6% (multi-agent), and driving policies can reduce surrounding vehicles’ arrival rates versus SUMO’s native traffic controller.
2026-09-30 · Yizheng Huang, Wensheng Lin, Lixin Li et al. · arXiv:2609.35936
This tutorial proposes embodied semantic communication (ESC) for collective autonomous agents: an explicit wireless link carries action-oriented embodied semantics (multimodal perceptual states, intrinsic kinematic/body capabilities, and collaborative intents) so heterogeneous receivers can parse, align, and ground them in local motor control within a closed-loop perception–action cycle. The paper also outlines system characteristics, environment-constrained pathways, supporting mathematical tools, and open challenges.
The authors review in-context learning (ICL) for robots, where demonstrations and interaction change deployed behavior without updating neural parameters during deployment. They organize prior work into four interfaces connecting context to execution: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution, analyzing how training, correspondence, and memory determine whether taught requirements transfer under changing objects, environments, and execution conditions.
2026-09-30 · He Zhu, Lusen Zhao, Kwan Man Cheng et al. · arXiv:2609.36171
SkillWeaver is an agentic framework for scalable robot data generation. Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and parameterizes Neural Interaction Skills (NIS)—reusable closed-loop policies for contact-rich manipulation—and uses verifier-guided tree search with reflection and experience memory to discover and verify long-horizon behaviors. It scales to 39.1K demonstrations across 14.1K scenes and distills them into visuomotor policies for generalization and sim-to-real transfer.
2026-09-30 · Shitong Wang, Zhongang Cai, Yuzhou Hong · arXiv:2609.36227
The authors argue that next-latent (one-step) prediction is not, by itself, a world model. They show that one-step regression identifies only a conditional mean (a point), while multi-step open-loop error depends on the residual innovation covariance and grows with horizon after perfect one-step fit. They analyze cases with nonlinear means and non-injective observations, and show that an isotropy penalty depends only on the embedding marginal and has zero transition gradient.
2026-09-30 · Antonio Pariente, Ignacio Boero, Nikolai Matni et al. · arXiv:2609.36305
The authors propose a JEPA-style world model where latent dynamics are restricted to a bilinear (then normalized) parameterization, enabling efficient gradient-based planning and structurally enforcing action recoverability to prevent representation collapse. They report that, on standard 2D/3D control tasks, bilinear-parameterized representations match or improve success while reducing planning time by nearly three orders of magnitude.
2026-09-30 · Anna Pavlenko, Bogdan Crivat, Brandon Haynes et al. · arXiv:2609.36323
The authors propose an AI Software Factory for Data Systems that automates SDLC stages (Targeting, Coding, Reviewing, Ops) using a World Model and an evolutionary coding system (Darwin) wrapped in a self-improvement loop. They report scaled Microsoft deployments (tens of repositories) yielding 3× engineering efficiency vs agentic coding and up to 22× token efficiency, plus OSS and production evidence.
2026-09-30 · Ke Fang, Yupu Yao, Lu Cheng · arXiv:2609.36333
ATLAS addresses a gap in latent world-model planning: regularizing only the latent marginal (anti-collapse) does not guarantee preservation of the relational geometry needed for goal-conditioned action selection. ATLAS transfers normalized pairwise structure from an informative encoder representation to the planning latent and calibrates the planning latent’s marginal using Wasserstein embedding matching (WEMReg). Instantiated in LeWM, it improves mean goal-reaching success, especially on higher-novelty TwoRoom episodes.
2026-09-30 · Iishaan Inabathini, Margaret M. Henderson · arXiv:2609.36366
The authors extend cross-attention fMRI encoding models from static images to naturalistic video. They use per-parcel joint spatiotemporal cross-attention (joint over space and time) over V-JEPA-2 video tokens, fitting to BOLD Responses from short video clips. Joint routing yields higher held-out prediction accuracy across higher visual regions and produces interpretable attention maps that track moving objects and differ across category-selective networks.