Robotics
Action-conditioned world models for robot control. (9 papers)
2026-09-30 · He Zhu, Lusen Zhao, Kwan Man Cheng et al. · arXiv:2609.36171
SkillWeaver is an agentic framework for scalable robot data generation. Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and parameterizes Neural Interaction Skills (NIS)—reusable closed-loop policies for contact-rich manipulation—and uses verifier-guided tree search with reflection and experience memory to discover and verify long-horizon behaviors. It scales to 39.1K demonstrations across 14.1K scenes and distills them into visuomotor policies for generalization and sim-to-real transfer.
Roboticsauto-summary
2026-09-30 · Bang Du, Yichen Xie, Shuqi Zhao et al. · arXiv:2609.36413
The authors propose RoboActualizer to turn a pretrained video world model into a robot policy by “actualizing” a single task-conditioned future. A frozen V-JEPA 2.1 encoder provides a prior over plausible future latents; a tiny flow-matching head with two DiT experts learns task-conditioned selection and action realization. The approach is trained on cached latents and deployed with low latency for real-time control.
Roboticsauto-summary
2026-09-30 · Jade Choghari, Pepijn Kooijmans, Mansi Agarwal et al. · arXiv:2609.36416
FineART is a densely annotated bimanual manipulation dataset (40,543 episodes, 1,718 hours, 533,913 subtasks across 151 tasks). The authors introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask and executes continuous actions. Dense subtask training improves success (32.0%→100.0% on spatial disambiguation; 16.0%→76.0% on a long-horizon task). Mid-training enables data-efficient cross-robot transfer, and the authors open-source dataset, model weights, and training code.
Roboticsauto-summary
2026-09-30 · Wenliang Guo, Zhanbo Huang, Yu Kong · arXiv:2609.36454
DynamicHOI reconstructs hand-object interaction (HOI) trajectories from monocular RGB video using geometry-grounded diffusion refinement plus coupled hand-object physics. It derives hand generalized forces via articulated inverse dynamics and object wrenches via Newton–Euler dynamics, couples them through contact-force transfer, and recovers active hand actuation regularized by a probabilistic prior to suppress mechanically implausible motion.
Roboticsauto-summary
2026-09-30 · Sen Wang, Liu Liu, Xinjiang Wang et al. · arXiv:2609.37721
CogWAM is a cognition-guided world-action model for language-conditioned long-horizon robot manipulation. It maintains a persistent Semantic State that stores completed task events and the active subtask, updating only on semantic transitions. Progress-conditioned WORLD and ACTION queries use this state to condition future-world prediction (training only) and action generation (inference), aiming to keep predictions and actions aligned with task progress.
Roboticscodeauto-summary
2026-09-30 · Wenbo Chen, Tianfu Li, Haoxuan Xu et al. · arXiv:2609.37793
MVG-WAM is a world–action model for robotic manipulation that represents synchronized multi-view observations as geometrically related projections of one physical world. It combines an epipolar-constrained global state with view-indexed geometric states routed to corresponding video regions, and uses multi-horizon future metric-depth supervision for scale grounding without depth decoding at deployment. The authors report average success rates of 99.1% (LIBERO) and 92.07% (RoboTwin 2.0), plus 91.3% success on Cobot Magic.
Roboticsauto-summary
2026-09-30 · Sicheng Xie, Yitong Chen, Haidong Cao et al. · arXiv:2609.37810
The authors introduce RoboSkill, a skill acquisition-and-reuse framework organized as an Explore–Execute–Evolve loop for embodied agents. The agent explores to gather missing task information, executes tasks with visual and tactile feedback while running reusable code, then evolves a skill library from session records. On LIBERO-10, they report higher first-episode success and lower runtime, and similar gains on real robots.
Roboticscodeauto-summary
2026-09-30 · Shenghe Zheng, Wenbo Li, Jiyao Zhang et al. · arXiv:2609.38059
WorldLine is an action-driven visual simulator for robotic manipulation. It learns transferable robot–object dynamics from 10,000+ hours of action-free robot videos, then grounds heterogeneous actions using 2,000+ hours of action trajectories across 10+ embodiments via an image-space action interface and multi-view/failure-enriched relational training. A robot-focused few-step distillation enables efficient causal rollout for policy evaluation and planning.
Roboticsauto-summary
2022-06-28 · Philipp Wu, Alejandro Escontrela, Danijar Hafner et al. · arXiv:2206.14176
DayDreamer applies the Dreamer algorithm to four physical robots that learn online in the real world without simulators, using the same hyperparameters across tasks: a quadruped, two robot arms and a wheeled robot.
Roboticsauto-summary