Pub-AI: AI in Science

Video World Models · 2026-09-30

Dexterous Tactile World Model

Ziyao Zeng, Xiatao Sun, Hao Wang, Yueyang Pan, Zhengxiang Yu, Fengyu Yang, Tianyu Liu, Zhiwen Fan, Daniel Rakita

arXiv:2609.34286PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

DTWM is a video world model for future-frame prediction of egocentric human manipulation using both RGB video and dexterous tactile glove signals from each hand. It conditions a pretrained video diffusion transformer by injecting a zero-initialized tactile residual at each hand’s location in the video tokens with block-causal masking so predicted frames cannot use future info. Compared to a matched vision-only baseline, it reduces hand-motion underestimation (23% to 9%) and hand-region perceptual error.

Why it matters

Manipulation outcomes often depend on contact events that are hard to observe visually but easier to sense through touch. The authors’ approach resolves “indexing mismatch” by aligning tactile readings with the corresponding hand location in the video token grid, enabling future-frame prediction to better track where and how hands move, and to use touch history even when no future touch is available at inference.

Method

  • Predict future RGB frames (frames 13–48) from past frames, hand skeletons, and tactile glove readings from both hands; no future hand pose/action/touch is provided.
  • Condition a pretrained video diffusion transformer with a zero-initialized residual injected into video tokens at each hand’s projected location using a normalized Gaussian footprint; a causal mask blocks future-token access.
  • Train with a flow-matching loss over chunked, block-causal latent predictions; evaluate against a matched vision-only baseline where tactile readings are zeroed.

Limitation

The tactile signal is standardized and clipped to [−8, 8], with hand-centered injection based on rendered skeleton hand location; specifics of generalization beyond the evaluated setting are not stated in the provided text.

Abstract (from arXiv)

World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.

Related papers