Driving · 2026-09-30
HelixWorld: A Real-time Interactive Audio-Visual World Model
Lei Ke, Jiahao Pan, Zeyue Tian, Jiaming Wang, Haoyuan Huang, Kam Man Wu, Pengjun Fang, Hongyu Liu, Chenyang Qi, Lin Wang, Ruibin Yuan, Weijia Chen, Fangneng Zhan, Qifeng Chen, Wei Xue, Yike Guo
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
HelixWorld is a real-time interactive audio-visual world model where visual scenes and camera-grounded stereo sound co-evolve under user interaction. It trains a bidirectional teacher on a curated high-fidelity stereo spatial audio-visual dataset with metric camera poses, then distills it into a few-step causal streaming student using online trajectory distillation and streaming fine-tuning for 24 FPS on a single NVIDIA H800 GPU.
Why it matters
The authors address interactive world models that are “strictly silent” by coupling continuous 6-DoF camera trajectories and actions to synchronized visual generation and spatial stereo audio. They also introduce HelixBench to evaluate whether synthesized sound fields physically track dynamic camera motion.
Method
- Curate a spatial audio-visual dataset with “true stereo acoustics and metric camera poses,” and discretize motions into an 81-class action vocabulary; use decoupled three-track captions (video-only, audio-only, joint audio-visual).
- Train a bidirectional teacher conditioned on continuous 6-DoF camera trajectories (PRoPE) and discrete actions (AdaLN-Zero), using cross-modal attention to bind visual geometry with acoustic features.
- Distill the teacher into a few-step causal student with causal initialization, online trajectory distillation loss (PF-ODE targets), and long-horizon streaming fine-tuning for drift-free joint audio-visual rollouts at 24 FPS.
Abstract (from arXiv)
World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.