Pub-AI: AI in Science

Video World Models · 2026-09-30

Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders

Jiaming Zhang, Wu Yang, Shuai Tao, Wulong Liu

arXiv:2609.31938PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

Cache-aware Conv3D lowering for embedded generative video VAEs: supported causal Conv3D calls are expressed as batched spatial Conv2D while preserving pretrained weights, causal-cache semantics, convolution parameters, bias placement, and output layout. On 64-GB Jetson AGX Orin in the Cosmos3-Edge image-to-video pipeline, the authors report ~7.32× VAE-decoder speedup and 2.21× end-to-end speedup at 25 frames.

Why it matters

The paper targets runtime efficiency for fixed pretrained causal video decoders on edge hardware, separating learned-model behavior from execution-operator representation. By preserving cache semantics and using guarded fallback for unsupported cases, the method aims to accelerate VAE decoding and maintain fast-path coverage inside complete pipelines, and to transfer across a distinct Wan-family VAE architecture.

Method

  • Lower supported cached causal Conv3D temporal taps to batched spatial Conv2D operations, preserving pretrained weights, temporal-cache ordering/semantics, convolution parameters, bias placement, and output layout.
  • Use guarded runtime procedure: unsupported cases return to original Conv3D route; strict mode terminates validation when unsupported or execution errors occur; route counters expose coverage/fallbacks.
  • Evaluate on Cosmos3-Edge complete I2V pipeline on Jetson AGX Orin, compare against native and full-coverage TensorRT, and test decoder-only transfer to LingBot-World’s Wan2.1 VAE.

Limitation

The authors do not evaluate other cache APIs, Conv3D classes, hardware platforms, resolutions, schedulers, and software stacks. They also note TensorRT is 1.36× faster on the clean 17-frame workload and do not exhaustively evaluate other compiler-/vendor-/manual Conv3D implementations.

Abstract (from arXiv)

Generative world models can provide visual rollouts for embodied planning, yet their feasibility on edge devices depends not only on the learned model but also on how the execution runtime represents its operations. We introduce a cache-aware lowering that expresses supported causal Conv3D calls as batched spatial Conv2D operations while preserving pretrained weights, temporal-cache semantics, convolution parameters, bias placement, and output layout. Across the complete Cosmos3-Edge image-to-video pipeline on a 64-GB NVIDIA Jetson AGX Orin, the proposed route accelerates VAE decoding by approximately $7\times$ and reduces complete-generation latency by more than $2\times$, while repeated decoder evaluations maintain complete fast-path coverage without fallbacks. The unchanged lowering also improves Cosmos3-Nano and transfers to LingBot-World's architecturally distinct Wan2.1 VAE. A clean-device comparison against fully specialized TensorRT shows that TensorRT provides a further $1.36\times$ steady-state improvement, but requires substantially greater per-module and per-runtime-state AOT specialization. Same-latent BF16 and FP32 evaluations characterize the finite-precision differences introduced by the alternative execution order. Together, these results position cache-aware lowering as a lightweight runtime optimization that recovers most of the available decoder acceleration without modifying the learned models themselves.

Related papers