Video World Models · 2026-09-30
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
LongLive-Plug introduces a once-for-all distillation framework that learns reusable LoRA adapters on a base video diffusion model. The adapters cover single-pass classifier-free guidance, few-step (four-step) sampling, and long-context error correction for causal autoregressive generation, enabling training-free plug-and-play deployment across compatible downstream models.
Why it matters
The authors address repeated per-target distillation costs by distilling reusable capabilities once per backbone family and reusing them without per-target retraining. They also decouple CFG and few-step behavior so guidance strength can be controlled at inference while preserving few-step generation.
Method
- Train functional LoRAs once on a frozen base model for reusable capabilities: CFG distillation (single-pass), few-step distillation (four-step sampling), and long-context distillation (error correction for causal AR rollouts).
- Decouple CFG control from few-step correction: learn separate CFG-only and few-step LoRAs so the CFG LoRA inference weight acts as a guidance dial while the few-step weight stays fixed.
- Deploy adapters without downstream training by merging LoRA updates into compatible descendant model weights; adapters transfer even if downstream models add conditioning branches or expand output channels.
Limitation
Once-for-all reuse requires compatible descendants of each base model. Long-context transfer requires existing causal AR inference, since LoRA updates alone do not change attention masks. CFG LoRA weights provide approximate guidance control and may require adjustment after transfer.
Abstract (from arXiv)
Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, our dedicated CFG LoRA provides text guidance control through its inference weight. Combining it with a few-step LoRA simultaneously preserves few-step generation and CFG controllability on downstream tasks. We verify training-free deployment on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation. The approach may support additional compatible models. Each capability can thus be distilled once per backbone family and reused without per-target retraining.
Related papers
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
- Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
- World2Motion: Turning Video World Models into 3D Human Motion Generators
- EVO-WAM: Evolving World Action Models through Video-Action Verification
- Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning