Video World Models · 2026-09-30
Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang, Yizhi Wang, Xinwei Huang, Minxuan Lin, Angtian Wang, Chongyang Ma, Fan Tang
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
The authors argue that token-wise load-balanced MoE routing forms a “uniformity trap” for video, scattering coherent spatiotemporal tokens across experts and causing fragmentation and distortion. They propose SplitMoE, splitting experts into semantic and generic roles with prototype-guided semantic routing and pull-push regularization, reporting faster convergence and improved routing coherence and video generation quality under matched activated-parameter budgets.
Why it matters
If MoE routing is misaligned with video’s spatiotemporal redundancy and long-tailed semantics, scaling can hurt generation. SplitMoE targets routing coherence by separating semantic abstraction from residual visual modeling, and the paper provides benchmark comparisons plus ablations on convergence and routing pathology, relevant to building large-scale video world models with sparse compute.
Method
- SplitMoE partitions MoE experts into semantic experts and generic experts; routing uses independent sigmoid scores per group with fixed Top-K active experts matching the standard baseline budget.
- Prototype-Guided Semantic Routing uses learnable prototypes optimized with router alignment (KL to VAE-feature induced soft targets) and prototype pull-push regularization in VAE feature space.
- The overall training objective combines flow matching with alignment and prototype losses; semantic experts use semantic alignment while generic experts use a loss-free load-balancing strategy for routing.
Limitation
Limitations include reliance on frozen visual features and sparse-routing overhead, motivating future work on adaptive partitioning, efficient distributed routing, and broader video generation tasks.
Abstract (from arXiv)
Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.
Related papers
- LongLive-Plug: Once-for-All Distillation for Video Generation
- Honeycomb: Constant-Size Scene Memory Representation for Video World Models
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
- World2Motion: Turning Video World Models into 3D Human Motion Generators
- Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning