Pub-AI: AI in Science

Video World Models · 2026-09-30

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang, Yizhi Wang, Xinwei Huang, Minxuan Lin, Angtian Wang, Chongyang Ma, Fan Tang

arXiv:2609.38140PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors argue that token-wise load-balanced MoE routing forms a “uniformity trap” for video, scattering coherent spatiotemporal tokens across experts and causing fragmentation and distortion. They propose SplitMoE, splitting experts into semantic and generic roles with prototype-guided semantic routing and pull-push regularization, reporting faster convergence and improved routing coherence and video generation quality under matched activated-parameter budgets.

Why it matters

If MoE routing is misaligned with video’s spatiotemporal redundancy and long-tailed semantics, scaling can hurt generation. SplitMoE targets routing coherence by separating semantic abstraction from residual visual modeling, and the paper provides benchmark comparisons plus ablations on convergence and routing pathology, relevant to building large-scale video world models with sparse compute.

Method

  • SplitMoE partitions MoE experts into semantic experts and generic experts; routing uses independent sigmoid scores per group with fixed Top-K active experts matching the standard baseline budget.
  • Prototype-Guided Semantic Routing uses learnable prototypes optimized with router alignment (KL to VAE-feature induced soft targets) and prototype pull-push regularization in VAE feature space.
  • The overall training objective combines flow matching with alignment and prototype losses; semantic experts use semantic alignment while generic experts use a loss-free load-balancing strategy for routing.

Limitation

Limitations include reliance on frozen visual features and sparse-routing overhead, motivating future work on adaptive partitioning, efficient distributed routing, and broader video generation tasks.

Abstract (from arXiv)

Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.

Related papers