Latent Dynamics · 2026-09-30
DSWM: Decomposed Spatio-Temporal World Model for Demand-Driven UAV Base Station Repositioning
Shengjie Zhong, Zhongliang Zhao, Jingxuan Chen, Xianbin Cao, Xinmei Qiang, Dapeng O. Wu, Tony Q. S. Quek
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
DSWM is a decomposed spatio-temporal world model for demand-driven UAV base station repositioning. It uses a recurrent state-space model trained with an EMA-based latent predictive objective plus variance regularization, a differentiable service simulator head (association, probabilistic LoS, Shannon rate, capped fulfillment), and observation-anchored CEM planning that executes only the first action of the best imagined rollout.
Why it matters
The authors frame UAV-BS repositioning as decision-time planning under partial, time-varying demand, arguing that the key bottleneck is how current observations are exploited at decision time rather than forecast accuracy. DSWM couples a latent dynamics model with a decomposed differentiable coverage/service physics simulator, then uses CEM over imagined rollouts anchored on the current observation.
Method
- Learn an RSSM world model of demand-service dynamics with a GRU deterministic state and diagonal-Gaussian latent state; train with an EMA-based latent predictive objective and variance-margin regularization (VICReg-style).
- Attach a differentiable, decomposed demand-service physics head that replays association, probabilistic LoS, Shannon rate chain, and capped per-cell fulfillment inside latent rollouts; penalize uncertainty via ensemble disagreement.
- Plan with observation-anchored CEM decision-time planning: imagine H-slot trajectories in the world model, anchor imagined demand on the current observation with mixing coefficient ρ=0.95, and execute only the first action.
Limitation
YJMob100K demand is a distinct-user presence proxy (5% sampling), not traffic bytes, so cross-dataset conclusions about absolute served ratios should be read with the declared semantics. Energy and fairness discriminate weakly under hover-dominated power and battery constraints (∼3% energy spread, Jain 0.888–1.000), so claims rest on served ratio.
Abstract (from arXiv)
Uncrewed aerial vehicle base stations (UAV-BSs) are expected to cover traffic demand that shifts across space and time, yet most repositioning schemes either re-solve an optimization problem per slot or learn reactive policies without an explicit demand model. We cast demand-driven fleet repositioning as latent-space decision-time planning and propose DSWM, a decomposed spatio-temporal world model: an agentic controller that perceives the demand field through a rolling observation window, retains operational context in a latent recurrent state, reasons about candidate motions by imagined rollouts under an uncertainty penalty, and coordinates the fleet through replanned first actions. DSWM learns a recurrent state-space model shaped by an exponential-moving-average (EMA) based latent predictive objective with variance regularization. It attaches a differentiable service simulator that replays the association, probabilistic line-of-sight channel, and Shannon rate chain inside latent rollouts. Planning uses a cross-entropy method whose imagined demand is anchored on the current observation window with mixing coefficient $\rho=0.95$. On a unified pipeline over three real datasets (Milan CDR (call detail record), Shanghai Telecom, YJMob100K) and 14 methods including five reproduced IEEE baselines, DSWM attains weekday served ratios of 0.889, 0.908, and 0.898, ranking first among non-ablated configurations on every dataset. On Milan it improves over the strongest non-learning baseline (Greedy, 0.780) by 0.109, a margin that comes from decision-time use of observations rather than prediction accuracy.
Related papers
- What Do Latent Predictive Vehicle Representations Retain? Measuring State, Geometry, and Local Response
- V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving
- Correct then Forecast: Observer State-Space Models for Time Series Forecasting
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
- Learning Latent Dynamics for Planning from Pixels