Pub-AI: AI in Science

Video World Models · 2026-09-30

MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception

Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu

arXiv:2609.33804PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

MinkowskiPE proposes Minkowski positional encoding for spatiotemporal perception. It assigns each token a spacetime coordinate and uses the relative spacetime displacement between tokens to parameterize Lorentz transformations applied to query/key features, making attention depend on position only through relative displacement and be translation-invariant. It retains the standard dot-product attention interface and is compatible with efficient attention implementations.

Why it matters

The authors target the challenge of explicitly modeling spatiotemporal coupling while keeping learning-based flexibility. MinkowskiPE builds inductive bias at the positional-representation level by using Minkowski geometry and Lorentz-transform parameterizations so that query–key attention depends only on relative spacetime displacement, not global coordinate translation.

Method

  • Represent spatiotemporal tokens as events x_i=(t_i,r_i) and define relative displacement Δx_ij=x_j−x_i.
  • Use joint temporal/spatial coordinates to parameterize one-parameter Lorentz transformations that are applied to query and key features (value features unchanged).
  • Multi-head attention via partitioning features into 2D blocks; a shared projection matrix maps event coordinates to scalar parameters ρ_i, yielding translation-invariant pairwise attention dependence on Δx_ij.

Limitation

The authors note that more expressive multi-parameter Lorentzian positional encodings would require additional choices of commuting subalgebras and are left for future work.

Abstract (from arXiv)

Modeling spatiotemporal coupling is a key challenge in building physical intelligence across scales, from microscopic to macroscopic. Existing models capture such structure broadly through physics-motivated dynamical formulations or learning-motivated architectures. The former provide stronger priors but may constrain flexibility, whereas the latter are more flexible but leave the spatiotemporal coupling largely implicit. We therefore seek an approach that combines flexible learning with an explicit geometric bias for jointly modeling time and space. To this end, we propose Minkowski Positional Encoding (MinkowskiPE), which uses joint temporal and spatial coordinates to parameterize Lorentz transformations applied to query and key features. With MinkowskiPE, the query-key attention score depends on position only through the relative spacetime displacement between the two tokens and is therefore invariant to global translation of the coordinates. This paradigm retains the standard dot-product attention interface and remains compatible with efficient attention implementations. We evaluate MinkowskiPE on microscopic molecular dynamics and macroscopic video prediction tasks, achieving the best results on all nine multi-trajectory molecular evaluations and reducing KTH video-prediction MSE by 9.9% relative to the best baseline while using roughly one-tenth as many parameters.

Related papers