Pub-AI: AI in Science

Video World Models · 2026-09-30

Honeycomb: Constant-Size Scene Memory Representation for Video World Models

Jack Wei Lun Shi, Kaichen Zhou, Haoyu Chen, Yufeng Weng, Keane Ong, Ruojin Cai, Hang Hua, Justin K. W. Yeoh, Mengyu Wang

arXiv:2609.37690PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

Honeycomb is a video world model that keeps persistent scene memory in a fixed-size HexMemory. HexMemory stores scene features as six fixed-dimension low-rank planes (three spatial, three spatiotemporal). A feed-forward writer updates only new video chunks by warping prior planes to expanded bounds and fusing with confidence-weighted pooling plus a learned residual. A reader reconstructs latent features from HexMemory to condition generation; the method avoids per-scene optimization and full-history reprocessing.

Why it matters

The authors address long-horizon video generation where earlier observations fall outside the model’s temporal context, hurting revisit consistency. Honeycomb’s fixed-size memory aims to preserve scene layout and appearance over revisits while keeping feature storage constant during generation, improving generation quality and revisit fidelity on WorldScore and RealEstate10K.

Method

  • Represent persistent scene features with HexMemory: six fixed-size feature planes (three spatial and three spatiotemporal) in a low-rank plane factorization updated over time.
  • Use a feed-forward writer that maps only the latest chunk’s latent observations into plane features; when bounds expand, warp previous planes to new bounds (preserving grid dimensions) and fuse via confidence-weighted pooling plus learned residual correction.
  • At each chunk, a memory reader reconstructs latent feature maps from HexMemory (with visibility masks) to condition a diffusion transformer, then write-back updates HexMemory from generated chunks.

Limitation

When evaluated without dynamic object filtering during memory writes, Honeycomb outperforms Spatia and LSM-World on overall WorldScore.

Abstract (from arXiv)

Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memory systems accumulate RGB observations or latent features, causing storage requirements to grow as generation proceeds. We introduce **Honeycomb**, a video world model built on **HexMemory**, a compact low-rank representation that stores scene features in a fixed-size memory comprising six spatial and spatiotemporal planes. A feed-forward writer maps each newly generated video chunk to plane features. As the spatial coverage or temporal range expands, HexMemory warps the existing planes while preserving their dimensions, then integrates new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latent features from HexMemory to condition subsequent video generation. Because the writer processes only observations from the latest chunk, Honeycomb avoids per-scene optimization and repeated processing of the full generation history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust consistency when revisiting previously observed regions, while maintaining constant feature-storage requirements throughout generation. Code and additional visualizations are available on our https://jackswl.github.io/honeycomb/.

Related papers