Driving · 2026-09-30
InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning
Hongpei Zheng, Hujun Yin
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
InsightMap is a framework that uses top-down maps as explicit spatial memory and action-conditioned prediction targets. It links historical views to labeled map locations and trains a shared multimodal backbone (based on BAGEL) to jointly predict navigation actions and post-action map generation, with QA and 3D grounding sharing the same spatial interface. The authors report improved navigation and competitive static spatial reasoning and grounding results.
Why it matters
The authors argue that action-label supervision alone does not explicitly train how a policy updates its map after actions. InsightMap uses action-conditioned post-action map prediction as auxiliary spatial supervision while supporting a unified spatial interface for navigation, visual question answering, situated reasoning, and 3D grounding via an aligned RGB-D pipeline.
Method
- Uses top-down maps as persistent spatial memory, linking historical first-person views to labeled map locations.
- Trains a shared multimodal backbone to predict navigation actions and action-conditioned post-action map generation from the observed spatial context (future-map generation used as auxiliary training supervision).
- Shares the same spatial interface for static scene reasoning (QA) and 3D grounding, using task-specific text outputs with map-based spatial context.
Abstract (from arXiv)
Language-guided navigation requires connecting partial observations to a persistent spatial reference and learning how actions change that representation. We introduce InsightMap, a framework that uses top-down maps as both explicit spatial memory and action-conditioned prediction targets. Historical views are linked to labeled map locations, and a shared multimodal backbone jointly learns navigation action prediction and post-action map generation. Map prediction provides auxiliary training supervision, while navigation inference decodes actions from the observed spatial context. An aligned RGB-D data pipeline supports a common interface for navigation, visual question answering, situated reasoning, and 3D grounding. On the validation-unseen splits of R2R-CE and RxR-CE, InsightMap achieves success rates (SR) of 56.9% and 54.9%, respectively. Adding map-prediction supervision improves R2R-CE SR by 4.3 and success weighted by path length (SPL) by 3.2 percentage points. On static spatial tasks, InsightMap achieves 103.7 CIDEr on ScanQA, 60.1% exact-match accuracy on SQA3D, and 53.1% grounding accuracy at 0.5 IoU on ScanRefer with detected object proposals. On Unitree Go2, it outperforms NaVid and NaVILA in hallway, lab, and office environments.
Related papers
- CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces
- RoboChrono: A Real Robot Benchmark for Streaming Task Understanding
- T$^2$Mem: Learning Test-Time Memory for Robotics
- FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents