Latent Dynamics · 2026-09-30
Cross-attention encoding models reveal dynamic spatiotemporal routing across human higher visual cortex
Iishaan Inabathini, Margaret M. Henderson
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
The authors extend cross-attention fMRI encoding models from static images to naturalistic video. They use per-parcel joint spatiotemporal cross-attention (joint over space and time) over V-JEPA-2 video tokens, fitting to BOLD Responses from short video clips. Joint routing yields higher held-out prediction accuracy across higher visual regions and produces interpretable attention maps that track moving objects and differ across category-selective networks.
Why it matters
Video-evoked neural responses depend on which spatial locations and temporal moments are most predictive, but many encoding models use fixed/pooled or factorized attention. By testing spatial, temporal, factorized, and joint spatiotemporal routing with shared training/readout settings, the authors identify where coupled routing improves predictions and provide stimulus-specific attention maps that can be compared to functional systems (e.g., lateral/face/body vs scene vs early visual).
Method
- Use a frozen self-supervised video model (V-JEPA-2 ViT-g/16) to map each clip to a grid of spatiotemporal tokens; learn a parcel-specific query that attends to tokens to produce a stimulus-dependent attended feature.
- Fit the cross-attention readout to fMRI responses from the BOLD Moments Dataset using a voxel-specific linear readout; compare joint spatiotemporal attention against spatial-only, temporal-only, factorized, mean-pool, and ridge regression baselines.
- Analyze interpretability by inspecting attention maps over space-time and localizing category-selective networks via attention-token IoU contrasts.
Abstract (from arXiv)
Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel responses to complex natural videos. However, the majority of video-computable encoding models predict responses using simple linear mappings from model tokens, overlooking the spatiotemporal structure shared by video representations and neural responses. Recent cross-attention encoding models address this limitation for static images, enabling flexible stimulus-dependent weighting of image content across space. Here, we extend this framework to naturalistic video, using per-parcel cross-attention to dynamically route features from a self-supervised video model (V-JEPA-2) across both space and time, fitting this model to fMRI responses to short video clips. We compare joint spatiotemporal attention with factorized and selectively constrained alternatives, and find that joint routing improves predictions of brain responses to held-out videos across higher visual regions, most consistently in lateral and dorsal visual areas associated with dynamic motion perception. Moreover, our method provides interpretable, stimulus-specific attention maps that dynamically follow moving objects, revealing which locations and temporal moments contribute to each neural response. We further show that attention maps from parcels in different category-selective networks (face-, body-, scene-selective) differentially weight content in accordance with expected semantic selectivity. Together, this work provides a new computational framework for understanding how visual information is adaptively weighted by cortical populations during dynamic visual perception.
Related papers
- MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
- Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models
- Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning
- RoboChrono: A Real Robot Benchmark for Streaming Task Understanding