Driving · 2026-09-30
T$^2$Mem: Learning Test-Time Memory for Robotics
Yize Liu, Huang Huang, Yining Hong, Zijian Du, Zhi Cao, Li Fei-Fei, Jiajun Wu
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
T²Mem is a test-time training framework that turns a memory-free pretrained vision-language-action policy into a robotics controller with fast-weight test-time memory. It uses an observation-grounded interface to encode vision-language history into compact fast weights via online self-supervised updates, and action supervision plus alternating memory–policy learning shapes what is stored and how it is used. On RoboMME (16 tasks), it improves average success from 17.93% to 56.83% over the memory-free base policy.
Why it matters
The authors target decision-relevant memory for robotics when the current observation lacks information needed for future actions. They aim to learn this capability from action demonstrations alone, without external reasoning models or memory-specific annotations, by integrating memory formation and its use into a single pretrained VLA and adapting the memory at test time through online self-supervised updates.
Method
- Integrates adaptive fast-weight memory into a single pretrained vision-language-action policy using a VLM-to-action-expert interface; memory adapts within an episode while policy slow parameters remain fixed.
- Builds memory with online test-time training: observation-grounded self-supervised updates write to fast weights and reads precede writes to condition subsequent actions.
- Trains with alternating memory–policy learning so action supervision shapes memory formation and the policy learns to use retrieved historical context.
Limitation
The paper states that swap tasks remain lower when the policy must track swaps and update object–location associations rather than simply recall a stored binding, suggesting limitations in the temporal reasoning of the base policy; additionally, low-level control can limit tasks such as InsertPeg.
Abstract (from arXiv)
Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use historical information from action demonstrations alone, without external memory support. We introduce T$^2$Mem, a framework that develops this capability within a pretrained vision-language-action policy, without external reasoning models or memory-specific annotations. T$^2$Mem uses test-time training to encode observation history into compact fast weights through online self-supervised updates, avoiding repeated processing of the full history. An observation-grounded interface extracts vision-language information for memory formation and supplies retrieved context to the action expert. Action supervision shapes what the memory learns to retain and use, while alternating memory-policy learning gives each component a fixed counterpart during optimization. Across 16 RoboMME tasks, T$^2$Mem improves average success from 17.93% to 56.83% over the memory-free base policy and outperforms the recurrent-memory methods reported in the benchmark, while controlled profiling indicates at least 3x inference speedup over explicit methods. Project website: https://yzliu84.github.io/T2MEM-project/
Related papers
- Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning
- RoboChrono: A Real Robot Benchmark for Streaming Task Understanding
- Copper-Policy: Focus on the Representation for Robust Robot Manipulation
- FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation
- RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation