Latent Dynamics · 2026-09-30
Self-Confirming Superposition Traps in Reinforcement Learning
Dai Shi, Andi Han, Feng Chen, Yiqun Duan, Junbin Gao, Jos\'e Miguel Hern\'andez-Lobato
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
The authors show that RL’s representation-fitting loop can sustain a lower-return policy: in a self-confirming superposition trap, globally optimal codes under the current policy share overlapping directions for features that rarely co-activate, yet interfere after an alternative action. Interventions that preserve access to neglected states or protect replay weights reduce interference and can improve control, including in DreamerV3–Crafter.
Why it matters
The paper identifies a mechanism where “globally optimal” representation fitting on policy-collected data does not preserve return ranking across actions, due to interference from shared capacity. The authors bound when ranking can stabilize a worse policy and derive a replay condition. The experiments then test mitigation via state access, replay reweighting, and overlap/entropy interventions, including for world-model agents.
Method
- The authors define self-confirming superposition traps in linear tied two-step task models, characterizing allocation (nonorthogonal encoder directions), reversal of return ranking, and local stability of the lower-return policy under exact adaptation.
- They analyze robustness and reversal distortion with finite-action models, deriving a bound and a sufficient replay condition that preserves the return advantage of the best separately adapted action under residual representation error.
- They test interventions in neural PPO (matched-feature data and feedback learning), state/replay access interventions on MiniGrid and DMControl, and protected world-model fitting in DreamerV3–Crafter.
Limitation
The authors state that the complete trap guarantees apply to the stated feature models and actor dynamics, and the finite-action envelope analysis does not transfer automatically to a general MDP parameterized by neural policy weights; additionally, changing critics, finite adaptation, and imperfect optimization can introduce additional dynamics.
Abstract (from arXiv)
Reinforcement learning (RL) trains representations on data selected by the agent's policy, which then uses the resulting returns to guide its next choices. We show that this loop can sustain a lower-return policy even when representation fitting is globally optimal on those data. In a self-confirming superposition trap, every optimal code assigns overlapping directions to features that rarely occur together under the current policy. An alternative action brings them together, causing interference that lowers its return and reinforces avoidance, although refitting to that action would yield more return at the same capacity. We characterize the dimensions admitting a trap in a tied two-step model and show separately that equal feature frequencies, continued visitation, and independent controller learning need not prevent it. Because fitting weights errors by visitation, an avoided action can lose its return advantage at little cost to the objective. In a finite-action model, we bound this distortion and derive a replay condition: sufficient training weight on the best separately adapted action preserves its ranking despite residual error. Neural PPO experiments show how the feedback develops during learning: agents initialized toward different actions develop different interference patterns, opposite mean return rankings, and different final policies at the same capacity. We therefore test whether retaining access to neglected states can improve control. Keeping these states in training reduces measured interference and improves sequential return, with gains even when the encoder is frozen. Related interventions on state access, replay weights, and feature overlap improve control on MiniGrid and DMControl. For agents that learn through a world model, protected fitting improves DreamerV3--Crafter's cumulative training scores at unchanged capacity.
Related papers
- Beyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation Learning
- Shaping Persistent Representations from Independent Interactions
- Don't Throw Away the Tail: Action Upcycling for Policy Acceleration
- SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation
- Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents