Driving · 2026-09-30
Beyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation Learning
Xuanlin Chen, Ziyue Wang, Xunlan Zhou, Yuan-yih Shang, Qiang Wu, Shenghua Wan
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
RECON improves model-based imitation learning by separating conservative policy learning from active real-environment data collection. It trains a main (conservative) policy for task execution and an explorer optimized using epistemic uncertainty conditioned on recoverability estimated from multi-step imagination with the main policy, prioritizing uncertain yet recoverable dynamics. Experiments across locomotion, navigation, and manipulation show consistent gains in interaction efficiency, imitation performance, and robustness.
Why it matters
The authors argue conservative MBIL controls model usage but does not address the recovery-relevant acquisition blind spot: if data are collected using the same conservative policy, uncertain recovery regions around the expert distribution may be underexplored. RECON targets those underexplored recovery regions by gating exploration with recoverability predicted through main-policy imagination, aiming to improve both imitation and the learned world model’s coverage for deployment.
Method
- Separate conservative main-policy learning (deployed at test) from explorer-based real data collection; both share the same world model, representation, discriminator, and replay.
- Estimate recoverability using multi-step Dreamer imagination: roll the explorer endpoint forward with the frozen main policy and score whether it returns toward expert-compatible behavior using discriminator-based compatibility.
- Optimize the explorer using uncertainty rewards gated by recoverability, while keeping explorer-only exploration effects during training (evaluation uses only the deterministic main policy).
Limitation
The paper notes that the recovery criterion motivates targeted (not global) model accuracy, since the residual term is unavoidable unless the gate upper-bounds recovery occupancy everywhere.
Abstract (from arXiv)
Model-based imitation learning (MBIL) improves real-environment interaction efficiency by optimizing policies on imagined rollouts from a learned world model. However, the gap between model-induced and real-environment occupancies makes policy learning sensitive to model error. Conservative MBIL mitigates model exploitation during policy optimization, but when real-environment interactions are collected by the same conservative policy, uncertain regions around the expert distribution remain insufficiently sampled. Generic uncertainty-driven exploration, on the other hand, may allocate interaction to novel but task-irrelevant dynamics. We propose REcoverability-CONditioned Exploration for Model-Based Imitation Learning (RECON). RECON separates conservative policy learning from active data collection by maintaining a main policy for task execution and an explorer for real-environment interaction. The explorer is optimized based on epistemic uncertainty conditioned on recoverability estimated from multi-step main-policy imagination, focusing data collection on unknown states from which the main policy can still return toward expert behavior. Experiments on locomotion, navigation and manipulation show consistent gains in interaction efficiency, imitation performance, and robustness, indicating that RECON directs real-environment interaction toward recovery regions around the expert distribution that are underexplored by prior methods, and thereby learns a world model better suited for imitation.
Related papers
- Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation
- SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation
- Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents
- Copper-Policy: Focus on the Representation for Robust Robot Manipulation
- Shaping Persistent Representations from Independent Interactions