Pub-AI: AI in Science

Robotics · 2026-09-30

Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents

Sicheng Xie, Yitong Chen, Haidong Cao, Shunlin Lu, Zuxuan Wu, Yu-Gang Jiang

arXiv:2609.37810PDFCode

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors introduce RoboSkill, a skill acquisition-and-reuse framework organized as an Explore–Execute–Evolve loop for embodied agents. The agent explores to gather missing task information, executes tasks with visual and tactile feedback while running reusable code, then evolves a skill library from session records. On LIBERO-10, they report higher first-episode success and lower runtime, and similar gains on real robots.

Why it matters

RoboSkill targets high execution costs in zero-shot robotic task solving by reusing skills across exploration and execution cycles. The loop explicitly connects physical interaction outcomes to evolving reusable skills, and the design choices (tactile feedback and executable code) are meant to reduce uncertainty and reasoning overhead during reuse. Reported improvements on LIBERO-10 and real robots support the value of accumulating and adapting skills rather than solving each task from scratch.

Method

  • Organize embodied behavior as an Explore–Execute–Evolve loop: explore to acquire missing information, execute with feedback, then consolidate session records and executed code into skills for subsequent cycles.
  • During execution, augment RGB/depth with tactile (pressure and force/torque) feedback; use an SDK so the agent can compose executable procedures (code) like depth point cloud conversion and table-height verification.
  • In evolution, the agent reviews accumulated records to create/update skill packages with textual guidance, reusable code, and supporting records; in reuse, it selects up to K=3 relevant skill packages based on metadata.
Abstract (from arXiv)

Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by reasoning and exploring the physical world from scratch. To reduce these costs, we introduce RoboSkill, a framework that connects skill acquisition and reuse through an Explore, Execute, Evolve loop. Within this loop, the agent explores to gather task-relevant information, executes tasks while adapting to feedback, and evolves its skill library based on execution records. It then reuses these skills to guide exploration and execution in the next cycle, closing the loop. To improve loop efficiency, we complement vision with tactile feedback to reduce uncertainty during physical interaction. We further augment textual guidance with reusable code to reduce reasoning overhead during skill reuse. On LIBERO-10, RoboSkill improves first-episode success rates by 12.5--25.0 percentage points and reduces average runtime by 7.6--72.4% across four agents. On real robots, it improves success rates by 8.3 percentage points and reduces average runtime for successful trials by at least 14.4%.

Related papers