Pub-AI: AI in Science

Driving · 2026-09-30

RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

Shifeng Bao, Fanding Huang, Yihan Lin, Youhe Feng, Guanlin Li, Chen Zhao, Yang Li, Jiawei He, Cheng Chi, Jing Zhang

arXiv:2609.37583PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

RoboHarn-Evo evolves Hierarchical Physical Knowledge (HPK) from physical interaction experience to improve long-horizon robotic manipulation without updating the base vision-language model or low-level executor. It uses dual loops: an inner execution loop that retrieves Task Knowledge and Action Knowledge, and an outer knowledge-update loop that revises entries via task-completion and physical-effect checks. Experiments on RMBench and RoboDojo report large held-out and transfer gains.

Why it matters

The authors target continual self-improvement for robotic manipulation while keeping the base vision-language model and executor fixed. By separating verification criteria at task and action levels (subtask completion vs. physical effects) and using evidence to revise only eligible knowledge, the approach aims to correct historical errors while improving performance on held-out configurations and enabling zero-shot cross-domain transfer.

Method

  • Introduce HPK with two levels: Task Knowledge for subtask selection/completion, and Action Knowledge for object-relative geometric strategies and their physical effects.
  • Use a dual-loop harness: inner execution retrieves Task and Action knowledge and grounds tool calls in the current scene; outer loop uses recorded observations/tool calls/physical outcomes to revise knowledge entries.
  • Maintain HPK via hierarchical verification: Task entries are checked against subtask completion, Action entries against intended physical effects, then update applicability and retrieval eligibility using evidence.
Abstract (from arXiv)

Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop harness that evolves Hierarchical Physical Knowledge (HPK) from physical experience. HPK couples two levels of reusable knowledge: Task Knowledge captures which subtask should be executed and when it is complete, while Action Knowledge captures object-relative geometric strategies and their physical effects. During execution, the agent retrieves knowledge at the corresponding decision level and grounds it in the current scene under the task goal. Across episodes, physical feedback is used to revise historical knowledge, update its applicability, and organize reusable entries for subsequent retrieval. Experiments on RMBench show that HPK improves average success by up to 24.2 percentage points across different agent models. With 80 interaction rollouts, held-out success rises from 48.3% to 75.0% for GPT-5.5 and from 70.0% to 88.3% for GPT-6. RoboHarn-Evo also resolves over 83% of historical knowledge errors while retaining 95.8% of valid knowledge, and transfers zero-shot from RMBench to RoboDojo with gains of 35.0 and 25.0 percentage points. These results demonstrate that physical interaction can be accumulated into reusable knowledge for improving subsequent manipulation.

Related papers