Driving · 2026-09-30
Lookahead-R: Budget-Aware Tool Retrieval via Execution-Centric Planning
Zongze Wu, Yani Guo, Runnan Li
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
Lookahead-R reframes tool retrieval as a resource-constrained sequential decision problem. It trains an execution-aware surrogate world model to predict tool execution success, latency cost, and semantic utility without invoking real APIs, then uses a cost-sensitive, uncertainty-guided Monte Carlo Tree Search to select tools under latency budgets. On ToolBench I3, the authors report NDCG@5 = 91.40%, outperforming ToolGen (90.16%) by 1.24%.
Why it matters
Tool retrieval bottlenecks LLM agents over large API ecosystems because semantic relevance and execution feasibility can diverge. The authors’ approach targets the semantic–functional gap by predicting execution outcomes and latency, enabling planning under strict latency/resource budgets instead of relying on static similarity ranking.
Method
- Reformulates tool retrieval as a Resource-Constrained Multi-Objective Planning problem and uses a learned execution-aware surrogate world model to predict execution outcome, utility, and latency without real API calls.
- Builds a cost-sensitive, uncertainty-guided Monte Carlo Tree Search that performs “virtual rollouts,” prunes high-cost/high-risk branches, and is designed as an anytime algorithm.
- Uses a Student–Teacher Mining protocol to synthesize execution traces (success, latency, error logs) to train the world model; ablations attribute key gains to explicit latency modeling.
Limitation
The provided text does not state limitations of the proposed method.
Abstract (from arXiv)
Tool retrieval is a critical bottleneck for LLM-based agents operating over large, heterogeneous API ecosystems. Existing approaches face an inherent trade-off: semantic retrievers are fast but suffer from the semantic-functional gap, while execution-based validation improves precision at the cost of prohibitive latency. We propose Lookahead-R, a planning-based framework that reformulates tool retrieval as a resource-constrained sequential decision-making problem. At its core, Lookahead-R introduces a lightweight execution-aware surrogate world model that jointly predicts tool execution success, latency cost, and semantic utility---without invoking real APIs. This world model drives a cost-sensitive, uncertainty-guided Monte Carlo Tree Search that navigates the tool space under strict budget constraints. Evaluated on the large-scale ToolBench benchmark, Lookahead-R achieves a superior accuracy-efficiency trade-off across all test scenarios. On the most challenging I3 split, it attains an NDCG@5 of 91.40\%, outperforming the state-of-the-art ToolGen (90.16\%) by 1.24\%. Ablation studies confirm that explicit latency modeling is the key discriminative signal for identifying high-quality tools under resource constraints.
Related papers
- The GUI Is Not the State: Diagnosing State Aliasing in GUI World Models
- Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents
- Towards an AI Software Factory for Data Systems
- What Must a World Model Distinguish for Planning?
- SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation