Pub-AI: AI in Science

Driving · 2026-09-30

VehicleArena: A Realistic Urban Environment for Multi-Agent Driving

Jie Yang, Jiajun Chen, Jiazheng Zhou, Mianqiu Huang, Yining Zheng, Yuxin Wang, Xipeng Qiu

arXiv:2609.35916PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

VehicleArena is a high-fidelity 3D urban-driving benchmark for independently operating multi-agent driving with passenger requests and cockpit/cabin actions in a single closed loop. It provides 112 held-out evaluation tasks (80 single-agent, 32 multi-agent). The authors report that the highest arrival rates across nine evaluated models are only 65.0% (single-agent) and 65.6% (multi-agent), and driving policies can reduce surrounding vehicles’ arrival rates versus SUMO’s native traffic controller.

Why it matters

The benchmark targets independent-objective physical coexistence: agents with independent goals interact through persistent physical consequences in shared traffic, not predefined collaboration/competition protocols. The authors emphasize separate evaluation of trip completion, passenger-request/cabin fulfillment, traffic externalities, and inference cost, showing that strong request/cabin scores do not reliably imply successful trip completion and that focal policies can create measurable externalities for surrounding vehicles.

Method

  • Builds a shared 3D urban simulation where Personal Agents issue passenger requests, Driving Agents execute driving and cabin operations, and a Judge verifies frozen criteria from recorded evidence.
  • Uses SUMO’s native driving policy as an oracle baseline, holding the same environment and surrounding-agent configurations while swapping only the ego controller; deadlines are based on last required SUMO arrival time and the last scheduled environmental event plus a margin.
  • Evaluates multiple outcomes separately: ego trip completion and driving quality, passenger-request fulfillment and cabin compliance, externalities on surrounding traffic, and model token use for efficiency.
Abstract (from arXiv)

Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction protocols, leaving such emergent physical coupling underexplored. We introduce VehicleArena, a 3D urban-driving benchmark for studying independently operating agents in a dynamic shared world. In VehicleArena, LLM-controlled agents must fulfill evolving passenger requests while navigating complex traffic, and each agent's driving decisions can reshape traffic flow, delays, risks, and subsequent observations for surrounding agents. The benchmark provides 112 evaluation tasks spanning single-agent and multi-agent driving. Across nine evaluated models, the highest arrival rates reach only 65.0% on single-agent tasks and 65.6% on multi-agent tasks, while strong passenger-request or cabin scores do not reliably translate into successful trip completion. Moreover, in matched multi-agent runs, every tested focal policy reduces the arrival rate of surrounding vehicles relative to the simulator's native traffic controller, revealing measurable externalities beyond the focal vehicle itself.

Related papers