Pub-AI: AI in Science

Driving · 2026-09-30

doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving

Parthib Roy, Yash Tandon, Marcus Blennemann, Giovanni Tapia Lopez, Angel Martinez-Sanchez, Mohan M. Trivedi, Ross Greer

arXiv:2609.38028PDFCode

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

doPlan is a publicly available real-world autonomous driving dataset with 5,154 human-written passenger instructions aligned to variable-duration nuPlan contexts (30.0–508.8 s). The authors report that language-conditioned models may change predictions without reliably improving directional instruction following, and that the maneuver referenced by passenger instructions often occurs far beyond common 5 s prediction horizons.

Why it matters

The dataset targets language as persistent task context, where intent can be multi-stage and deferred. The authors’ diagnostics show that instruction sensitivity (e.g., increased response to language) does not necessarily yield directional compliance, and that evaluation windows can miss when requested maneuvers actually occur. This motivates models and evaluation that connect persistent intent to successive planning decisions over longer horizons.

Method

  • Build doPlan from the official nuPlan training split; collect human passenger instructions via a taxi-test style protocol aligned to variable-length driving windows (30.0–508.8 s).
  • Evaluate four language-conditioned driving models (OpenEMMA, AutoVLA, Alpamayo 1.5, Alpamayo 2.0) at a common evaluation time point across language conditions to isolate effects of language input.
  • Use temporal grounding diagnostics for matched future turns and analyze language conditioning sensitivity via inference-time navigation-guidance ablation (Alpamayo 2.0).

Limitation

doPlan does not independently revalidate every retained instruction through a second round of manual review, and it does not distinguish referenced information visible at the evaluation time point from prior knowledge; temporal offsets are relative to the selected evaluation time point.

Abstract (from arXiv)

Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving datasets largely focus on short, localized interactions, leaving these longer-horizon forms of passenger intent comparatively underexplored. We introduce doPlan, to our knowledge the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context. Built on nuPlan, doPlan contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned context over 50.9 hours of unique driving, with annotation windows ranging from 30.0 to 508.8 s. The annotations capture immediate, deferred, event-conditioned, persistent, and multi-stage passenger intent. The dataset, annotation interface, and supporting resources are publicly available at https://github.com/Mi3-Lab/doPlan. We evaluate four language-conditioned driving models and find that sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction. More broadly, among 2,161 examples with a matched future maneuver, the first associated maneuver occurs a median of 24.6 s after the evaluation point, and only 9.8% occur within the models' common 5 s prediction horizon. These findings highlight the need to connect persistent passenger intent with successive planning decisions. doPlan provides a setting for studying how unresolved goals can be retained, grounded in evolving scenes, and tracked across multiple stages, including how a planner determines when a future goal becomes relevant to the current plan.

Related papers