Pub-AI: AI in Science

Driving · 2026-09-30

RoboChrono: A Real Robot Benchmark for Streaming Task Understanding

Yuzhou Wu, Longteng Fan, Zimeng Li, Yu Wanchan, Ting Zhang, Yiyang Ma, Shihao Li, Wei Ying, Jianbin Qin, Jiajian Jing, Fangwen Chen, Yifan Wu, Zichen Zhang, Ruiqi Yang, Weibin Kong, Yihang Xu, Haoran Liu, Zonghang He, Xuyang Liu, YiFan Xiong, Siteng Huang, Tao Xu, Zhuo Xu, Long Chen, Ruoxiang Li

arXiv:2609.36605PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

RoboChrono is a real-robot benchmark for streaming task understanding with 39 scenarios and 34,713 evaluation instances, using a causal observation boundary (only visual observations up to the query time). Across 18 multimodal models under zero-shot evaluation, the authors report large task-dependent capability differences, especially on temporal ordering, view matching, and action time localization, plus input-ablations showing different dependencies on visual evidence.

Why it matters

The authors argue that aggregate scores can hide distinct failure modes in streaming robot manipulation understanding. RoboChrono provides a diagnostic evaluation over seven complementary tasks (recognition, alignment, temporal grounding) under restricted observation histories, and controlled ablations to examine how visual evidence, temporal order, and task/action priors affect current-state recognition vs next-action prediction vs temporal grounding.

Method

  • Benchmark with causal observation boundary: at time t, models receive only the observed video prefix O_{≤t} and must answer without future frames.
  • Seven evaluation tasks (multiple-choice accuracy for six tasks; Action Time Localization uses Recall@1 at tIoU≥0.5).
  • Zero-shot evaluation of 18 vision-language models, plus input ablations using matched questions and frame/observation perturbations.

Limitation

A model–task run with a response parse failure rate of 50% or higher is treated as an unsuccessful run and is excluded from reported comparisons.

Abstract (from arXiv)

Understanding ongoing robot manipulation requires models to interpret visual observations in relation to interaction history and task progress. We introduce RoboChrono, a benchmark for streaming task understanding comprising 39 scenarios and 34,713 evaluation instances, constructed from real robot executions and complementary bare-hand human recordings. The benchmark evaluates seven tasks grouped into recognition, alignment, and temporal grounding, covering action understanding and anticipation, visual correspondence, temporal ordering, and action localization. Zero-shot evaluation of 18 vision-language models reveals substantial differences across tasks. GPT-6-Astra achieves 98.3% accuracy on Frame Matching but 68.3% on Frame Ordering, while RynnBrain1.1-122B-A10B exhibits a larger gap, reaching 95.4% and 32.9%, respectively. Input ablations on matched questions with five open-weight models further reveal distinct dependencies on visual evidence: removing visual observations reduces Current Action Recognition accuracy by 22.1 percentage points, whereas Next Action Prediction decreases by only 0.7 points. These findings show that strong visual matching does not consistently coincide with strong temporal ordering, and suggest that next-action prediction can be supported by task and action priors even when visual evidence is unavailable. RoboChrono provides a diagnostic setting for examining these differences, highlighting the need for capability-specific evaluation beyond aggregate scores when assessing task understanding in robot manipulation.

Related papers