Pub-AI: AI in Science

Latent Dynamics · 2026-09-30

Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning

Merve Atasever, Keyan Azbijari, Cagan Bakirci, Bo-Ruei Huang, Tolga Izdas, Zahra Shahrooei, Richard Yang, Erdem Biyik, Jyotirmoy V. Deshmukh

arXiv:2609.37519PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

Video2STL converts observation-only videos into parametric Signal Temporal Logic (STL) specifications for robot learning. A VLM extracts an embodiment-independent semantic event trace, then generates a bank of parametric STL formulas; robot trajectories ground predicate thresholds and temporal bounds. The method uses two timescales: rolling-window robustness for dense local rewards and a causal monitor over a retained long-horizon formula for progress rewards, enabling cross-embodiment transfer.

Why it matters

The authors argue that prior video-to-reward methods hide or entangle temporal structure, making tasks harder to inspect, ground, and reuse—especially in cross-embodiment settings. Video2STL aims to provide an interpretable intermediate representation by separating symbolic task structure from embodiment-specific numerical grounding, then leveraging STL robustness semantics to produce RL reward signals with explicit temporal satisfaction boundaries.

Method

  • Use a two-stage VLM pipeline: video → embodiment-independent semantic event trace → bank of parametric STL (PSTL) formulas with symbolic temporal parameters and no robot-specific numerical thresholds.
  • Ground predicate thresholds and temporal bounds from successful robot trajectories using quantiles, and use a disjoint held-out set for expert-consistency filtering to discard incompatible formulas.
  • Construct RL reward with two timescales: dense local rewards from rolling-window quantitative robustness of short-horizon formulas, plus one-time progress rewards from a causal monitor over a retained long-horizon specification.

Limitation

The authors only report experiments for their instantiated manipulation tasks and quadruped locomotion settings; they do not state guarantees or results beyond these domains.

Abstract (from arXiv)

Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched demonstrations. A central challenge is deciding what information should be transferred from the video to the robot. Existing approaches commonly convert visual observations into scalar similarity or value signals, or ask foundation models to directly generate reward code. These approaches can make the temporal structure of a task difficult to inspect, ground, and reuse. We present Video2STL, a framework that converts observation-only videos into parametric Signal Temporal Logic (STL) specifications and uses the resulting formal representation for robot learning. A vision-language model extracts an embodiment-independent semantic event trace and constructs a bank of symbolic temporal specifications. The model determines the task structure, while numerical predicate thresholds and temporal bounds are grounded from successful robot trajectories. For policy learning, we separate short- and long-timescale temporal information: short-horizon specifications provide dense rewards through rolling-window quantitative robustness, while a causal monitor over a retained long-horizon specification provides one-time progress rewards for valid temporal prefixes. The same representation supports cross-embodiment transfer from human or animal videos to robot control. Across four manipulation tasks, Video2STL achieves $85.8\%$ average success-once and $67.0\%$ success-at-end, compared with $81.5\%/59.5\%$ for native dense PPO and $65.0\%/42.3\%$ for Text2Reward; in quadruped locomotion, Qwen-3.8 and GPT-5.6-based Video2STL policies achieve $100\%$ success across velocities from $0.3$ to $2.1\,\mathrm{m/s}$ while remaining competitive in high-speed energy efficiency. Project webpage: \href{https://video2stl.github.io/}{video2stl}.

Related papers