Pub-AI: AI in Science

Driving · 2026-09-30

Stochastic World Models for Verifying Vision-Based Neural Feedback Systems

I. Samuel Akinwande, Mykel J. Kochenderfer, Clark Barrett

arXiv:2609.38120PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors train stochastic, physically grounded perception surrogates (world models) for verifying vision-based neural feedback systems. Their verification procedure combines falsification, adaptive refinement, symbolic, and backward analyses. On an emergency braking benchmark, they report resolving the entire state space (38% previously unresolved) with a GAN surrogate, and resolving over 80% of the state space on the RGB benchmark with a world-model surrogate.

Why it matters

Vision-based neural feedback verification depends on how well the perception surrogate captures sensor variation while staying tractable for closed-loop analysis. The paper argues that stochastic surrogates with physically grounded latent factors can be bounded by standard verifiers, and pairs this with a verification procedure that aims to reduce conservatism via refinement, symbolic reasoning, and backward reasoning.

Method

  • Formalize vision-based neural feedback systems as state-based systems where the sensor is a generative model mapping (state, latent) to observations.
  • Train a stochastic world model renderer with physically grounded latents (a box of physical conditions) using FiLM modulation, with layers bounded by standard verifiers and McCormick relaxation for FiLM products.
  • Verify surrogates via a partitioned-cell procedure using falsification, forward analysis with abstractions/optimization, adaptive refinement, symbolic analysis, and backward analysis.

Limitation

The authors do not claim soundness against floating-point error inside the bounding library.

Abstract (from arXiv)

Verifying a vision-based neural feedback system requires a model of the observations its controller acts upon. Such a model must capture the variation the sensor produces, while remaining tractable for closed-loop analysis. Generative adversarial networks (GANs) have served as perception surrogates, but they are large, reproduce complex scenes poorly, and are hard to verify. We explore stochastic world models as a richer class of perception surrogates. We train a world model with physically grounded latents, built from operations that standard verifiers bound. It reproduces held-out frames more faithfully than GAN surrogates with up to 130 times as many parameters. To verify these surrogates, we develop a procedure that combines falsification, adaptive refinement, symbolic, and backward analyses. On an emergency braking benchmark with a GAN surrogate, our procedure resolves the entire state space, 38% of which the state-of-the-art verifier left unresolved. On the RGB version of the benchmark, where no verification results have previously been reported, our procedure resolves over 80% of the state space with a world model surrogate.

Related papers