Pub-AI: AI in Science

Latent Dynamics · 2026-09-30

Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality

Linhao Wang, Yiyan Fan, Dongjin Huang

arXiv:2609.32322PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors show that two world models with similar total prediction error can yield very different planning outcomes when their errors occur on different state dimensions. They introduce Decision-Relevant Prediction Error (DRPE), measuring multi-step error only on decision-relevant state dimensions, and an iso-error protocol that varies error allocation while keeping total error fixed.

Why it matters

Prediction error alone can be a weak proxy for control quality because not all state dimensions equally affect decisions. The paper argues for evaluating world models using DRPE (and error allocation tests) so the errors that matter for planning are directly measured, including how the “relevant” dimensions can depend on the task and how deeper imagination amplifies decision-relevant errors.

Method

  • Construct a controlled 8×8 gridworld with 12 state dimensions, where 6 are decision-relevant and 6 are decision-irrelevant, and evaluate 55 controlled/learned models with a common standardized planner.
  • Introduce DRPE: a multi-step prediction error that weights errors only on decision-relevant dimensions (via a relevance weight vector).
  • Use an iso-error protocol: vary error allocation between relevant and irrelevant dimensions while matching total multi-step prediction error, then compare planning success.

Limitation

The authors state that DRPE is not sufficient across bias and noise structures, nor at event level granularity; a complete description would require the distribution of errors over events relevant to decisions.

Abstract (from arXiv)

World models are typically trained and evaluated by prediction error, assuming that more accurate predictions lead to better decisions. We show that this assumption can fail because models with similar total error can differ substantially in planning performance when their errors occur on different state dimensions. We introduce Decision-Relevant Prediction Error (DRPE), which measures prediction error on the state dimensions that affect decisions. We also develop an iso-error evaluation protocol that varies error allocation while keeping total error fixed. In a factored gridworld with known state relevance and a standardized planner, we evaluate 55 controlled and learned models across different error levels and allocations. Total prediction error is weakly related to planning success (Spearman $\rho=-0.25$), whereas DRPE is strongly predictive ($\rho=-0.84$; $-0.98$ within the controlled family). Models with only a 1\% difference in total error can differ by 60 percentage points in planning success (97\% vs 37\%). The relevant error also depends on the task, with model rankings reversing across tasks at the same total error. Deeper imagination further amplifies decision-relevant errors, while learned models exhibit systematic bias on rare but decision-critical events. We formalize sufficient conditions under which DRPE correctly ranks models and total prediction error cannot.

Related papers