Pub-AI: AI in Science

Driving · 2026-09-30

AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving

Haoran Zhu, Wancong Zhang, Yann LeCun, Anna Choromanska

arXiv:2609.34085PDFCode

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

AD-E2E-JEPA is an action-conditioned JEPA world model for end-to-end autonomous driving that enables goal-conditioned zero-shot planning without training a driving policy. It adds a SIGReg-regularized learnable projector that compresses patch embeddings, reducing planning latency (0.8s for an 8-frame rollout over 256 trajectories) while retaining planning performance on NAVSIMv2.

Why it matters

The authors target the efficiency–planning trade-off in JEPA-based driving world models, and evaluate in a goal-conditioned zero-shot planning setup that avoids mixing world-model quality with policy learning. They report multiple planning and geodesic metrics beyond success-rate, and show projector pretraining can transfer to downstream imitation learning.

Method

  • Adapt JEPA-WM to end-to-end autonomous driving with a frozen DINOv3 backbone and an AdaLN-style predictor for action-conditioned latent prediction.
  • Introduce AD-E2E-JEPA: a SIGReg-regularized learnable projector on patch embeddings to reduce spatial patch count and embedding dimension, applied to both context-history and target-future branches.
  • Enable goal-conditioned zero-shot planning by selecting from a (subsampled) trajectory vocabulary using world-model rollouts toward a goal defined by ground-truth future observations; evaluate with planning and reliability metrics.
Abstract (from arXiv)

Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end autonomous driving (E2EAD). To isolate world-model quality from policy learning, we employ a goal-conditioned zero-shot planning setting that evaluates these models using ground-truth future observations as goals, without training any driving policy. We find that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. To address this trade-off, we propose \textbf{AD-E2E-JEPA}, which introduces a SIGReg-regularized learnable projector applied to projected patch embeddings. The projector reduces the number of planning patches by $16\times$ and the embedding dimension by $4\times$, achieving a $100\times$ inference speedup while retaining planning performance, with a 0.8-second runtime for an 8-frame rollout over 256 candidate trajectories. \textit{Without} training any driving policy, the world model itself reaches the goals located 20 meters away on average within the displacement of respectively 4.0/2.8 meters, using world-model rollouts over trajectory vocabularies of respectively 256/8,192 candidates. On the NAVSIMv2 benchmark, it achieves 67.3/72.9 EPDMS with multiplicative safety metrics and 84.1/86.5 EPDMS$^{\dagger}$ without them in goal-conditioned zero-shot planning. Experiments further show that the self-supervised pretrained projector improves downstream imitation learning performance from 80.2 to 85.4 EPDMS. The source code is available at https://github.com/HaoranZhuExplorer/AD-E2E-JEPA

Related papers