Video World Models · 2026-09-30
Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware
Rajit Rajpal, Shahbuland Matiana, Liew Wei Pyn, Anmol Agarwal, Ryan Craig, Andrew Lapp, Mithun Hunsur, Sami BuGhanem, Scottie Fox, Aaron Sanders Carson Poole, Irene Park, Dave Rossi, Spencer Frazier, Louis Castricato
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
Waypoint-1.5 is a real-time diffusion video world model for interactive generation on consumer GPUs. It uses a single-stream causal Diffusion Transformer over TAEHV 1.5 latents and is conditioned on full keyboard+mouse inputs. Trained on 100,000 hours of control-aligned gameplay, it produces playable rollouts with a KV-cache runtime and distillation that reduces denoising to 4 steps per frame. It is benchmarked via latent FPS throughput.
Why it matters
The authors target interactivity under strict latency/throughput constraints on consumer hardware. They evaluate real-time performance using latent FPS across multiple consumer GPUs and report that the INT8-quantized 720P model meets a 30 latent FPS threshold on GPUs from RTX 5060 Ti upward, with a faster 360P variant also meeting the threshold on tested hardware including RTX 3060.
Method
- Model: a 1.28B-parameter single-stream causal Diffusion Transformer (DiT) over TAEHV 1.5 latents, with frame-causal attention and a static rolling KV cache.
- Training: Diffusion Forcing pretraining with per-frame independent noise levels and controller conditioning (button states, mouse displacement, scroll), plus post-training Self-Forcing DMD distillation.
- Runtime: autoregressive generation producing one latent frame at a time, with pre-allocated ring-buffer KV-cache and ahead-of-time compilation (torch.compile) for throughput-critical steps; INT8 quantization for benchmarks.
Limitation
The paper notes that safety in interactive world models presents inherent tensions between creative freedom and content restriction, and that they do not claim to have resolved these tensions.
Abstract (from arXiv)
We present Waypoint 1.5, a real-time diffusion world model for interactive video generation on consumer-grade hardware. Unlike general video diffusion models, interactive world models (iWMs) must respond to dense user controls under strict latency and throughput constraints. Waypoint 1.5 is pre-trained on 100,000 hours of diverse, control-aligned video game data across hundreds of games, and generates playable video conditioned on full keyboard and mouse input. The model includes two resolution variants that run across a wide spectrum of consumer hardware. To characterize this unique setting, we distinguish rendered FPS, latent FPS, and control rate. We describe the data pipeline, architecture, training methodology, and runtime system behind Waypoint 1.5. We evaluate interactivity through latency and throughput. Finally, we discuss the safety and ethics considerations unique to iWMs.