←→ switch sample · Space pause · R restart · Z zoom · ,. step one frame · [] jump to an anchor · 123 length
TL;DR
-
Hidden cache-only forwards
Autoregressive video diffusion generates chunk by chunk. To remember past chunks, prior methods run extra forwards that rebuild a clean or less-noisy KV cache but never advance the video.
-
In-flight KV reuse
Every denoising forward already computes its chunk's KV. FlashForward hands it straight to the next chunk, removing context re-encoding, so with one GPU per denoising stage, different chunks run concurrently.
-
Sparse clean anchors
Reused history is noisy, so appearance and motion drift. A planner generates sparse clean anchors ahead of time, giving every chunk clean guidance from both the past and the future.
-
Faster, higher quality
With up to four GPUs on videos of 20 s or longer: 1.16–1.69× faster than HiAR and 1.42–2.92× faster than Self-Forcing, with higher VBench scores that stay stable up to 65 s.
Qualitative comparisons
Each sample shows one prompt generated by Self-Forcing, HiAR and Ours, played in sync. Choose a length, tick the methods you want to watch, switch samples with the arrows, and zoom in to inspect details.
Every clip on this page is heavily compressed to keep a small size. This compression may cause banding, smeared detail, and lost texture. Please try Chrome if the video cannot be loaded.