FlashForward

In-Flight KV Cache with Clean Anchors
for Faster Autoregressive Video Diffusion

1Meta2S-Lab, Nanyang Technological University

TL;DR

Qualitative comparisons

Each sample shows one prompt generated by Self-Forcing, HiAR and Ours, played in sync. Choose a length, tick the methods you want to watch, switch samples with the arrows, and zoom in to inspect details.

Every clip on this page is heavily compressed to keep a small size. This compression may cause banding, smeared detail, and lost texture. Please try Chrome if the video cannot be loaded.

Methods
Sample 01 20s

←→ switch sample · Space pause · R restart · Z zoom · ,. step one frame · [] jump to an anchor · 123 length

How fast it runs

In our renderer, every denoising forward both advances the video and publishes the memory the next chunk reads, so no pass is spent on cache updates alone and successive chunks overlap across GPUs. Pick a setting to see the measured per-video latency.

Model
Resolution

How the video is generated

The planner auto-regressively constructs a sparse set of clean anchors. The renderer renders every chunk with same-stage history and two-sided clean anchors, one denoising stage per GPU, removing the dense KV cache update model forward. Play the diagram to step through a full pass.

0 / 19

How the model is trained

One backbone learns both roles in a single packed forward, then a second phase distils it onto its own rollouts. Play the diagram to step through both phases.

Issues we observed and solved in our framework

Three cheaper designs we tried first. Each one fails in a way the final design avoids or greatly reduces; the clips at the top of this page are what we ship instead. All three are the four-step model.

Stage-matched history only

The renderer reads only its own same-stage KV, with no clean anchors. Nothing holds neighbouring chunks to a shared structure, so appearance and motion drift across chunk boundaries.

Spliced anchors, planner KV

Keeping the planner's anchor latents in the output is cheaper, but those frames came out of a different pass: a seam appears at every anchor position.

Spliced anchors, renderer KV

Re-encoding the anchors through the renderer first shrinks the seam, but visible luminance jumps remain at the anchor boundaries.

We therefore use the planner's anchors only as conditioning and let the renderer generate every output latent, which stabilizes the drift and greatly reduces the seam.

BibTeX

@article{wang2026flashforward,
  title   = {In-Flight {KV} Cache with Clean Anchors for Faster Autoregressive Video Diffusion},
  author  = {Wang, Yikai and Han, Xiao and Xu, Mengmeng and P{\'e}rez, Juan C. and Douratsos, Yiannis and He, Sen and Zhou, Zijian and Zhang, Fei and An, Zhaochong and P{\'e}rez-R{\'u}a, Juan-Manuel and Loy, Chen Change and Xiang, Tao},
  journal = {arXiv preprint arXiv:2609.32540},
  year    = {2026}
}
Magnify