HLA-WM: Hybrid Linear Attention for
Long-Horizon Video World Models

Code MethodResultsVideos
Long-range forgetting and revisit consistency.
Long-range forgetting and revisit consistency.

Abstract

Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memory cost. However, we identify severe long-range forgetting in Gated DeltaNet (GDN), where information from distant but relevant scenes is progressively attenuated by subsequent state updates. To address this limitation, we propose HLA-WM, a training-free hybrid linear-attention framework that combines coarse-grained geometry-guided retrieval with fine-grained recurrent linear-state computation. HLA-WM exploits the affine structure of GDN to cache compact chunk-wise transition summaries, retrieve scene-relevant historical chunks using camera geometry, and recompose them into query-specific recurrent states. On the 60-second SANA-WM-Bench, HLA-WM improves all six aggregate revisit-consistency and camera-control metrics of the base autoregressive generator without additional training, including a 0.74 dB PSNR gain and a 28.5% reduction in rotation error. The improvements persist after downstream refinement and generalize to MBench-A, where HLA-WM consistently improves all three revisit-consistency metrics across all four subsets and all evaluated inference modes over 547 samples. At a 60-second context, HLA-WM reduces historical-state memory by 12× relative to full KV caching while incurring at most a 1.6% reduction in inference throughput. These results demonstrate that selectively addressable recurrent memory can improve long-range scene recall while preserving the efficiency advantages of GDN.

Method

HLA-WM overview. Geometry selection, chronological state recomposition, and aligned KV retrieval.
HLA-WM overview.

HLA-WM retains chunk-wise GDN states during generation and uses geometry to retrieve relevant history and recompose the GDN state.

Quantitative Results

SANA-WM-Bench, Top1 retrieval. Choose a trajectory split to compare Stage 1 and both refinement modes.

Pipeline Method PSNR ↑ SSIM ↑ LPIPS ↓ RotErr ° ↓ TransErr ↓ CamMC ↓ FPS ↑
Stage 1 SANA-WM 9.18 0.1729 0.6327 19.7040 2.2625 2.3974 22.403
Stage 1 HLA-WM 9.91 0.1999 0.6049 13.5864 2.0928 2.1798 22.053
AR refinement SANA-WM 14.44 0.2770 0.5738 20.9315 2.1587 2.3121 8.280
AR refinement HLA-WM 14.82 0.2872 0.5682 20.2933 2.1910 2.3240 8.211
Bi. refinement SANA-WM 13.04 0.2945 0.5786 11.3848 1.9893 2.0545 7.701
Bi. refinement HLA-WM 13.56 0.3148 0.5614 8.8823 1.8637 1.9127 7.700

Benchmark results. Higher is better for PSNR, SSIM, and FPS; lower is better for the other metrics.

MBench-A · 547 samples
Pipeline Method PSNR ↑ SSIM ↑ LPIPS ↓
Stage 1 SANA-WM 10.17 0.2545 0.6251
Stage 1 HLA-WM 11.01 0.2848 0.5823
AR refinement SANA-WM 12.36 0.2996 0.6554
AR refinement HLA-WM 12.73 0.3122 0.6447
Bi. refinement SANA-WM 12.94 0.3310 0.5597
Bi. refinement HLA-WM 13.52 0.3431 0.5470

Ablation results.