HLA: Expressive Hybrid Linear Attention
via Chunk-Wise Dynamic Mixing

Code MethodResults
Motivating observations. Long histories call for selective, query-dependent access.
Motivating observations. Long histories call for selective, query-dependent access.

Abstract

Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fixed with respect to input content and therefore cannot adapt historical access to each query. We introduce Hybrid Linear Attention (HLA), a query-dependent chunk-level attention mechanism for Gated DeltaNet (GDN). HLA represents each completed chunk as an exact affine state transition and computes content-dependent routing gates from compact, self-attentively pooled representatives. Each gate interpolates the corresponding historical transition with the identity map, controlling both the chunk’s additive memory and its transformation of earlier states. Effective-support regularization further encourages concentrated routing for sparse inference.

We evaluate HLA under both pretrained adaptation and from-scratch training. Across Qwen3.5 models from 0.8B to 9B, HLA consistently improves over native GDN and fixed chunk mixing, with gains of up to 5.57 percentage points on LongBench-V2 and 3.97 points on RULER. In a controlled from-scratch 1.3B setting trained for 100B tokens with a 4K context, HLA also improves RULER performance from 4K to 32K, with gains increasing from 0.83 points at 4K to 4.22 points at 32K. These results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries.

Method

HLA architecture: chunk summaries, query-dependent routing, and affine state recomposition
Chunk-wise dynamic mixing · HLA architecture
Ãᵢ(q) = (1 − gᵢ(q))I + gᵢ(q)Aᵢ   ·   B̃ᵢ(q) = gᵢ(q)Bᵢ

A zero gate skips a chunk’s transition; a unit gate retains it. Gating both the multiplicative and additive terms lets HLA preserve earlier information when a historical chunk is irrelevant.

Summarize

Represent each completed chunk by an affine state transition and pooled routing keys.

Route

Compute query-dependent gates from compact pooled keys.

Recompose

Interpolate historical transitions with identity and compose them in chronological order.

Quantitative Results

GDN and HLA use the same 1.3B-scale architecture, 100B-token training budget, and 4K training context. Scores below are RULER macro averages over all 13 tasks, with 500 matched examples per task and context length.

Context GDN HLA (ours)
4K 24.63 25.46
8K 15.02 17.69
16K 7.98 11.05
32K 3.65 7.87

RULER macro averages.