Late-Stage Patching Full-Attention Transformers to Gated DeltaNets

Anupam Nayak*1Arian Raje*1Anthony Fei2Akaash Parthasarathy1Mohamed Abdelfattah2Gauri Joshi1

1Carnegie Mellon University 2Cornell University

*Equal contribution

TL;DR

Q: Can we efficiently convert pretrained full-attention transformers into hybrid Gated DeltaNet (GDN) models while preserving the reasoning and long-context capabilities of the original model?

One-line summary: Prolonged On-Policy Distillation recovers long-context reasoning and retrieval alongside short-context capabilities, enabling pretrained full-attention transformers to be efficiently (< 0.01% pretraining budget) converted into hybrid GDN models while retaining much of the base full-attention transformer model’s capabilities.

Primer on Gated DeltaNets

A standard Transformer layer uses full causal self-attention. For a hidden state h_t \in \mathbb{R}^d, the layer produces

q_t = f_Q(h_t), \qquad k_t = f_K(h_t), \qquad v_t = f_V(h_t)

with q_t, k_t \in \mathbb{R}^{d_k} and v_t \in \mathbb{R}^{d_v}. At token t, the layer retains the keys and values seen so far in a KV cache, K_{\leq t} = [k_1, \ldots, k_t]^{\top} \in \mathbb{R}^{t \times d_k} and V_{\leq t} = [v_1, \ldots, v_t]^{\top} \in \mathbb{R}^{t \times d_v}. The attention output is

o_t = \mathrm{softmax}\!\left(\frac{q_t K_{\leq t}^{\top}}{\sqrt{d_k}}\right) V_{\leq t} \;\in\; \mathbb{R}^{d_v}

Every query interacts explicitly with all preceding keys. Across a sequence of length n, the total attention computation grows as O(n^2), while the KV cache grows linearly with sequence length.

A Gated DeltaNet (GDN) layer[3] instead compresses the history into a fixed-size state S_t \in \mathbb{R}^{d_v \times d_k}, rather than storing the full KV cache. As in attention, the layer produces

q_t = f_Q(h_t), \qquad k_t = f_K(h_t), \qquad v_t = f_V(h_t)

along with data-dependent forget and update gates

\alpha_t = f_\alpha(h_t) \in (0,1), \qquad \beta_t = f_\beta(h_t) \in (0,1)

The compressed state evolves as

S_t = \alpha_t S_{t-1} + \beta_t\left(v_t - \alpha_t S_{t-1} k_t\right) k_t^{\top}

and the layer output is obtained directly from this state,

o_t = S_t q_t \;\in\; \mathbb{R}^{d_v}

Instead of explicitly retaining every previous k_t and v_t, the GDN carries forward the fixed-size memory S_t, with \alpha_t controlling retention of the previous state and \beta_t controlling how strongly the new key-value association is written. Because the state size is fixed, the computational cost of a GDN layer grows as O(n) over a sequence of length n, compared with the O(n^2) cost of a full-attention layer. 11The equations above describe a single head and omit the details that make these layers work in practice; activation functions, the short causal convolutions applied to q, k and v, the normalization and output-gating layers, and the multi-head structure of both attention and GDN. fQ, fK, fV, fα and fβ stand in for the learned projections and their associated machinery.

Hybrid Models

Recent frontier models increasingly rely on mixing GDNs and full-attention transformers than choosing between them. Examples include Kimi K3[1] (Moonshot AI, 2.8T parameters, 1M-token context) and Qwen3.8[2] (Alibaba, 2.4T MoE).

A hybrid of full-attention and GDN layers can be substantially more efficient to serve than a pure attention model, particularly at long-context lengths, because most of its layers replace a KV cache that grows with every token and an attention computation that revisits the entire history at every step with a fixed-size memory and a constant per-token update.

The intuition behind the hybrid design is that attention and GDN layers provide different forms of memory. GDN layers can carry information over long spans through a compact recurrent state, making them efficient at preserving and updating context without attending over the entire history. Full-attention layers, meanwhile, retain direct access to individual past token positions, which is useful when the model needs to recover a specific detail or model an exact interaction across a long distance. Combining the two therefore offers efficient long-range memory through GDNs while preserving the precise retrieval capabilities of full-attention.

Proposed Pretrained Transformer to GDN conversion pipeline

STAGE 0

Architecture Morphing

Morphing is the first stage, aimed at converting an off-the-shelf pretrained full-attention model into a hybrid architecture, before any training begins. Every 4th layer retains its original softmax full-attention unchanged, while the attention blocks in the remaining layers are replaced with Gated DeltaNet[3] modules matching the Qwen3-Next[4] design. We apply this transformation to two base architectures: Qwen3-4B[5] and MiMo-7B-RL[6].

Qwen3-4B 36 layers · hidden 2560
0
3
7
11
15
19
23
27
31
35
MiMo-7B-RL 36 layers · hidden 4096
0
3
7
11
15
19
23
27
31
35
27 GDN layers 9 full-attention layers — 3, 7, 11 … 35 Both models use the same 3:1 convention — every fourth layer keeps full-attention, the rest carry a fixed-size recurrent state.

Parameters modified

Inside one GDN layer
componentQwen3-4BMiMo-7B
Query  Q−5.2M−8.4M
Key  K+2.6M+4.2M
Value  V+7.9M+12.6M
Output gate  Z+10.5M+16.8M
Gates  α, β+164K+262K
Convolution+33K+33K
Misc−64−6.0K
Net per layer +15.9M+25.5M
× 27 GDN layers +430M+687M
Whole model
parametersQwen3-4BMiMo-7B
Base model4.02B7.62B
Hybrid4.55B8.31B
Per full-attention layer+10.5M0
× 9 full-attention layers+94M0
Per GDN layer+15.9M+25.5M
× 27 GDN layers+430M+687M
Total added +524M+687M

Component by component22 Many of these architectural choices exist to match the configuration used in Qwen3-Next — the output projection on the full-attention layers, for instance; rather than because our method requires them.

  1. Query q — merged: The base model has 32 query heads; while GDN has 16. Each GDN query is the average of two heads in the base model, which results in loss of information.
  2. Key k and Value v — replicated: The base model has 8 key/value heads shared across its 32 query heads; GDN uses 16 key and 32 value heads. Each of the base model's head is copied out across the GDN heads that map onto it.
  3. Output projection, layer norms, embeddings and linear layers — copied unchanged: GDN has to project its output back into the residual stream, and its output width happens to equal the attention block's, so that weight transfers verbatim in both models. Everything outside the attention block transfers the same way — the linear layers, the layer norms and the embeddings are all copied exactly, in both kinds of layer. Only the attention block is replaced, which is why none of them appear in the table.
  4. Output gate — new: The GDN layers scale its own output by a learned gate. At 10.5M per layer on Qwen and 16.8M on MiMo it is roughly two thirds of everything conversion adds. However, this is added to all layers including full-attention in Qwen3-4B to match Qwen3-Next while only to GDN in MiM0-7B-RL.
  5. α and β — new: GDN specific parameters, gated: α decides how much of the memory to keep and β how strongly to write into it. They are one small projection, about 1% of the parameters added.
  6. Convolution — new: A short causal convolution runs over q, k and v before the recurrence. About 33K weights per layer in either model.
  7. Misc: The per-head decay terms, the timestep bias and the GDN block's own norm come to under 200 weights. Qwen3's per-head query and key norms (256 weights), and MiMo's query, key and value biases (6.1K, since MiMo's Qwen2-family attention carries biases where Qwen3's does not) are removed.
STAGE 1

Initialization: Off-Policy Representation Alignment

This is the first stage of our recovery pipeline. Throughout all stages, we use the original base model as the teacher and the morphed hybrid model as the student. The goal of this stage is to provide a strong initialization for subsequent training. Concretely, we train each student GDN layer to match the output of the corresponding teacher full-attention layer, while feeding both layers the hidden representation produced by the teacher’s preceding layer (hence called off-policy). This encourages the morphed model GDN layers to mimic the teacher full-attention layer's outputs by minimizing distribution shift between the teacher and student representations.

STAGE 2

Short context recovery: Off-Policy Distillation

Stage 1 calibrates each GDN block independently, using hidden-state inputs generated by the teacher model. While this provides a strong layerwise initialization, it does not account for interactions between the newly introduced GDN blocks: each block is optimized on teacher-generated representations and not the representations it will receive from preceding student layers at inference time. Consequently, approximation errors and representation shifts may accumulate when the student is composed end-to-end.

Stage 2 is the first stage that trains the language model end-to-end distilling the teacher’s output distribution into the complete student model, thereby encouraging globally consistent behavior across the full network. It runs in two substages: first at short context, then at the teacher's full window.

Stage 2a: Short-Context Distillation

Stage 2b: Long-Context Distillation

Many prior approaches keep the retained full-attention layers frozen throughout training while the GDN layers learn how to update and use the fixed size memory state. However, we freeze them only during stage 1, while the newly introduced blocks are being initialized. In stage 3, which is the longest stage of our training pipeline, the retained attention layers are updated jointly with the rest of the model.

This is important because the full-attention layer in pure transformers and hybrid models differ in the role those layers must play. Before conversion, every attention layer contributes directly to long-range interaction at its own depth. After conversion, the retained full-attention layers become the only points in the network where information can be accessed beyond the compressed state, and must therefore support long-range retrieval on behalf of the converted layers as well. Their pretrained weights were not optimized for this altered role, so they must adapt during recovery training. We observe that the best results are obtained when Stage 2 serves as a gradual bridge between stages 1 and 3, jointly updating the full model while using a lower learning rate for the retained full-attention layers so that they adapt on a slower timescale than the newly introduced GDN blocks.

STAGE 3

Long-Context Recovery: Prolonged On-Policy Distillation33The name and the philosophy are inspired by ProRL[16], which extends the RL post-training phase well beyond the usual budget because reasoning capability continues to improve rather than saturating early.

Stage 3 is the final and most important stage of the recovery pipeline. As in the earlier stages, the original base model serves as the teacher and the morphed hybrid as the student, with the objective of recovering the capabilities lost during conversion.

This is also the stage where long-context reasoning is restored. After Stage 2, the student can retrieve information from long inputs, but it still struggles to reason effectively over long contexts. Its AIME accuracy remains at 0, while MATH-500 scores under a 32K-token budget are only around 2.0%, substantially worse than performance with shorter 1–4K token budgets and thinking disabled. In other words, giving the model more room to reason actually hurts performance without stage 3.

This gap between long-context retrieval and long-context generation comes from a difference in distribution shift. During retrieval, nearly all of the tokens the model conditions on are provided in the input, so teacher-forced training covers the relevant distribution reasonably well. During long-form generation, however, the model repeatedly conditions on its own outputs. After only a few thousand generated tokens, it can enter regions of the sequence distribution that are rarely, if ever, represented in static training corpora. Off-policy training would need to anticipate and cover this space in advance, which is impractical for open-ended generation. Training directly on the student's own generated trajectories instead exposes the model to exactly the states it will encounter at inference time. Stage 3 additionally also improves long form retrieval performance.

Two learning-rate curves over training progress. Cosine rises through a short warmup then falls smoothly to zero. WSD rises through the same warmup, holds flat across the middle, then drops linearly at the end.
Figure 1. Cosine against warmup-stable-decay. Cosine begins shedding learning rate immediately; WSD holds it flat and spends it all in a short decay at the end.
Figure 1 from Wen et al., 2024, on warmup-stable-decay learning rates.
Figure 2. Typical loss curves for cosine and WSD for a supervised learning objective. Through the stable phase WSD runs above cosine, once the decay begins it falls sharply matching the cosine loss curve. Image taken from Wen et al., Understanding Warmup-Stable-Decay Learning Rates[7].
Training loss, reverse KL, plotted against tokens from 200M to 2B. Fourteen overlapping WSD curves, one per arm, each descending and ending in a steeper decay drop.
Figure 3. Training loss curves for WSD at each training horizon, for Qwen3-4B. Each arm branches from the pre-decay state of the one before it.

Other Methods

Existing methods for converting pretrained full-attention transformer layers into linear attention layers typically use some combination of layerwise alignment (Stage 1), end-to-end off-policy distillation (Stage 2)[12][13][14][15][20], and continual pretraining on additional data[21]. Some methods additionally perform some form of layer selection for linearization before training[15][20][22].

4In distillation, the data/prompts mediate the transfer between the teacher and student; therefore, from an information-theoretic perspective, the procedure is not entirely free of knowledge transfer from the data itself.We intentionally omit the continual pretraining stage to avoid transferring additional capabilities beyond those already present in the language model.4 In particular, if the morphed hybrid student were to exactly match the teacher distribution, the objectives in all three stages would reduce to zero. The continual learning stage sometimes uses 30–50x more tokens than our whole pipeline[21], often approaching modern midtraining budgets.

Main results

Commonsense Reasoning Suite : Short Context

Each linearized model is compared against its corresponding teacher. The final column reports the average across all six tasks; percentages in parentheses denote performance retained relative to the teacher average. This part is easy to recover even without Stage 3. We donot beat the baselines here and pay the short context tax for long-context performance. We think that Commonsense performance can be recovered somewhat cheaply with additional SFT as done by some of the baselines.

Model PIQA HellaSwag ARC-E ARC-C Winogrande MMLU Avg.
RADLADS[12]
Qwen2.5-7B-Instruct80.280.580.955.070.574.373.6
QRWKV6-7B-Instruct79.978.979.355.971.464.271.6 (97.3%)
Llamba[13]
Llama-3.2-1B-Instruct74.961.663.937.661.446.057.6
Llamba-1B73.761.865.237.561.931.555.3 (96.0%)
Llamba[13]
Llama-3.2-3B-Instruct76.871.671.046.468.760.765.9
Llamba-3B78.074.073.646.271.750.165.6 (99.6%)
Zebra-Llama[14]
Llama-3.2-3B-Instruct76.871.671.046.468.760.765.9
Zebra-Llama-3B77.372.775.852.365.751.765.9 (100.1%)
Zebra-Llama[14]
Llama-3.1-8B-Instruct81.479.579.955.673.668.473.1
Zebra-Llama-8B79.979.077.557.671.359.170.7 (96.8%)
HALO[15]
Qwen3-1.7B72.060.469.543.061.660.261.1
HypeNet-2B72.057.267.342.761.739.556.7 (92.8%)
HALO[15]
Qwen3-4B74.968.578.553.865.870.168.6
HypeNet-5B76.366.175.950.667.252.464.8 (94.4%)
OPAL
Qwen3-4B74.968.578.553.865.870.168.6
OPAL-Qwen3 (ours)72.358.565.041.663.057.259.6 (86.9%)
OPAL
MiMo-7B-RL-053074.564.968.145.061.657.261.9
OPAL-MiMo (ours)72.559.962.739.558.854.658.0 (93.7%)

Needle-in-a-haystack: Long Form Retrieval

Retrieval accuracy across context lengths, grouped by task type. The average is the macro-average over all twelve task–length combinations; percentages in parentheses denote performance relative to the corresponding teacher. One can unlock non trivial performance here without stage 3. Some baselines such as Zebra-Llama[14] perform well in shorter contexts (4–8K) but severely underperform at longer contexts (16-32K) and in the multi-key/query settings compared to the corresponding full-attention model, which has near perfect retrieval scores across the board. OPAL on the other hand has scores that match the corresponding full-attention model.

Model Single needle Multi-key Multi-query Avg.
4K 8K 16K 32K 4K 8K 16K 32K 4K 8K 16K 32K
RADLADS[12]
Qwen2.5-7B-Instruct100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0100.0
QRWKV6-7B-Instruct0.00.00.00.01.00.00.00.00.00.00.00.00.1 (0.1%)
Llamba[13]
Llama-3.2-1B-Instruct100.0100.0100.0100.094.093.071.067.098.594.293.296.592.3
Llamba-1B0.00.00.00.00.00.00.00.00.00.00.00.00.0 (0.0%)
Llamba[13]
Llama-3.2-3B-Instruct100.0100.0100.0100.098.095.094.087.0100.099.287.568.594.1
Llamba-3B2.00.00.00.03.03.00.00.00.80.00.00.00.7 (0.8%)
Zebra-Llama[14]
Llama-3.2-3B-Instruct100.0100.0100.0100.098.095.094.087.0100.099.287.568.594.1
Zebra-Llama-3B85.061.052.047.055.039.028.010.064.224.223.03.541.0 (43.6%)
Zebra-Llama[14]
Llama-3.1-8B-Instruct100.0100.0100.0100.0100.0100.099.099.0100.0100.0100.0100.099.8
Zebra-Llama-8B100.098.096.074.073.040.040.027.086.870.066.543.067.9 (68.0%)
HALO[15]
Qwen3-1.7B100.0100.0100.0100.099.0100.099.099.0100.0100.0100.0100.099.8
HypeNet-2B73.076.066.073.028.022.024.018.027.319.018.515.538.4 (38.5%)
HALO[15]
Qwen3-4B100.0100.0100.0100.0100.0100.0100.099.0100.0100.0100.0100.099.9
HypeNet-5B88.090.090.096.034.037.033.028.049.555.849.245.058.0 (58.0%)
OPAL
Qwen3-4B100.0100.0100.0100.0100.0100.0100.099.8100.0100.0100.0100.0100.0
OPAL-Qwen3 (ours)100.0100.0100.0100.0100.0100.0100.099.6100.0100.0100.0100.0100.0 (100.0%)
OPAL
MiMo-7B-RL-0530100.0100.0100.0100.099.699.098.892.899.999.798.376.297.0
OPAL-MiMo (ours)100.0100.0100.0100.096.896.894.883.8100.0100.0100.0100.097.7 (100.7%)

Mathematical reasoning : Long-Context Generation

This is the stage where Prolonged On-Policy Distillation yields the largest gains, with the improvements over baselines being most pronounced on harder benchmarks such as AIME, where most baselines fail to recover any performance. In fact the numbers reported below correspond to premature early stopped checkpoints. Extended training still shows signs of decreasing training loss.

Eval details
  • GSM: non-thinking mode with greedy decoding, generation length of 1024 under zero-shot prompting. GSM symbolic numbers are an average of the "main", "p1" and "p2" splits weighted by number of samples.
  • MATH-500: 32K generation length with thinking mode, temperature 0.7; reports avg@8 numbers.
  • AIME: thinking mode, temperature 0.6, generation length 32K; reports pass@1 (avg@8) and pass@8 (temp 0.6).
The final column averages the eight displayed metrics, with percentages in parentheses denoting performance relative to the corresponding teacher.
Model GSM8K GSM-Plus GSM-Symbolic MATH-500 AIME’24 AIME’25 Avg.
pass@1 pass@8 pass@1 pass@8
RADLADS[12]
Qwen2.5-7B-Instruct91.972.783.476.410.426.77.523.349.0
QRWKV6-7B-Instruct27.416.89.324.00.43.30.43.310.6 (21.6%)
Llamba[13]
Llama-3.2-1B-Instruct38.524.926.225.80.86.70.00.015.4
Llamba-1B28.112.913.08.60.00.00.00.07.8 (51.0%)
Llamba[13]
Llama-3.2-3B-Instruct71.158.359.539.02.513.30.00.030.5
Llamba-3B50.727.524.87.80.43.30.00.014.3 (47.0%)
Zebra-Llama[14]
Llama-3.2-3B-Instruct71.158.359.539.02.513.30.00.030.5
Zebra-Llama-3B62.738.836.720.00.43.30.00.020.2 (66.4%)
Zebra-Llama[14]
Llama-3.1-8B-Instruct86.167.073.248.62.910.00.00.036.0
Zebra-Llama-8B61.340.335.324.40.86.70.43.321.6 (59.9%)
HALO[15]
Qwen3-1.7B83.563.365.074.613.830.010.423.345.5
HypeNet-2B1.13.81.24.40.00.00.83.31.8 (4.0%)
HALO[15]
Qwen3-4B92.372.883.285.425.843.322.143.358.5
HypeNet-5B14.915.88.711.80.00.00.00.06.4 (10.9%)
OPAL
Qwen3-4B92.372.883.296.473.886.763.780.081.1
OPAL-Qwen3 (ours)83.263.865.992.052.580.040.463.367.6 (83.4%)
OPAL
MiMo-7B-RL-053080.261.671.697.073.383.367.983.377.3
OPAL-MiMo (ours)84.863.565.793.865.483.347.573.372.2 (93.4%)

How reasoning returns, stage by stage

The tables above compare finished models. The tables below trace the same models across the pipeline, which is where the long-context reasoning improvement is visible.

Long-form retrieval recovers faster and is essentially saturated after the first 200M tokens.

Qwen3-4B — mathematical reasoning across stages
Checkpoint MATH-500 AIME’24 think @32K AIME’25 think @32K
think @32K no-think @4K pass@1pass@8 pass@1pass@8
Teacher (Qwen3-4B)96.484.273.886.763.780.0
Stage 10.06.00.00.00.00.0
Stage 2a3.442.00.00.00.00.0
Stage 2b2.040.60.00.00.00.0
OPD 200M83.668.123.350.018.833.3
OPD 400M87.070.233.363.331.746.7
OPD 600M89.671.036.763.331.250.0
OPD 800M89.373.237.166.732.150.0
OPD 1B90.574.445.873.335.856.7
OPD 1.2B90.876.643.870.038.360.0
OPD 1.4B91.175.445.076.740.863.3
OPD 1.6B91.075.948.373.342.166.7
OPD 1.8B91.675.550.476.740.863.3
OPD 2B92.076.952.580.040.463.3
MiMo-7B-RL — mathematical reasoning across stages
Checkpoint MATH-500 AIME’24 think @32K AIME’25 think @32K
think @32K no-think @4K pass@1pass@8 pass@1pass@8
Teacher (MiMo-7B-RL)97.071.273.383.367.983.3
Stage 10.81.80.00.00.00.0
Stage 2a20.216.40.43.32.13.3
Stage 2b17.016.40.00.01.73.3
OPD 200M88.264.435.463.328.746.7
OPD 400M90.668.447.576.730.456.7
OPD 600M92.868.051.276.735.856.7
OPD 800M94.065.652.573.340.463.3
OPD 1B94.466.458.873.340.060.0
OPD 1.2B94.870.856.280.042.160.0
OPD 1.4B95.469.262.183.339.263.3
OPD 1.6B93.868.262.183.347.573.3
OPD 1.8B93.470.263.383.349.273.3
OPD 2B93.870.865.483.347.573.3

Converting to Mamba-3 Hybrid

5We do not intend to draw conclusions about whether one target architecture is inherently easier to convert to, or performs better after conversion, than another since we do not separately optimize an extensive set of hyperparameters for each architecture. More generally, however, we observe that conversions involving fewer randomly initialized parameters after morphing tend to be easier, suggesting that the pipeline may benefit from an architectural conversion curriculum that introduces changes to the backbone progressively rather than all at once (see discussion).The same OPAL conversion pipeline is used with Mamba-3[17] in place of GDN as the recurrent layer, producing OPAL-Qwen3-Mamba and OPAL-MiMo-Mamba.5

Mathematical reasoning
OPD Tokens MATH-500 AIME’24 AIME’25
GDNMamba3 GDNMamba3 GDNMamba3
Qwen3-4B
200M83.687.223.326.218.826.7
400M87.087.633.335.431.732.5
600M89.689.236.738.831.235.4
800M89.389.637.139.232.135.0
1B90.591.045.850.035.835.0
MiMo-7B-RL-0530
200M88.287.835.437.128.725.0
400M90.690.547.547.130.430.0
600M92.891.651.251.235.839.6
800M94.092.452.555.040.441.2
1B94.492.958.857.540.040.0
Commonsense reasoning
Model PIQA HellaSwag ARC-E ARC-C Winogrande MMLU Avg.
Qwen3-4B
Teacher74.968.578.553.865.870.168.6
OPAL-Qwen3-Mamba73.461.168.445.262.459.861.7 (89.9%)
MiMo-7B-RL-0530
Teacher74.564.968.145.061.657.261.9
OPAL-MiMo-Mamba72.659.760.340.458.651.457.2 (92.4%)
Needle-in-a-haystack retrieval
Model Single needle Multi-key Multi-query Avg.
4K8K16K32K 4K8K16K32K 4K8K16K32K
Qwen3-4B
Teacher100.0100.0100.0100.0100.0100.0100.099.8100.0100.0100.0100.0100.0
OPAL-Qwen3-Mamba100.0100.0100.0100.0100.099.899.699.4100.0100.0100.0100.099.9 (99.9%)
MiMo-7B-RL-0530
Teacher100.0100.0100.0100.099.699.098.892.899.999.798.376.297.0
OPAL-MiMo-Mamba100.0100.0100.0100.098.099.095.087.0100.0100.0100.0100.098.3 (101.3%)

Discussion

We establish prolonged on-policy distillation as an effective approach for converting trained Transformer layers into efficient architectures such as Mamba-3 and Gated DeltaNet, while recovering both short- and long-context capabilities. Our method requires no additional teacher beyond the original full-attention Transformer and can successfully hybridize off-the-shelf pretrained models, including those that have undergone extensive RL post-training, with only a modest additional training budget of roughly 3B tokens. Moreover, performance continues to improve as the on-policy stage is prolonged, with gains remaining unsaturated at the end of our training budget. This brings us to the question:

Why convert?: Organizations with sufficient compute may prefer to pretrain the target architecture directly, while compute-constrained users may simply fine-tune an existing hybrid model such as Qwen3.5[18].

In this sense, our result is primarily a demonstration that architecture conversion that retains long-context abilities from transformer layers to GDN or Mamba-3 like layers is possible even late in training, including after RL post-training, at relatively low additional cost. This may become more useful as architectural innovation outpaces pretraining cycles, allowing already-trained models to adopt newer architectures without restarting from scratch. The concept could also be particularly relevant for architectures that are substantially harder or more expensive to pretrain, such as models with deep-memory or TITANS-style[19] modules.

Looking forward, several directions are particularly promising. Beyond extending the approach to more sophisticated architectures such as looped Transformers, architectural conversion itself could be treated as a curriculum, gradually modifying the backbone rather than introducing a large architectural shift at once. More ambitiously, one could study scaling laws for architecture conversion: how the required recovery compute depends on model scale, architectural distance, context length, and the fraction of layers being replaced. Such a characterization could help turn post-hoc architecture conversion into a more predictable and general mechanism for adapting pretrained models to new computational constraints.

References

  1. Kimi Team, Moonshot AI. Kimi K3: Open Frontier Intelligence. Technical report, 2026. 2.8T-parameter MoE built on Kimi Delta Attention and Attention Residuals, 1M-token context. arxiv.org/abs/2607.24653
  2. Qwen Team, Alibaba. Qwen3.8-2.4T-A95B — 2.4T parameters, 95B active: Model card and accompanying blog post, 2026. huggingface.co/Qwen/Qwen3.8-2.4T-A95Bqwen.ai/blog?id=qwen3.8
  3. Songlin Yang, Jan Kautz, Ali Hatamizadeh. Gated Delta Networks: Improving Mamba2 with Delta Rule. ICLR 2025. arxiv.org/abs/2412.06464
  4. Qwen Team, Alibaba. Qwen3-Next —. Model card, 2025. huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct
  5. Qwen Team, Alibaba. Qwen3 Technical Report. 2025. Covers the Qwen3 family, including the 4B model used as a teacher here. arxiv.org/abs/2505.09388
  6. Xiaomi LLM-Core Team. MiMo: Unlocking the Reasoning Potential of Language Model — From Pretraining to Posttraining. 2025. MiMo-7B-RL is the RL-tuned model from this report. arxiv.org/abs/2505.07608
  7. Kaiyue Wen, Zhiyuan Li, Jason Wang, David Hall, Percy Liang, Tengyu Ma. Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective. 2024. arxiv.org/abs/2410.05192
  8. Kevin Lu et al., Thinking Machines Lab. On-Policy Distillation. 2025. thinkingmachines.ai/blog/on-policy-distillation
  9. Shengding Hu et al. MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies. 2024. arxiv.org/abs/2404.06395
  10. Etash Guha et al. OpenThoughts: Data Recipes for Reasoning Models. 2025. arxiv.org/abs/2506.04178
  11. Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, Reynold Xin. Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM. 2023. databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm
  12. Daniel Goldstein, Eric Alcaide, Janna Lu, Eugene Cheah. RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale. 2025. arxiv.org/abs/2505.03005
  13. Aviv Bick, Tobias Katsch, Nimit Sohoni, Arjun Desai, Albert Gu. Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing. 2025. arxiv.org/abs/2502.14458
  14. Mingyu Yang, Mehdi Rezagholizadeh, Guihong Li, Vikram Appia, Emad Barsoum. Zebra-Llama: Towards Extremely Efficient Hybrid Models. 2025. openreview.net/forum?id=l42UGsdrNn
  15. Yingfa Chen, Zhen Leng Thai, Zihan Zhou, Zhu Zhang, Xingyu Shen, Shuo Wang, Chaojun Xiao, Xu Han, Zhiyuan Liu. Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts. 2026. arxiv.org/abs/2601.22156
  16. Mingjie Liu et al. ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models. 2025. arxiv.org/abs/2505.24864
  17. Aakash Lahoti, Kevin Y. Li, Berlin Chen, Caitlin Wang, Aviv Bick, J. Zico Kolter, Tri Dao, Albert Gu. Mamba-3: Improved Sequence Modeling using State Space Principles. 2026. arxiv.org/abs/2603.15569
  18. Qwen Team. Qwen3.5: Towards Native Multimodal Agents. February 2026. qwen.ai/blog?id=qwen3.5
  19. Ali Behrouz, Peilin Zhong, Vahab Mirrokni. Titans: Learning to Memorize at Test Time. 2024. openreview.net/forum?id=8GjSf9Rh7Z
  20. Disen Lan, Jianbin Zheng, Yuxi Ren, Xin Xia, Xuanda Wang, Xuefeng Xiao, Xipeng Qiu, Yu Cheng. Morphing into Hybrid Attention Models. 2026. arxiv.org/abs/2606.30562
  21. Aditya Chattopadhyay, Elvis Nunez, Prannay Kaul, Benjamin Bowman, Evan Becker, Luca Zancato, David Thomas, Wei Xia, Stefano Soatto. Priming: Hybrid State Space Models From Pre-trained Transformers. 2026. arxiv.org/abs/2605.08301
  22. Yanhong Li, Songlin Yang, Shawn Tan, Mayank Mishra, Rameswar Panda, Jiawei Zhou, Yoon Kim. Distilling to Hybrid Attention Models via KL-Guided Layer Selection. 2025. arxiv.org/abs/2512.20569
  23. Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Y. Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, Ion Stoica. DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL. 2025. notion.site/DeepScaleRhuggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset
  24. ByteDance Seed, Tsinghua AIR. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. 2025. arxiv.org/abs/2503.14476huggingface.co/datasets/BytedTsinghua-SIA/DAPO-Math-17k