Q: Can we efficiently convert pretrained full-attention transformers into hybrid Gated DeltaNet (GDN) models while preserving the reasoning and long-context capabilities of the original model?
- Gap: Existing conversion methods largely recover short-context performance, while long-context retrieval and reasoning remain substantially degraded. Many methods essentially score 0 while using thinking mode on long-context reasoning benchmarks such as AIME. Some additionally rely on subsequent SFT or preference-alignment stages beyond architectural initialization and distillation.
- Method: We recover these capabilities through a three-stage recovery pipeline after architectural initialization: (i) Off-Policy Representation Alignment (102M tokens) to initialize the newly replaced GDN layers, (ii) Off-Policy Distillation (~900M tokens) to recover short-context performance and bootstrap long-context behavior, and (iii) Prolonged On-Policy Distillation (2B tokens), to recover long-context retrieval, generation, and reasoning. Throughout recovery, we use only the original pretrained full-attention transformer as opposed to a bigger model as the teacher, with no additional SFT, DPO or RLVR.
- Results: Across Qwen3-4B and MiMo-7B-RL, our proposed method OPAL (On-Policy Attention Linearization) recovers 83% and 93% of the full-attention transformer’s long-context reasoning performance respectively, and matches or surpasses its long-context retrieval performance (100% and 100.7%), achieving perfect scores in most NIAH tasks, all within ~3B total distillation tokens. Performance continues to improve as we prolong the on-policy stage, with training loss still decreasing at our final 2B-token checkpoint. The conversion pipeline also works with other architectures such as Mamba-3.
One-line summary: Prolonged On-Policy Distillation recovers long-context reasoning and retrieval alongside short-context capabilities, enabling pretrained full-attention transformers to be efficiently (< 0.01% pretraining budget) converted into hybrid GDN models while retaining much of the base full-attention transformer model’s capabilities.
Primer on Gated DeltaNets
Primer on Gated DeltaNets
A standard Transformer layer uses full causal self-attention. For a hidden state h_t \in \mathbb{R}^d, the layer produces
with q_t, k_t \in \mathbb{R}^{d_k} and v_t \in \mathbb{R}^{d_v}. At token t, the layer retains the keys and values seen so far in a KV cache, K_{\leq t} = [k_1, \ldots, k_t]^{\top} \in \mathbb{R}^{t \times d_k} and V_{\leq t} = [v_1, \ldots, v_t]^{\top} \in \mathbb{R}^{t \times d_v}. The attention output is
Every query interacts explicitly with all preceding keys. Across a sequence of length n, the total attention computation grows as O(n^2), while the KV cache grows linearly with sequence length.
A Gated DeltaNet (GDN) layer[3] instead compresses the history into a fixed-size state S_t \in \mathbb{R}^{d_v \times d_k}, rather than storing the full KV cache. As in attention, the layer produces
along with data-dependent forget and update gates
The compressed state evolves as
and the layer output is obtained directly from this state,
Instead of explicitly retaining every previous k_t and v_t, the GDN carries forward the fixed-size memory S_t, with \alpha_t controlling retention of the previous state and \beta_t controlling how strongly the new key-value association is written. Because the state size is fixed, the computational cost of a GDN layer grows as O(n) over a sequence of length n, compared with the O(n^2) cost of a full-attention layer. 11The equations above describe a single head and omit the details that make these layers work in practice; activation functions, the short causal convolutions applied to q, k and v, the normalization and output-gating layers, and the multi-head structure of both attention and GDN. fQ, fK, fV, fα and fβ stand in for the learned projections and their associated machinery.
Hybrid Models
Recent frontier models increasingly rely on mixing GDNs and full-attention transformers than choosing between them. Examples include Kimi K3[1] (Moonshot AI, 2.8T parameters, 1M-token context) and Qwen3.8[2] (Alibaba, 2.4T MoE).
A hybrid of full-attention and GDN layers can be substantially more efficient to serve than a pure attention model, particularly at long-context lengths, because most of its layers replace a KV cache that grows with every token and an attention computation that revisits the entire history at every step with a fixed-size memory and a constant per-token update.
The intuition behind the hybrid design is that attention and GDN layers provide different forms of memory. GDN layers can carry information over long spans through a compact recurrent state, making them efficient at preserving and updating context without attending over the entire history. Full-attention layers, meanwhile, retain direct access to individual past token positions, which is useful when the model needs to recover a specific detail or model an exact interaction across a long distance. Combining the two therefore offers efficient long-range memory through GDNs while preserving the precise retrieval capabilities of full-attention.
Proposed Pretrained Transformer to GDN conversion pipeline
Architecture Morphing
Morphing is the first stage, aimed at converting an off-the-shelf pretrained full-attention model into a hybrid architecture, before any training begins. Every 4th layer retains its original softmax full-attention unchanged, while the attention blocks in the remaining layers are replaced with Gated DeltaNet[3] modules matching the Qwen3-Next[4] design. We apply this transformation to two base architectures: Qwen3-4B[5] and MiMo-7B-RL[6].
Parameters modified
| component | Qwen3-4B | MiMo-7B |
|---|---|---|
| Query Q | −5.2M | −8.4M |
| Key K | +2.6M | +4.2M |
| Value V | +7.9M | +12.6M |
| Output gate Z | +10.5M | +16.8M |
| Gates α, β | +164K | +262K |
| Convolution | +33K | +33K |
| Misc | −64 | −6.0K |
| Net per layer | +15.9M | +25.5M |
| × 27 GDN layers | +430M | +687M |
| parameters | Qwen3-4B | MiMo-7B |
|---|---|---|
| Base model | 4.02B | 7.62B |
| Hybrid | 4.55B | 8.31B |
| Per full-attention layer | +10.5M | 0 |
| × 9 full-attention layers | +94M | 0 |
| Per GDN layer | +15.9M | +25.5M |
| × 27 GDN layers | +430M | +687M |
| Total added | +524M | +687M |
Component by component22 Many of these architectural choices exist to match the configuration used in Qwen3-Next — the output projection on the full-attention layers, for instance; rather than because our method requires them.
- Query q — merged: The base model has 32 query heads; while GDN has 16. Each GDN query is the average of two heads in the base model, which results in loss of information.
- Key k and Value v — replicated: The base model has 8 key/value heads shared across its 32 query heads; GDN uses 16 key and 32 value heads. Each of the base model's head is copied out across the GDN heads that map onto it.
- Output projection, layer norms, embeddings and linear layers — copied unchanged: GDN has to project its output back into the residual stream, and its output width happens to equal the attention block's, so that weight transfers verbatim in both models. Everything outside the attention block transfers the same way — the linear layers, the layer norms and the embeddings are all copied exactly, in both kinds of layer. Only the attention block is replaced, which is why none of them appear in the table.
- Output gate — new: The GDN layers scale its own output by a learned gate. At 10.5M per layer on Qwen and 16.8M on MiMo it is roughly two thirds of everything conversion adds. However, this is added to all layers including full-attention in Qwen3-4B to match Qwen3-Next while only to GDN in MiM0-7B-RL.
- α and β — new: GDN specific parameters, gated: α decides how much of the memory to keep and β how strongly to write into it. They are one small projection, about 1% of the parameters added.
- Convolution — new: A short causal convolution runs over q, k and v before the recurrence. About 33K weights per layer in either model.
- Misc: The per-head decay terms, the timestep bias and the GDN block's own norm come to under 200 weights. Qwen3's per-head query and key norms (256 weights), and MiMo's query, key and value biases (6.1K, since MiMo's Qwen2-family attention carries biases where Qwen3's does not) are removed.
Initialization: Off-Policy Representation Alignment
This is the first stage of our recovery pipeline. Throughout all stages, we use the original base model as the teacher and the morphed hybrid model as the student. The goal of this stage is to provide a strong initialization for subsequent training. Concretely, we train each student GDN layer to match the output of the corresponding teacher full-attention layer, while feeding both layers the hidden representation produced by the teacher’s preceding layer (hence called off-policy). This encourages the morphed model GDN layers to mimic the teacher full-attention layer's outputs by minimizing distribution shift between the teacher and student representations.
-
Objective: A single teacher forward per batch yields every hidden state,
where h^{T}_{l} is the input to layer l and
h^{T}_{l+1} is the output.
h^{T}_{0} = \mathrm{Embed}(x), \qquad h^{T}_{l+1} = f^{T}_{l}\!\left(h^{T}_{l}\right)Let f^{S}_{l} denote the functional map for the layer l of the student model we minimize\mathcal{L} \;=\; \sum_{l \,\in\, \mathrm{GDN}} \;\left\lVert\, f^{S}_{l}\!\left(h^{T}_{l}\right) \;-\; h^{T}_{l+1} \,\right\rVert_{2}^{2}Every layer's input is the teacher's h^{T}_{l}, not the student's output from the preceding layer, so \partial\mathcal{L}/\partial\theta_{l} depends on layer l alone. The objective decomposes into 27 independent layer-wise feature-matching problems, each minimizing a squared reconstruction error. The gradient flows through the inherited q, k, v, output projection, and the parameters with no counterpart in attention: the forget gate \alpha, the update gate \beta, the output gate z, the short convolution and the decay terms — which is 42.1M per layer on Qwen3-4B and 67.4M on MiMo-7B, or 1.14B and 1.82B across the 27. Everything outside those blocks is frozen. The reason for using a layerwise loss here is to avoid compounding approximation errors caused from distribution shift at the input across the depth of the network
-
Data and Optimizer:
Data: Both arms use the same 102M-token mix: 70% FineWeb-Edu, 20% Python code, 10% open-web-math. Optimizer: AdamW, no weight decay, betas 0.9 and 0.95, gradients clipped at global norm 1.0. Learning rate warmup over 3% of steps then cosine decay to 10% with peak LR 3e-3 on Qwen3-4B, 1e-3 on MiMo-7B with a global batch of 32K tokens on Qwen3-4B (~3,050 steps) and 16K on MiMo-7B (~6,100 steps). Weight decay is off because there is a fixed right answer here. Decay nudges every weight a little toward zero at each step — it is a regulariser, meant to keep a model from fitting noise in its training data too closely. But these blocks are being fitted to a target function that already exists, starting from weights copied out of a working model. There is no noise to guard against, so shrinking them only drags the fit away from the answer.
Short context recovery: Off-Policy Distillation
Stage 1 calibrates each GDN block independently, using hidden-state inputs generated by the teacher model. While this provides a strong layerwise initialization, it does not account for interactions between the newly introduced GDN blocks: each block is optimized on teacher-generated representations and not the representations it will receive from preceding student layers at inference time. Consequently, approximation errors and representation shifts may accumulate when the student is composed end-to-end.
Stage 2 is the first stage that trains the language model end-to-end distilling the teacher’s output distribution into the complete student model, thereby encouraging globally consistent behavior across the full network. It runs in two substages: first at short context, then at the teacher's full window.
-
Objective: Both student and teacher models use same text as the input, and at every position the
student's output distribution is trained to match the teacher distribution by minimizing the forward KL Loss:
\mathcal{L}_{2} \;=\; \mathbb{E}_{x \sim \mathcal{D}}\left[\, \sum_{t} \mathrm{KL}\!\left(p^{T}(\cdot \mid x_{<t}) \;\middle\|\; p^{S}(\cdot \mid x_{<t})\right) \right]Forward KL, with the teacher as the reference distribution, is mode-covering: it penalizes the student for assigning too little probability wherever the teacher places mass, encouraging the student to cover the full range of outputs considered plausible by the teacher. The training data consists of ordinary pretraining-style packed corpus text rather than prompt–response pairs or teacher-generated samples. The teacher does not generate any tokens; instead, it provides its predictive distribution at each position in the fixed corpus. Because the trajectories are drawn from an external data distribution rather than sampled from the teacher, this stage is off-policy.
Stage 2a: Short-Context Distillation
-
Data and Optimizer:
Data: 600M tokens of pretraining-style text — web, code and maths — packed into 4K-token sequences, one pass, no repetition. Optimizer: AdamW, betas 0.9 and 0.95, gradients clipped at 1.0. Warmup over 3% of steps, then cosine decay to a 10%. Global batch 2M tokens. Both the full-attention and GDN layers use a cosine learning rate schedule decayed to 10% over the training horizon, while the peak learning rate of the full-attention layers is fixed to 10% of that of the GDN layers. The peak learning rate used for the GDN layers is 1.5e-4 for Qwen3-4B and 5e-5 for MiMo-7B.
Stage 2b: Long-Context Distillation
-
Data and Optimizer:
Data: 294M tokens packed into 32K-token sequences, chosen so that those sequences are genuinely long documents — including books, long web pages, and mathematics content and not short documents concatenated together. The mix is approximately 70% long documents and 30% short-context replay, with the replay included to prevent short-context behavior from regressing while the recurrent state learns to retain information over longer inputs. The 32K context length is set by the teacher's own context window — the teacher cannot provide supervision beyond what it can itself read. Optimizer: AdamW, betas 0.9 and 0.95, clip 1.0, 3% warmup, cosine decay to 10%, global batch 2M tokens. Peak LR 1e-4 for the GDN group on Qwen3-4B and 2.5e-5 on MiMo-7B, and the peak LR of transformer layers is 10% of the GDN layers.
Many prior approaches keep the retained full-attention layers frozen throughout training while the GDN layers learn how to update and use the fixed size memory state. However, we freeze them only during stage 1, while the newly introduced blocks are being initialized. In stage 3, which is the longest stage of our training pipeline, the retained attention layers are updated jointly with the rest of the model.
This is important because the full-attention layer in pure transformers and hybrid models differ in the role those layers must play. Before conversion, every attention layer contributes directly to long-range interaction at its own depth. After conversion, the retained full-attention layers become the only points in the network where information can be accessed beyond the compressed state, and must therefore support long-range retrieval on behalf of the converted layers as well. Their pretrained weights were not optimized for this altered role, so they must adapt during recovery training. We observe that the best results are obtained when Stage 2 serves as a gradual bridge between stages 1 and 3, jointly updating the full model while using a lower learning rate for the retained full-attention layers so that they adapt on a slower timescale than the newly introduced GDN blocks.
Long-Context Recovery: Prolonged On-Policy Distillation33The name and the philosophy are inspired by ProRL[16], which extends the RL post-training phase well beyond the usual budget because reasoning capability continues to improve rather than saturating early.
Stage 3 is the final and most important stage of the recovery pipeline. As in the earlier stages, the original base model serves as the teacher and the morphed hybrid as the student, with the objective of recovering the capabilities lost during conversion.
This is also the stage where long-context reasoning is restored. After Stage 2, the student can retrieve information from long inputs, but it still struggles to reason effectively over long contexts. Its AIME accuracy remains at 0, while MATH-500 scores under a 32K-token budget are only around 2.0%, substantially worse than performance with shorter 1–4K token budgets and thinking disabled. In other words, giving the model more room to reason actually hurts performance without stage 3.
This gap between long-context retrieval and long-context generation comes from a difference in distribution shift. During retrieval, nearly all of the tokens the model conditions on are provided in the input, so teacher-forced training covers the relevant distribution reasonably well. During long-form generation, however, the model repeatedly conditions on its own outputs. After only a few thousand generated tokens, it can enter regions of the sequence distribution that are rarely, if ever, represented in static training corpora. Off-policy training would need to anticipate and cover this space in advance, which is impractical for open-ended generation. Training directly on the student's own generated trajectories instead exposes the model to exactly the states it will encounter at inference time. Stage 3 additionally also improves long form retrieval performance.
-
Objective: The student generates a continuation from a prompt. The frozen
teacher then scores those exact tokens, and the loss is the divergence between the two
distributions at each generated position:
\mathcal{L}_{3} \;=\; \mathbb{E}_{y \sim \pi^{S}}\left[\, \sum_{t} \mathrm{KL}\!\left(p^{S}(\cdot \mid y_{<t}) \;\middle\|\; p^{T}(\cdot \mid y_{<t})\right) \right]Reverse KL is mode-seeking, and it punishes the student for assigning high probability where the teacher does not. The prompt is context only; loss is taken on generated positions, so generated tokens and loss tokens are the same count. Generation length starts at 512 and doubles whenever the smoothed divergence stops improving, out to 16K and later 32K, so the model is supervised at depths it will actually be used at.
- Why does Stage 3 have to be prolonged: On-policy distillation is most effective when the teacher and student induce reasonably similar output distributions[8] — that is, the student must produce good trajectories often enough that there is some positive signal for the teacher to reinforce. A well trained model such as off the shelf Qwen3-4B checkpoint used as a student starts inside that regime and converges quickly. Our model post stage 2 does not. The student is very weak at long-context generation when this stage 3 begins, so good trajectories are rare, and the teacher is the base model itself rather than a larger one as is the case typically in standard on policy distillation. The same-size teacher is a deliberate constraint rather than a limitation we worked around. A stronger teacher could introduce capabilities that the original pre-Stage-0 model never possessed, turning the process into knowledge transfer rather than capability recovery. Hence as a result of this weak student - weak teacher combination, Stage 3 requires prolonged training for effective recovery.
- Pretraining-like step: Each optimizer step consumes a fixed budget of 128K generated tokens — the same for Qwen3-4B and MiMo-7B — split across the trainer processes. An update is defined by a token count, as in pretraining, and not by a number of prompts as in post-training distillation/RL. How many sequences that budget buys is whatever the current generation length allows. The learning rate is scheduled on consumed tokens rather than on steps, so the horizon ladder described below can change sequence length without disturbing it.
- Generation-length curriculum: Generation length starts at 512 tokens and doubles every phase to 16K over the first 600M tokens, then to 32K over the next 400M. The generation length is increased when the loss plateaus and doesn't make any meaningful progress. Within each phase, we track the lowest smoothed KL seen so far. A 2% improvement resets the counter and sets a new record; otherwise, the counter increments. After 30 steps without sufficient improvement, the generation cap doubles and all phase statistics reset. This curriculum is more efficient than starting with long rollouts: with a fixed token budget, longer generations mean fewer samples, noisier gradients, and slower sampling. Early in training, the student also struggles to stay coherent over long horizons, making much of the extra text useless.
-
Data:
Data: The data used for the first 1B tokens in this stage consists of prompts drawn from approximately 65% OpenThoughts[10], 20% Dolly-15k[11], and 15% synthetic long-context retrieval tasks, with all answers generated on-policy by the student model. For Stage 3 (1B–2B tokens), OPAL-MiMo replaces the mathematical reasoning data with a mixture of 43.5% DeepScaleR[23] and 16.5% DAPO-17K[24], while increasing the retrieval data proportion to 20%. OPAL-Qwen3 retains the same data mixture used during the first 1B tokens. During inference time retrieval, the model must recover a fact from a long-context in the prompt and retain it while continuing to generate its response. These tasks are introduced only once the generation cap reaches 1024 tokens, since shorter rollouts do not provide enough generation length for retrieval and subsequent use of the retrieved information to become meaningful. Sampling is performed at temperature 1.0, so the student is trained on the same distribution from which its answers are generated. - Optimizer and learning-rate schedule: For the prolonged on-policy distillation phase, we use a Warmup-Stable-Decay (WSD) learning-rate schedule[9] (Figure 1). WSD is predominantly a pretraining LR schedule. It belongs to a family of infinite LR schedules designed for continual pretraining. We adopt it here specifically because our post-training regime is unusually long and is extended incrementally in blocks of 100–200M tokens, without knowing in advance when the loss will saturate. WSD is well suited to this setting because it defines an effectively open-ended schedule: the stable phase can be extended cheaply as additional training is added, and the final decay can be triggered only once we decide to stop. In contrast, cosine decay assumes a fixed training horizon and generally requires the schedule to be specified from the outset. This flexibility is the main reason we prefer WSD despite cosine decay being the standard choice in post-training. We parameterize the schedule by the number of tokens trained on rather than the number of generated sequences or optimization steps, since rollout lengths vary and an equal generation or step budget can correspond to substantially different amounts of training data. All GDN and full-attention layers are trained using the same learning-rate schedule. We use AdamW with a peak learning rate of 2e-5 for Qwen3-4B and 1e-5 for MiMo-7B, and fix the global batch size to 2M tokens. The learning rate is linearly warmed up for 150 steps, while the decay phase is reconfigured at each training increment and spans approximately 10% of the cumulative Stage 3 steps reached at that point. All results reported in this work use 2B tokens of Stage 3 training unless stated. We note that the loss has not yet saturated at 2B tokens (Figure 3), suggesting that extending Stage 3 further is likely to yield additional gains.
-
Asynchronous sampling:
Asynchronous sampling: Generation and training run concurrently rather than in alternation. We maintain multiple trainers and a single sampler copy. Sampler processes hold an inference copy of the student and do nothing but produce continuations; trainer processes hold the training copy and do only training via forward and backward pass on generations. The two are coupled by a queue of rollouts and a weight publish: after every optimizer step the trainer writes out its current weights, the samplers hot-load them. The process is asynchronous in the sense that the sampler does not wait for the trainer; multiple trainers per sampler ensure that the sampler never idles through backward passes due to a long queue.
Other Methods
Other Methods
Existing methods for converting pretrained full-attention transformer layers into linear attention layers typically use some combination of layerwise alignment (Stage 1), end-to-end off-policy distillation (Stage 2)[12][13][14][15][20], and continual pretraining on additional data[21]. Some methods additionally perform some form of layer selection for linearization before training[15][20][22].
4In distillation, the data/prompts mediate the transfer between the teacher and student; therefore, from an information-theoretic perspective, the procedure is not entirely free of knowledge transfer from the data itself.We intentionally omit the continual pretraining stage to avoid transferring additional capabilities beyond those already present in the language model.4 In particular, if the morphed hybrid student were to exactly match the teacher distribution, the objectives in all three stages would reduce to zero. The continual learning stage sometimes uses 30–50x more tokens than our whole pipeline[21], often approaching modern midtraining budgets.
Main results
Commonsense Reasoning Suite : Short Context
Each linearized model is compared against its corresponding teacher. The final column reports the average across all six tasks; percentages in parentheses denote performance retained relative to the teacher average. This part is easy to recover even without Stage 3. We donot beat the baselines here and pay the short context tax for long-context performance. We think that Commonsense performance can be recovered somewhat cheaply with additional SFT as done by some of the baselines.
| Model | PIQA | HellaSwag | ARC-E | ARC-C | Winogrande | MMLU | Avg. |
|---|---|---|---|---|---|---|---|
| RADLADS[12] | |||||||
| Qwen2.5-7B-Instruct | 80.2 | 80.5 | 80.9 | 55.0 | 70.5 | 74.3 | 73.6 |
| QRWKV6-7B-Instruct | 79.9 | 78.9 | 79.3 | 55.9 | 71.4 | 64.2 | 71.6 (97.3%) |
| Llamba[13] | |||||||
| Llama-3.2-1B-Instruct | 74.9 | 61.6 | 63.9 | 37.6 | 61.4 | 46.0 | 57.6 |
| Llamba-1B | 73.7 | 61.8 | 65.2 | 37.5 | 61.9 | 31.5 | 55.3 (96.0%) |
| Llamba[13] | |||||||
| Llama-3.2-3B-Instruct | 76.8 | 71.6 | 71.0 | 46.4 | 68.7 | 60.7 | 65.9 |
| Llamba-3B | 78.0 | 74.0 | 73.6 | 46.2 | 71.7 | 50.1 | 65.6 (99.6%) |
| Zebra-Llama[14] | |||||||
| Llama-3.2-3B-Instruct | 76.8 | 71.6 | 71.0 | 46.4 | 68.7 | 60.7 | 65.9 |
| Zebra-Llama-3B | 77.3 | 72.7 | 75.8 | 52.3 | 65.7 | 51.7 | 65.9 (100.1%) |
| Zebra-Llama[14] | |||||||
| Llama-3.1-8B-Instruct | 81.4 | 79.5 | 79.9 | 55.6 | 73.6 | 68.4 | 73.1 |
| Zebra-Llama-8B | 79.9 | 79.0 | 77.5 | 57.6 | 71.3 | 59.1 | 70.7 (96.8%) |
| HALO[15] | |||||||
| Qwen3-1.7B | 72.0 | 60.4 | 69.5 | 43.0 | 61.6 | 60.2 | 61.1 |
| HypeNet-2B | 72.0 | 57.2 | 67.3 | 42.7 | 61.7 | 39.5 | 56.7 (92.8%) |
| HALO[15] | |||||||
| Qwen3-4B | 74.9 | 68.5 | 78.5 | 53.8 | 65.8 | 70.1 | 68.6 |
| HypeNet-5B | 76.3 | 66.1 | 75.9 | 50.6 | 67.2 | 52.4 | 64.8 (94.4%) |
| OPAL | |||||||
| Qwen3-4B | 74.9 | 68.5 | 78.5 | 53.8 | 65.8 | 70.1 | 68.6 |
| OPAL-Qwen3 (ours) | 72.3 | 58.5 | 65.0 | 41.6 | 63.0 | 57.2 | 59.6 (86.9%) |
| OPAL | |||||||
| MiMo-7B-RL-0530 | 74.5 | 64.9 | 68.1 | 45.0 | 61.6 | 57.2 | 61.9 |
| OPAL-MiMo (ours) | 72.5 | 59.9 | 62.7 | 39.5 | 58.8 | 54.6 | 58.0 (93.7%) |
Needle-in-a-haystack: Long Form Retrieval
Retrieval accuracy across context lengths, grouped by task type. The average is the macro-average over all twelve task–length combinations; percentages in parentheses denote performance relative to the corresponding teacher. One can unlock non trivial performance here without stage 3. Some baselines such as Zebra-Llama[14] perform well in shorter contexts (4–8K) but severely underperform at longer contexts (16-32K) and in the multi-key/query settings compared to the corresponding full-attention model, which has near perfect retrieval scores across the board. OPAL on the other hand has scores that match the corresponding full-attention model.
| Model | Single needle | Multi-key | Multi-query | Avg. | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 4K | 8K | 16K | 32K | 4K | 8K | 16K | 32K | 4K | 8K | 16K | 32K | ||
| RADLADS[12] | |||||||||||||
| Qwen2.5-7B-Instruct | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 |
| QRWKV6-7B-Instruct | 0.0 | 0.0 | 0.0 | 0.0 | 1.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.1 (0.1%) |
| Llamba[13] | |||||||||||||
| Llama-3.2-1B-Instruct | 100.0 | 100.0 | 100.0 | 100.0 | 94.0 | 93.0 | 71.0 | 67.0 | 98.5 | 94.2 | 93.2 | 96.5 | 92.3 |
| Llamba-1B | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 (0.0%) |
| Llamba[13] | |||||||||||||
| Llama-3.2-3B-Instruct | 100.0 | 100.0 | 100.0 | 100.0 | 98.0 | 95.0 | 94.0 | 87.0 | 100.0 | 99.2 | 87.5 | 68.5 | 94.1 |
| Llamba-3B | 2.0 | 0.0 | 0.0 | 0.0 | 3.0 | 3.0 | 0.0 | 0.0 | 0.8 | 0.0 | 0.0 | 0.0 | 0.7 (0.8%) |
| Zebra-Llama[14] | |||||||||||||
| Llama-3.2-3B-Instruct | 100.0 | 100.0 | 100.0 | 100.0 | 98.0 | 95.0 | 94.0 | 87.0 | 100.0 | 99.2 | 87.5 | 68.5 | 94.1 |
| Zebra-Llama-3B | 85.0 | 61.0 | 52.0 | 47.0 | 55.0 | 39.0 | 28.0 | 10.0 | 64.2 | 24.2 | 23.0 | 3.5 | 41.0 (43.6%) |
| Zebra-Llama[14] | |||||||||||||
| Llama-3.1-8B-Instruct | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.0 | 99.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.8 |
| Zebra-Llama-8B | 100.0 | 98.0 | 96.0 | 74.0 | 73.0 | 40.0 | 40.0 | 27.0 | 86.8 | 70.0 | 66.5 | 43.0 | 67.9 (68.0%) |
| HALO[15] | |||||||||||||
| Qwen3-1.7B | 100.0 | 100.0 | 100.0 | 100.0 | 99.0 | 100.0 | 99.0 | 99.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.8 |
| HypeNet-2B | 73.0 | 76.0 | 66.0 | 73.0 | 28.0 | 22.0 | 24.0 | 18.0 | 27.3 | 19.0 | 18.5 | 15.5 | 38.4 (38.5%) |
| HALO[15] | |||||||||||||
| Qwen3-4B | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.9 |
| HypeNet-5B | 88.0 | 90.0 | 90.0 | 96.0 | 34.0 | 37.0 | 33.0 | 28.0 | 49.5 | 55.8 | 49.2 | 45.0 | 58.0 (58.0%) |
| OPAL | |||||||||||||
| Qwen3-4B | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.8 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 |
| OPAL-Qwen3 (ours) | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.6 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 (100.0%) |
| OPAL | |||||||||||||
| MiMo-7B-RL-0530 | 100.0 | 100.0 | 100.0 | 100.0 | 99.6 | 99.0 | 98.8 | 92.8 | 99.9 | 99.7 | 98.3 | 76.2 | 97.0 |
| OPAL-MiMo (ours) | 100.0 | 100.0 | 100.0 | 100.0 | 96.8 | 96.8 | 94.8 | 83.8 | 100.0 | 100.0 | 100.0 | 100.0 | 97.7 (100.7%) |
Mathematical reasoning : Long-Context Generation
This is the stage where Prolonged On-Policy Distillation yields the largest gains, with the improvements over baselines being most pronounced on harder benchmarks such as AIME, where most baselines fail to recover any performance. In fact the numbers reported below correspond to premature early stopped checkpoints. Extended training still shows signs of decreasing training loss.
Eval details
- GSM: non-thinking mode with greedy decoding, generation length of 1024 under zero-shot prompting. GSM symbolic numbers are an average of the "main", "p1" and "p2" splits weighted by number of samples.
- MATH-500: 32K generation length with thinking mode, temperature 0.7; reports avg@8 numbers.
- AIME: thinking mode, temperature 0.6, generation length 32K; reports pass@1 (avg@8) and pass@8 (temp 0.6).
| Model | GSM8K | GSM-Plus | GSM-Symbolic | MATH-500 | AIME’24 | AIME’25 | Avg. | ||
|---|---|---|---|---|---|---|---|---|---|
| pass@1 | pass@8 | pass@1 | pass@8 | ||||||
| RADLADS[12] | |||||||||
| Qwen2.5-7B-Instruct | 91.9 | 72.7 | 83.4 | 76.4 | 10.4 | 26.7 | 7.5 | 23.3 | 49.0 |
| QRWKV6-7B-Instruct | 27.4 | 16.8 | 9.3 | 24.0 | 0.4 | 3.3 | 0.4 | 3.3 | 10.6 (21.6%) |
| Llamba[13] | |||||||||
| Llama-3.2-1B-Instruct | 38.5 | 24.9 | 26.2 | 25.8 | 0.8 | 6.7 | 0.0 | 0.0 | 15.4 |
| Llamba-1B | 28.1 | 12.9 | 13.0 | 8.6 | 0.0 | 0.0 | 0.0 | 0.0 | 7.8 (51.0%) |
| Llamba[13] | |||||||||
| Llama-3.2-3B-Instruct | 71.1 | 58.3 | 59.5 | 39.0 | 2.5 | 13.3 | 0.0 | 0.0 | 30.5 |
| Llamba-3B | 50.7 | 27.5 | 24.8 | 7.8 | 0.4 | 3.3 | 0.0 | 0.0 | 14.3 (47.0%) |
| Zebra-Llama[14] | |||||||||
| Llama-3.2-3B-Instruct | 71.1 | 58.3 | 59.5 | 39.0 | 2.5 | 13.3 | 0.0 | 0.0 | 30.5 |
| Zebra-Llama-3B | 62.7 | 38.8 | 36.7 | 20.0 | 0.4 | 3.3 | 0.0 | 0.0 | 20.2 (66.4%) |
| Zebra-Llama[14] | |||||||||
| Llama-3.1-8B-Instruct | 86.1 | 67.0 | 73.2 | 48.6 | 2.9 | 10.0 | 0.0 | 0.0 | 36.0 |
| Zebra-Llama-8B | 61.3 | 40.3 | 35.3 | 24.4 | 0.8 | 6.7 | 0.4 | 3.3 | 21.6 (59.9%) |
| HALO[15] | |||||||||
| Qwen3-1.7B | 83.5 | 63.3 | 65.0 | 74.6 | 13.8 | 30.0 | 10.4 | 23.3 | 45.5 |
| HypeNet-2B | 1.1 | 3.8 | 1.2 | 4.4 | 0.0 | 0.0 | 0.8 | 3.3 | 1.8 (4.0%) |
| HALO[15] | |||||||||
| Qwen3-4B | 92.3 | 72.8 | 83.2 | 85.4 | 25.8 | 43.3 | 22.1 | 43.3 | 58.5 |
| HypeNet-5B | 14.9 | 15.8 | 8.7 | 11.8 | 0.0 | 0.0 | 0.0 | 0.0 | 6.4 (10.9%) |
| OPAL | |||||||||
| Qwen3-4B | 92.3 | 72.8 | 83.2 | 96.4 | 73.8 | 86.7 | 63.7 | 80.0 | 81.1 |
| OPAL-Qwen3 (ours) | 83.2 | 63.8 | 65.9 | 92.0 | 52.5 | 80.0 | 40.4 | 63.3 | 67.6 (83.4%) |
| OPAL | |||||||||
| MiMo-7B-RL-0530 | 80.2 | 61.6 | 71.6 | 97.0 | 73.3 | 83.3 | 67.9 | 83.3 | 77.3 |
| OPAL-MiMo (ours) | 84.8 | 63.5 | 65.7 | 93.8 | 65.4 | 83.3 | 47.5 | 73.3 | 72.2 (93.4%) |
How reasoning returns, stage by stage
The tables above compare finished models. The tables below trace the same models across the pipeline, which is where the long-context reasoning improvement is visible.
Long-form retrieval recovers faster and is essentially saturated after the first 200M tokens.
| Checkpoint | MATH-500 | AIME’24 think @32K | AIME’25 think @32K | |||
|---|---|---|---|---|---|---|
| think @32K | no-think @4K | pass@1 | pass@8 | pass@1 | pass@8 | |
| Teacher (Qwen3-4B) | 96.4 | 84.2 | 73.8 | 86.7 | 63.7 | 80.0 |
| Stage 1 | 0.0 | 6.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| Stage 2a | 3.4 | 42.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| Stage 2b | 2.0 | 40.6 | 0.0 | 0.0 | 0.0 | 0.0 |
| OPD 200M | 83.6 | 68.1 | 23.3 | 50.0 | 18.8 | 33.3 |
| OPD 400M | 87.0 | 70.2 | 33.3 | 63.3 | 31.7 | 46.7 |
| OPD 600M | 89.6 | 71.0 | 36.7 | 63.3 | 31.2 | 50.0 |
| OPD 800M | 89.3 | 73.2 | 37.1 | 66.7 | 32.1 | 50.0 |
| OPD 1B | 90.5 | 74.4 | 45.8 | 73.3 | 35.8 | 56.7 |
| OPD 1.2B | 90.8 | 76.6 | 43.8 | 70.0 | 38.3 | 60.0 |
| OPD 1.4B | 91.1 | 75.4 | 45.0 | 76.7 | 40.8 | 63.3 |
| OPD 1.6B | 91.0 | 75.9 | 48.3 | 73.3 | 42.1 | 66.7 |
| OPD 1.8B | 91.6 | 75.5 | 50.4 | 76.7 | 40.8 | 63.3 |
| OPD 2B | 92.0 | 76.9 | 52.5 | 80.0 | 40.4 | 63.3 |
MiMo-7B-RL — mathematical reasoning across stages
| Checkpoint | MATH-500 | AIME’24 think @32K | AIME’25 think @32K | |||
|---|---|---|---|---|---|---|
| think @32K | no-think @4K | pass@1 | pass@8 | pass@1 | pass@8 | |
| Teacher (MiMo-7B-RL) | 97.0 | 71.2 | 73.3 | 83.3 | 67.9 | 83.3 |
| Stage 1 | 0.8 | 1.8 | 0.0 | 0.0 | 0.0 | 0.0 |
| Stage 2a | 20.2 | 16.4 | 0.4 | 3.3 | 2.1 | 3.3 |
| Stage 2b | 17.0 | 16.4 | 0.0 | 0.0 | 1.7 | 3.3 |
| OPD 200M | 88.2 | 64.4 | 35.4 | 63.3 | 28.7 | 46.7 |
| OPD 400M | 90.6 | 68.4 | 47.5 | 76.7 | 30.4 | 56.7 |
| OPD 600M | 92.8 | 68.0 | 51.2 | 76.7 | 35.8 | 56.7 |
| OPD 800M | 94.0 | 65.6 | 52.5 | 73.3 | 40.4 | 63.3 |
| OPD 1B | 94.4 | 66.4 | 58.8 | 73.3 | 40.0 | 60.0 |
| OPD 1.2B | 94.8 | 70.8 | 56.2 | 80.0 | 42.1 | 60.0 |
| OPD 1.4B | 95.4 | 69.2 | 62.1 | 83.3 | 39.2 | 63.3 |
| OPD 1.6B | 93.8 | 68.2 | 62.1 | 83.3 | 47.5 | 73.3 |
| OPD 1.8B | 93.4 | 70.2 | 63.3 | 83.3 | 49.2 | 73.3 |
| OPD 2B | 93.8 | 70.8 | 65.4 | 83.3 | 47.5 | 73.3 |
Converting to Mamba-3 Hybrid
Converting to Mamba-3 Hybrid
5We do not intend to draw conclusions about whether one target architecture is inherently easier to convert to, or performs better after conversion, than another since we do not separately optimize an extensive set of hyperparameters for each architecture. More generally, however, we observe that conversions involving fewer randomly initialized parameters after morphing tend to be easier, suggesting that the pipeline may benefit from an architectural conversion curriculum that introduces changes to the backbone progressively rather than all at once (see discussion).The same OPAL conversion pipeline is used with Mamba-3[17] in place of GDN as the recurrent layer, producing OPAL-Qwen3-Mamba and OPAL-MiMo-Mamba.5
| OPD Tokens | MATH-500 | AIME’24 | AIME’25 | |||
|---|---|---|---|---|---|---|
| GDN | Mamba3 | GDN | Mamba3 | GDN | Mamba3 | |
| Qwen3-4B | ||||||
| 200M | 83.6 | 87.2 | 23.3 | 26.2 | 18.8 | 26.7 |
| 400M | 87.0 | 87.6 | 33.3 | 35.4 | 31.7 | 32.5 |
| 600M | 89.6 | 89.2 | 36.7 | 38.8 | 31.2 | 35.4 |
| 800M | 89.3 | 89.6 | 37.1 | 39.2 | 32.1 | 35.0 |
| 1B | 90.5 | 91.0 | 45.8 | 50.0 | 35.8 | 35.0 |
| MiMo-7B-RL-0530 | ||||||
| 200M | 88.2 | 87.8 | 35.4 | 37.1 | 28.7 | 25.0 |
| 400M | 90.6 | 90.5 | 47.5 | 47.1 | 30.4 | 30.0 |
| 600M | 92.8 | 91.6 | 51.2 | 51.2 | 35.8 | 39.6 |
| 800M | 94.0 | 92.4 | 52.5 | 55.0 | 40.4 | 41.2 |
| 1B | 94.4 | 92.9 | 58.8 | 57.5 | 40.0 | 40.0 |
| Model | PIQA | HellaSwag | ARC-E | ARC-C | Winogrande | MMLU | Avg. |
|---|---|---|---|---|---|---|---|
| Qwen3-4B | |||||||
| Teacher | 74.9 | 68.5 | 78.5 | 53.8 | 65.8 | 70.1 | 68.6 |
| OPAL-Qwen3-Mamba | 73.4 | 61.1 | 68.4 | 45.2 | 62.4 | 59.8 | 61.7 (89.9%) |
| MiMo-7B-RL-0530 | |||||||
| Teacher | 74.5 | 64.9 | 68.1 | 45.0 | 61.6 | 57.2 | 61.9 |
| OPAL-MiMo-Mamba | 72.6 | 59.7 | 60.3 | 40.4 | 58.6 | 51.4 | 57.2 (92.4%) |
| Model | Single needle | Multi-key | Multi-query | Avg. | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 4K | 8K | 16K | 32K | 4K | 8K | 16K | 32K | 4K | 8K | 16K | 32K | ||
| Qwen3-4B | |||||||||||||
| Teacher | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.8 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 |
| OPAL-Qwen3-Mamba | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.8 | 99.6 | 99.4 | 100.0 | 100.0 | 100.0 | 100.0 | 99.9 (99.9%) |
| MiMo-7B-RL-0530 | |||||||||||||
| Teacher | 100.0 | 100.0 | 100.0 | 100.0 | 99.6 | 99.0 | 98.8 | 92.8 | 99.9 | 99.7 | 98.3 | 76.2 | 97.0 |
| OPAL-MiMo-Mamba | 100.0 | 100.0 | 100.0 | 100.0 | 98.0 | 99.0 | 95.0 | 87.0 | 100.0 | 100.0 | 100.0 | 100.0 | 98.3 (101.3%) |
Discussion
We establish prolonged on-policy distillation as an effective approach for converting trained Transformer layers into efficient architectures such as Mamba-3 and Gated DeltaNet, while recovering both short- and long-context capabilities. Our method requires no additional teacher beyond the original full-attention Transformer and can successfully hybridize off-the-shelf pretrained models, including those that have undergone extensive RL post-training, with only a modest additional training budget of roughly 3B tokens. Moreover, performance continues to improve as the on-policy stage is prolonged, with gains remaining unsaturated at the end of our training budget. This brings us to the question:
Why convert?: Organizations with sufficient compute may prefer to pretrain the target architecture directly, while compute-constrained users may simply fine-tune an existing hybrid model such as Qwen3.5[18].
In this sense, our result is primarily a demonstration that architecture conversion that retains long-context abilities from transformer layers to GDN or Mamba-3 like layers is possible even late in training, including after RL post-training, at relatively low additional cost. This may become more useful as architectural innovation outpaces pretraining cycles, allowing already-trained models to adopt newer architectures without restarting from scratch. The concept could also be particularly relevant for architectures that are substantially harder or more expensive to pretrain, such as models with deep-memory or TITANS-style[19] modules.
Looking forward, several directions are particularly promising. Beyond extending the approach to more sophisticated architectures such as looped Transformers, architectural conversion itself could be treated as a curriculum, gradually modifying the backbone rather than introducing a large architectural shift at once. More ambitiously, one could study scaling laws for architecture conversion: how the required recovery compute depends on model scale, architectural distance, context length, and the fraction of layers being replaced. Such a characterization could help turn post-hoc architecture conversion into a more predictable and general mechanism for adapting pretrained models to new computational constraints.
References
- Kimi Team, Moonshot AI. Kimi K3: Open Frontier Intelligence. Technical report, 2026. 2.8T-parameter MoE built on Kimi Delta Attention and Attention Residuals, 1M-token context. arxiv.org/abs/2607.24653
- Qwen Team, Alibaba. Qwen3.8-2.4T-A95B — 2.4T parameters, 95B active: Model card and accompanying blog post, 2026. huggingface.co/Qwen/Qwen3.8-2.4T-A95Bqwen.ai/blog?id=qwen3.8
- Songlin Yang, Jan Kautz, Ali Hatamizadeh. Gated Delta Networks: Improving Mamba2 with Delta Rule. ICLR 2025. arxiv.org/abs/2412.06464
- Qwen Team, Alibaba. Qwen3-Next —. Model card, 2025. huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct
- Qwen Team, Alibaba. Qwen3 Technical Report. 2025. Covers the Qwen3 family, including the 4B model used as a teacher here. arxiv.org/abs/2505.09388
- Xiaomi LLM-Core Team. MiMo: Unlocking the Reasoning Potential of Language Model — From Pretraining to Posttraining. 2025. MiMo-7B-RL is the RL-tuned model from this report. arxiv.org/abs/2505.07608
- Kaiyue Wen, Zhiyuan Li, Jason Wang, David Hall, Percy Liang, Tengyu Ma. Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective. 2024. arxiv.org/abs/2410.05192
- Kevin Lu et al., Thinking Machines Lab. On-Policy Distillation. 2025. thinkingmachines.ai/blog/on-policy-distillation
- Shengding Hu et al. MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies. 2024. arxiv.org/abs/2404.06395
- Etash Guha et al. OpenThoughts: Data Recipes for Reasoning Models. 2025. arxiv.org/abs/2506.04178
- Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, Reynold Xin. Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM. 2023. databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm
- Daniel Goldstein, Eric Alcaide, Janna Lu, Eugene Cheah. RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale. 2025. arxiv.org/abs/2505.03005
- Aviv Bick, Tobias Katsch, Nimit Sohoni, Arjun Desai, Albert Gu. Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing. 2025. arxiv.org/abs/2502.14458
- Mingyu Yang, Mehdi Rezagholizadeh, Guihong Li, Vikram Appia, Emad Barsoum. Zebra-Llama: Towards Extremely Efficient Hybrid Models. 2025. openreview.net/forum?id=l42UGsdrNn
- Yingfa Chen, Zhen Leng Thai, Zihan Zhou, Zhu Zhang, Xingyu Shen, Shuo Wang, Chaojun Xiao, Xu Han, Zhiyuan Liu. Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts. 2026. arxiv.org/abs/2601.22156
- Mingjie Liu et al. ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models. 2025. arxiv.org/abs/2505.24864
- Aakash Lahoti, Kevin Y. Li, Berlin Chen, Caitlin Wang, Aviv Bick, J. Zico Kolter, Tri Dao, Albert Gu. Mamba-3: Improved Sequence Modeling using State Space Principles. 2026. arxiv.org/abs/2603.15569
- Qwen Team. Qwen3.5: Towards Native Multimodal Agents. February 2026. qwen.ai/blog?id=qwen3.5
- Ali Behrouz, Peilin Zhong, Vahab Mirrokni. Titans: Learning to Memorize at Test Time. 2024. openreview.net/forum?id=8GjSf9Rh7Z
- Disen Lan, Jianbin Zheng, Yuxi Ren, Xin Xia, Xuanda Wang, Xuefeng Xiao, Xipeng Qiu, Yu Cheng. Morphing into Hybrid Attention Models. 2026. arxiv.org/abs/2606.30562
- Aditya Chattopadhyay, Elvis Nunez, Prannay Kaul, Benjamin Bowman, Evan Becker, Luca Zancato, David Thomas, Wei Xia, Stefano Soatto. Priming: Hybrid State Space Models From Pre-trained Transformers. 2026. arxiv.org/abs/2605.08301
- Yanhong Li, Songlin Yang, Shawn Tan, Mayank Mishra, Rameswar Panda, Jiawei Zhou, Yoon Kim. Distilling to Hybrid Attention Models via KL-Guided Layer Selection. 2025. arxiv.org/abs/2512.20569
- Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Y. Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, Ion Stoica. DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL. 2025. notion.site/DeepScaleRhuggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset
- ByteDance Seed, Tsinghua AIR. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. 2025. arxiv.org/abs/2503.14476huggingface.co/datasets/BytedTsinghua-SIA/DAPO-Math-17k