Research · 2 min read

SURGE Extracts Stronger Policies From Finished RL Runs Without Retraining

An arXiv paper posted 1 October 2026 shows SURGE can fuse checkpoints from completed reinforcement learning histories into policies that beat the best training checkpoint on math and coding benchmarks without additional training or extra inference tokens.

By Classy AI News · October 5, 2026

SURGE Extracts Stronger Policies From Finished RL Runs Without Retraining

What changed

Researchers posted Does Scaling Reinforcement Learning Really Require More Training? to arXiv on 1 October 2026 (identifier 2610.01133). The work introduces SURGE (Scaling Up RL via Gradient free Eigenspace fusion), a method that constructs deployable policies from an existing reinforcement learning training history without extending training or increasing per query inference compute.

The team calls this policy space scaling: expanding the set of policies reachable from a fixed RL run. SURGE pairs a high accuracy anchor checkpoint with a donor checkpoint, fuses complementary components in eigenspace, and selects how much of the anchor update to retain using a fixed target without testing every candidate policy on benchmarks.

Reported results span three histories: two 1.5B mathematical reasoning runs (DeepSeek and Nemotron) and one 7B coding run (OLMo). SURGE improved benchmark average accuracy over both input checkpoints while using fewer reasoning tokens than the anchor alone. On DeepSeek AIME24 the method reached 54.17% versus a measured native maximum of 50.83%. On OLMo HumanEval+ it reached 83.7% versus 82.8%. The authors argue these gains exceed what the original training curves suggested was available from each run.

Research desk with monitors showing training curves

Why it matters

Teams that already spent millions of GPU hours on reasoning RL may recover capability they thought was capped at the best saved checkpoint. That shifts capital allocation: post training fusion could delay or shrink a full retrain when benchmarks stall late in a run.

The method also challenges the default assumption that scaling reasoning always means more RL steps or more inference time. If stored histories are reusable assets, MLOps teams need retention policies for intermediate checkpoints and optimizer states, not just the final export.

Who is affected

Applied AI leads running math or code RL should audit whether archived checkpoints from 2025 and 2026 runs are still stored with enough metadata to pair anchor and donor models.

Inference platform owners benefit if fused policies deliver higher accuracy with fewer reasoning tokens, directly lowering serving cost for agent products.

Research labs publishing RL scaling laws should treat policy space methods as a separate axis from compute scaled training.

What to do next

Before launching another full RL cycle on a stalled reasoning model, task one engineer to reproduce SURGE style fusion on your last completed run using the public paper specification. Compare fused policy accuracy and token use against your production checkpoint on two held out suites.

What to watch

Whether any frontier lab ships a production fusion tool or documents SURGE style gains on models above 7B parameters before 31 December 2026. Replication at larger scale would confirm the approach is more than a small model artifact.

Whiteboard with matrix sketches in a machine learning lab

Sources

  1. Primary. arXiv, Does Scaling Reinforcement Learning Really Require More Training? (1 October 2026). Defines SURGE, reports DeepSeek, Nemotron, and OLMo benchmark numbers, and states no additional training or inference compute is required.
  2. Secondary. Hugging Face community discussion on arXiv 2610.01133 (October 2026). Independent restatement of the policy space scaling claim and benchmark table values.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.