The Surrogate Policy: SLPO Brings Outcome-Reward RL to Latent Reasoning and Unlocks Test-Time Scaling Below the Token Layer
SLPO brings outcome-reward reinforcement learning to latent reasoners via a surrogate policy interface and adaptive stopping head, improving Pass@8 and Pass@16 in all 12 evaluated backbone–dataset settings.
Chain-of-thought reasoning scaled impressively in 2025 and 2026—but every decoded token costs money, and much of what models write supports linguistic coherence rather than problem-solving state. Latent reasoning carries intermediate computation as continuous vectors in hidden space, often matching explicit CoT at far shorter horizons. The missing piece has been reinforcement learning: without a tractable per-step likelihood, outcome rewards could not assign credit across latent trajectories.
A new preprint closes that gap. Surrogate Latent Policy Optimization (SLPO), posted to arXiv on July 27, 2026 (v2), brings outcome-reward RL to autoregressive latent reasoners—and reports consistent test-time scaling gains across every backbone–dataset pair the authors evaluated.
The optimization interface latent models lacked
Explicit CoT reasoners benefit from RL with verifiable rewards (RLVR) because each reasoning step is a token sampled from the vocabulary distribution. Policy gradients can assign trajectory-level credit through a tractable per-step likelihood, and rollout length is inherently variable.
Latent reasoners propagate continuous vectors that bypass the vocabulary distribution at intermediate steps. As authors Runyang You, Zhiyuan Liu, Yongqi Li, and Wenjie Li (Hong Kong Polytechnic University and Sichuan University) note, existing latent systems largely remain imitation-bound—trained to align with compressed or rendered explicit CoT—while fixed thinking budgets freeze the compute horizon RL would otherwise optimize.
SLPO introduces two components:
- A surrogate likelihood in hidden space that converts rollout advantages into credit over vector-based transitions.
- A stopping head, cold-started with correctness supervision, that outcome-reward optimization refines into a variable-horizon policy.
Together, they enable standard algorithms such as RLOO or GRPO over complete latent trajectories.
Results: Pass@8 and Pass@16 rise everywhere tested
The paper evaluates two continuous latent reasoners across two backbones and three held-out benchmarks—12 backbone–dataset settings in total. SLPO improves Pass@8 and Pass@16 in all 12, with gains of up to 12.07 percentage points.
The effect persists across both RLOO and GRPO and transfers to soft-token latent inference. The learned stopping policy allocates longer latent trajectories to harder instances, converting a fixed thinking budget into difficulty-adaptive computation.
That last point matters for deployment economics. If latent RL learns when to stop—not just what to compute—operators may trade fixed token caps for variable compute that concentrates on hard queries.
Why this is not just another RL recipe
Prior latent-reasoning work (Coconut, CoLaR, Regular, and related lines cited in the paper) demonstrated that continuous intermediate states can match or surpass explicit CoT with shorter horizons. SLPO's contribution is infrastructural: it identifies the missing optimization interface between outcome rewards and latent transitions, then supplies a differentiable surrogate policy and adaptive stopping.
The authors release code at github.com/ModalityDance/SLPO.
Limits the paper acknowledges implicitly
The evaluation scope is benchmark-bound: three held-out datasets and two continuous latent backbones. Whether SLPO generalizes to tool-using agents or multimodal latent stacks remains open. The paper also does not claim to eliminate the interpretability trade-off—latent trajectories remain harder to audit than explicit CoT strings.
Still, for labs betting that the next reasoning efficiency gain lives below the token layer, SLPO offers a concrete path from imitation-only latent models to outcome-optimized test-time scaling.
### Sources
- arXiv — SLPO: Scaling Latent Reasoning via a Surrogate Policy (July 27, 2026)
- GitHub — ModalityDance/SLPO (2026)