Timing Is a Measurement Problem: RL for Code Optimization Finally Works
Meta researchers show why RL for code speed fails under noisy timing rewards — and how DMC-Optim plus adapted GRPO lifts Qwen 7B and CWM 32B optimization pass rates without sacrificing correctness.
Reinforcement learning has proven it can teach language models to write correct code: generate a program, run hidden tests, reward passes. Teaching models to write fast code looked like a small tweak — add execution time to the reward. In practice, it breaks.
A paper posted to arXiv on July 28, 2026 — "Reinforcement Learning for Code Optimization" (2607.25970) by Pierre Chambon and colleagues — documents why timing rewards collapse under measurement noise, reward sparsity, and GRPO instability, and how to fix it.
Why speed is harder than correctness
Once timing drives the reward, the authors report, generated solutions become barely faster — and more fail outright. Small sandbox measurement errors overwhelm the signal. Standard RLVR setups that work for correctness do not survive the noisier timed-execution setting without redesign.
Three-stage fix: benchmark, reward, learning
The team built DMC-Optim with large optimization test suites and a calibrated timing sandbox. They composed correctness and speed rewards in the RL environment and used an offline simulator to predict promising configurations before expensive rollouts. They adapted GRPO and evaluation protocols for sparser, noisier timed rewards.
Verified gains on real models
On DMC-Optim, optimization-aware configurations improved strict top-50% pass@1 from 18.0% to 31.3% on Qwen 2.5 7B and from 30.7% to 50.4% on CWM 32B — while preserving pure-correctness scores. At stricter top-30% cutoffs, CWM 32B saw roughly 125% relative improvement.
When the timing sandbox was deliberately degraded, robust optimization RL reached 100–200% improvement over standard RLVR depending on the metric. On LiveCodeBench (LCB), CWM 32B won up to 83% of median-sample speed comparisons against standard RLVR.
Half the human optimization rate — and that may be the point
Relative to the fastest correct human submissions per problem, the best RL runs reached about 14% versus 28% for humans on complexity-class improvements. The gap is real; so is the direction. Code optimization via RL is no longer a naive extension of correctness training — it requires its own benchmark, reward engineering, and learning recipe.
For agentic coding stacks racing toward faster inference, this paper is a reminder: speed is a measurement problem before it is a modeling problem.
### Sources
- arXiv — Reinforcement Learning for Code Optimization (July 28, 2026)
- arXiv — Reinforcement Learning for Code Optimization PDF (July 28, 2026)