Research · 2 min read

Timing Is a Measurement Problem: RL for Code Optimization Finally Works

Meta researchers show why RL for code speed fails under noisy timing rewards — and how DMC-Optim plus adapted GRPO lifts Qwen 7B and CWM 32B optimization pass rates without sacrificing correctness.

By Classy AI News · July 30, 2026

Timing Is a Measurement Problem: RL for Code Optimization Finally Works

Reinforcement learning has proven it can teach language models to write correct code: generate a program, run hidden tests, reward passes. Teaching models to write fast code looked like a small tweak — add execution time to the reward. In practice, it breaks.

A paper posted to arXiv on July 28, 2026 — "Reinforcement Learning for Code Optimization" (2607.25970) by Pierre Chambon and colleagues — documents why timing rewards collapse under measurement noise, reward sparsity, and GRPO instability, and how to fix it.

Purple and blue abstract digital light pattern

Why speed is harder than correctness

Once timing drives the reward, the authors report, generated solutions become barely faster — and more fail outright. Small sandbox measurement errors overwhelm the signal. Standard RLVR setups that work for correctness do not survive the noisier timed-execution setting without redesign.

Three-stage fix: benchmark, reward, learning

The team built DMC-Optim with large optimization test suites and a calibrated timing sandbox. They composed correctness and speed rewards in the RL environment and used an offline simulator to predict promising configurations before expensive rollouts. They adapted GRPO and evaluation protocols for sparser, noisier timed rewards.

Verified gains on real models

On DMC-Optim, optimization-aware configurations improved strict top-50% pass@1 from 18.0% to 31.3% on Qwen 2.5 7B and from 30.7% to 50.4% on CWM 32B — while preserving pure-correctness scores. At stricter top-30% cutoffs, CWM 32B saw roughly 125% relative improvement.

When the timing sandbox was deliberately degraded, robust optimization RL reached 100–200% improvement over standard RLVR depending on the metric. On LiveCodeBench (LCB), CWM 32B won up to 83% of median-sample speed comparisons against standard RLVR.

Person holding a blue light bulb representing innovation

Half the human optimization rate — and that may be the point

Relative to the fastest correct human submissions per problem, the best RL runs reached about 14% versus 28% for humans on complexity-class improvements. The gap is real; so is the direction. Code optimization via RL is no longer a naive extension of correctness training — it requires its own benchmark, reward engineering, and learning recipe.

For agentic coding stacks racing toward faster inference, this paper is a reminder: speed is a measurement problem before it is a modeling problem.

Close-up of a smartphone with abstract blue background

### Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.