Research · 2 min read

Looped Flows Lift Recurrent Reasoning to 58.8 Percent on ARC AGI 1

A 10 September arXiv paper trains looped recurrent models with local denoising objectives and reports state of the art looped model scores on ARC AGI 1 and ARC AGI 2.

By Classy AI News · September 11, 2026

Looped Flows Lift Recurrent Reasoning to 58.8 Percent on ARC AGI 1

What changed

Researchers posted Thinking with Looped Flows to arXiv on 10 September 2026 (arXiv:2609.11801), proposing looped flows: a training method that uses local denoising objectives to teach recurrent hidden state updates that transfer computation across inference steps. The authors report 58.8 percent test accuracy on ARC AGI 1 and 12.2 percent on ARC AGI 2 across six reasoning benchmarks, which they describe as outperforming prior state of the art looped models overall.

Looped models run multiple internal updates at inference time, but standard training often backpropagates through only one or a few steps, which weakens early updates that must support later ones. Looped flows tie denoising objectives across noise levels with shared noise so gradients incentivize states that remain useful across recurrent passes. Inference integrates a probability flow velocity field coupled with those recurrent states, allowing finer temporal grids and multiple valid outputs from different initial noise samples.

The benchmark suite includes two multi solution tasks where several answers are valid, which matters for planning systems that must propose alternatives rather than collapse to one completion.

Server racks representing compute for test time reasoning workloads

Why it matters

ARC style reasoning benchmarks remain a practical stress test for models that must generalize beyond memorized patterns. A looped architecture that improves ARC AGI 2 scores without requiring full unrolled backprop through every inference step matters for teams building test time compute systems on fixed training budgets. If the denoising trick generalizes, it offers a path to scale inference time reasoning without linearly scaling training memory.

Product teams shipping agents with internal deliberation loops should care because the method explicitly targets truncated backprop regimes that match production constraints. Research leaders comparing diffusion style training to recurrent reasoning now have a published ARC AGI 2 number to benchmark against before investing in custom looped stacks.

Who is affected

Applied research leads evaluating recurrent or diffusion hybrid architectures for internal reasoning modules should read the benchmark table and ablations. Inference platform engineers sizing GPU memory for long unrolls may prefer methods that train with truncated gradients but still benefit from many test time steps. Benchmark owners tracking ARC AGI leaderboards should note the looped flows entry and whether submissions disclose compute at inference time.

What to do next

Reproduce the ARC AGI 1 configuration on your internal reasoning slice before committing architecture changes. Compare looped flows against your current best looped baseline under matched inference step budgets. If you ship multi answer UX, test whether multiple noise seeds produce meaningfully diverse valid plans on your domain tasks rather than cosmetic variation.

What to watch

Watch for peer review or workshop acceptance follow ups on arXiv:2609.11801. Watch whether independent labs replicate ARC AGI 2 gains under disclosed compute budgets. Watch if authors release code or checkpoints referenced in the paper trail.

Motherboard close up representing efficient recurrent inference hardware

Sources

  1. Primary. arXiv, Thinking with Looped Flows (2609.11801) (10 September 2026). Methods, benchmarks, and reported ARC AGI scores.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.