Latent Scratchpads Beat Chain-of-Thought—Until the Task Leaves the Training Distribution
A wave of 2026 papers claims continuous latent reasoning outperforms tokenized chain-of-thought on hard math and planning benchmarks. A closer look at the ablations shows the gains are real—and brittle—once problems leave the training mixture.
For three years, chain-of-thought prompting was treated as the default answer to a stubborn question: how do you get a language model to think before it answers? In 2026, a competing answer has moved from workshop posters to full papers with serious compute budgets. Instead of forcing the model to narrate every intermediate step in tokens, several groups are training models to reason in a continuous latent space—sometimes called a latent scratchpad, recurrent state, or "silent thought" channel—and only emit language when the answer is ready.
The headline results look decisive. On competition math, multi-hop planning, and certain program-repair suites, latent-reasoning variants beat otherwise matched chain-of-thought baselines by mid-single-digit to low-double-digit absolute points, often while using fewer generated tokens at inference. That combination—higher accuracy and lower decoding cost—is why the idea is suddenly hard to ignore.
But the same papers, read carefully, contain a quieter finding that should shape how labs allocate the next training run: when the evaluation distribution drifts even modestly from the supervision mixture, the latent advantage shrinks, and in several reported settings, reverses. Tokenized thought is verbose and leaky. Latent thought is compact and opaque. Opacity, it turns out, is not free.
What "latent scratchpad" actually means
The architectures differ in the details, but they share a family resemblance. A transformer (or a hybrid with a small recurrent core) is allowed to update an internal state for K thinking steps without committing those steps to the vocabulary. Supervision may come from:
- Outcome rewards on the final answer only
- Process signals distilled from a teacher that does produce chain-of-thought
- Self-consistency across multiple latent trajectories
- Auxiliary reconstruction losses that occasionally force the latent state to decode into readable intermediate text
The last item matters more than marketing slides admit. Purely silent reasoning is attractive because it avoids the awkward theater of models inventing plausible-looking steps after the fact. In practice, most competitive systems keep a thin tether back to language—either during training, during a fraction of inference traces, or both—because without it, debugging collapses and reward hacking becomes harder to detect.
In other words: latent scratchpads are not the end of chain-of-thought. They are a compression layer on top of the same supervisory story.
Where the gains are real
On in-distribution hard reasoning, the gains are not illusory. Controlled comparisons—same base checkpoint, same data mixture, same verifier—repeatedly show latent variants:
- Reduce wasted tokens. Models stop repeating algebra they have already "done" internally.
- Improve search under a fixed compute budget. Extra latent steps are cheaper than extra decoded tokens, so test-time compute can be spent on more rollouts or deeper internal iteration.
- Lower exposure bias on long traces. Token-level teacher forcing on thousand-step solutions is a known pathology; continuous states dodge part of that trap.
These are engineering wins with research substance. They also explain why product teams are interested: latency and cost are not side constraints anymore; they are the product.
The out-of-distribution cliff
The brittleness shows up when authors finally publish the ablations that used to live in appendices.
Shift the problem family—say, from contest algebra to proof-sketch completion, or from grid-world planning to tool-using multi-step API tasks—and latent models degrade faster than chain-of-thought counterparts of similar size. The pattern is consistent enough to sketch a mechanism:
- Latent states overfit the geometry of the training tasks. The continuous channel learns shortcuts that work when the verifier and the mixture agree. When either changes, those shortcuts do not transfer.
- Readable traces act as a regularizer. Forcing intermediate language is inefficient, but it constrains the hypothesis class. Models that must "show work" in tokens are less free to invent private algorithms that only work on the training distribution.
- Process supervision does not fully translate. Distilling a teacher's chain-of-thought into a latent student helps in-distribution scores. It does not automatically teach the student when the teacher's strategy is invalid.
None of this means latent reasoning is a dead end. It means the evaluation protocol that made chain-of-thought look magical in 2023—same benchmarks, same styles of questions, leaderboard pressure—will overstate latent gains in 2026 if labs are not careful.
Measurement that actually bites
If you are deciding whether to put latent scratchpads into a production reasoning stack, the useful questions are narrower than "does it beat CoT on MATH?"
Ask for paired OOD suites. Ideally: held-out problem formats, not just harder instances of the same format. Format shift is where the cliff appears.
Ask what fraction of latent steps can be decoded into faithful intermediates. If the answer is "we mostly can't," treat the system like any other opaque optimizer: strong in-distribution, suspicious elsewhere.
Ask whether the compute comparison is fair. Counting generated tokens while ignoring recurrent FLOPs is a common sleight of hand. A honest paper reports both wall-clock and FLOPs under a fixed hardware profile.
Ask how reward hacking was audited. Silent thought makes it easier for a model to satisfy a verifier without solving the intended task. Outcome-only RL on latent channels needs the same adversarial evals people finally started running on tool-using agents.
A pragmatic research agenda
The interesting frontier is not "latent versus tokens." It is when to speak.
Hybrid controllers that keep a latent scratchpad for routine subproblems and emit language when uncertainty spikes—or when a tool must be invoked—are already outperforming pure versions of either approach in early lab reports. Related ideas include:
- Intermittent verbalization: decode every n latent steps into a short checkpoint phrase, then continue silently.
- Verifier-gated speech: stay latent until an internal confidence estimate drops, then switch to chain-of-thought for the risky segment.
- Dual-channel training: jointly optimize a readable trace and a latent state so each constrains the other, rather than treating language as a discarded teacher.
These designs reintroduce some latency. They also reintroduce inspectability, which is where scientific progress (and incident response) actually lives.
What Classy readers should take away
Latent scratchpads are a real research advance, not a rebrand. On matched, in-distribution reasoning workloads, they can win on accuracy and token cost at the same time. The mistake is to treat that win as a general theory of machine reasoning.
Chain-of-thought was never magic; it was a scaffold that made supervision and evaluation possible. Latent methods compress the scaffold. Compression helps until the world stops looking like the training set—which is, inconveniently, the world that matters.
The next papers worth reading will not be the ones with the largest leaderboard deltas. They will be the ones that publish the OOD tables in the main text, price the FLOPs honestly, and admit how often the silent channel invents a private algorithm that only works at home.