The Expanding Canvas: How Expanding Flow Maps Make Output Size a Learnable Variable
Sophia Tang and Pranam Chatterjee's Expanding Flow Maps framework decouple generative modeling from fixed canvas sizes, pairing an expand operator with a transport map to grow and denoise state across molecular, graph, and language domains.
For a decade, the generative modeling stack has treated canvas size as a precondition. Diffusion models commit to a fixed pixel grid. Flow maps on discrete text assume a predetermined sequence length. Even the most flexible architectures typically pad, truncate, or batch to a maximum dimension before sampling begins.
That assumption works until it doesn't — and increasingly, it doesn't. Molecular conformers grow atom by atom. Drug-like graphs arrive with unknown node counts. Language models are asked to produce sentences whose length is part of the answer, not an input constraint. The mismatch between fixed-canvas generators and variable-size targets is not a minor engineering inconvenience. It is a structural limitation in how we parameterize transport from noise to data.
On July 23, 2026, Sophia Tang and Pranam Chatterjee of the University of Pennsylvania posted a paper to arXiv that proposes a direct fix. Expanding Flow Maps (EFMs) — and their multi-step parent construction, Expanding Generative Flows (EFlows) — recast generation as a process that grows the state space while denoising it. Output size becomes a learned, controllable degree of freedom rather than a hard-coded hyperparameter.
The Fixed-Canvas Problem
Flow maps have become one of the most promising routes to few-step generation. By learning to jump directly between timesteps along a continuous-time interpolant — rather than integrating hundreds of small ODE steps — models can sample in one to four function evaluations while retaining much of the quality of their multi-step teachers. Recent work has extended the idea to discrete simplex states, enabling flow-map language models and categorical generators.
But as Tang and Chatterjee note in their introduction, existing flow maps operate on a fixed canvas: a continuous space of dimension d, or a discrete sequence of fixed length L, chosen before inference starts. That rigidity conflicts with tasks where the size of the output is itself a quantity to be modeled — variable-resolution structures, audio of arbitrary duration, text of unknown length, and multimodal data whose components carry different intrinsic dimensionality.
The central question their work addresses is blunt: How can we enable few-step generation of both continuous and discrete data with adaptive dimensionality at inference?
Expand, Then Transport
The answer decomposes generation into two learnable operations that compose at every timestep jump.
The expand operator augments the current state with new coordinates or tokens, lifting a lower-dimensional sample into a higher-dimensional space via conditional noise. Concretely, given a source state x<sub>s</sub> living in ℝ<sup>d(s)</sup> and a target dimensionality d(t) > d(s), the operator draws augmented noise ε from a tractable conditional distribution and inserts it according to a placement scheme. The paper describes three instantiations: concatenation (appending new coordinates), positional insertion (interleaving tokens at chosen positions), and child expansion (attaching new coordinates to parent entries — useful for coarse-to-fine molecular generation).
The transport map then pushes the expanded state forward along the interpolant toward the target distribution — the familiar flow-map step, but now operating in the augmented space.
Together, these two stages yield a single map that jointly expands and denoises. Standard fixed-dimensional flow maps emerge as the special case where the expand operator is the identity.
Tang and Chatterjee formalize the continuous construction as a piecewise-deterministic Markov process (PDMP): smooth denoising transport interleaved with jump kernels that increase dimensionality. Each newly inserted coordinate receives a local time coordinate, initialized at insertion and projected onto a common [0, 1] clock — a detail that matters when different parts of the state enter the interpolant at different global times.
On the discrete side, expansion becomes token or node insertion along a sequence or graph. A learned insertion head predicts how many elements to add at each gap; a transport map denoises categories toward the target one-hot distribution. The framework extends to variable-size graph generation by treating molecules as pairs of node and edge category matrices that grow from empty to full size.
Three Domains, One Recipe
The empirical program is deliberately cross-modal. Rather than claiming a single benchmark win, the authors evaluate EFlow and EFM on three tasks that stress different expansion mechanics.
Coarse-to-fine molecular conformers
On GEOM-QM9 and GEOM-Drugs, EFlow treats conformer generation as a continuous expanding interpolant: heavy-atom backbones are denoised and frozen, then hydrogens are inserted and resolved around the fixed scaffold. Using the GeoDiff dual-encoder velocity field for fair comparison against multi-step diffusion baselines that require 500–1,000 denoising steps, EFlow reaches competitive or superior coverage and RMSD metrics with 16–50× fewer steps.
On GEOM-QM9 at a 0.5 Å threshold, EFlow achieves 88.51% coverage-recall and 53.07% coverage-precision at 20 steps. With an optional coordinate refiner (EFlow+R), precision sharpens further — the paper reports the best average-minimum-RMSD-precision of any compared method at 0.4876 Å on QM9 at 30 steps. On the larger GEOM-Drugs molecules (up to 181 atoms), EFM+R attains the single best score on all four reported metrics.
After distillation into few-step EFMs, the framework produces conformers in as few as one to four steps, with EFM+R at 14 steps taking the top position across GEOM-Drugs recall and precision columns in the authors' Table 1.
Variable-size molecular graphs
For discrete graph generation on QM9, EFlow grows molecules from an empty graph using a learned insertion head while a DeFoG graph transformer denoises node and edge categories. At a 100-step budget, EFlow reports FCD 0.116 versus DeFoG's 0.134, with validity and uniqueness above 95%. As the step budget shrinks, the gap widens: at four steps, EFlow still produces 91.7% valid molecules where DeFoG falls to 53.6% validity.
The distilled EFM student is where variable-size expansion pays off most visibly in the few-step regime. At two steps, EFM achieves FCD 0.40 compared with categorical flow maps' 0.49. At one step, EFM's FCD remains 0.44 while CFM degrades to 2.14 — nearly five times higher — with EFM retaining 97.7% uniqueness versus CFM's 91.8%.
Variable-length language modeling
The language experiment is perhaps the most conceptually demanding: EFlow assigns each token an independent insertion time and learns per-gap insertions, allowing sequences to grow from empty to full length while a single network denoises token categories. On the One Billion Word benchmark (LM1B), EFlow improves on the fixed-length flow language model baseline at every step budget evaluated, despite solving the strictly harder variable-length problem.
Generative perplexity under GPT-2-Large falls from 133.62 at 64 steps to 103.63 at 1,024 steps, beating FLM at all compared budgets — for example, 103.63 versus 111.36 at 1,024 steps. The authors note that sample entropy stays flat around 4.14–4.16, slightly below FLM's 4.30–4.36 range, while remaining coherent.
In the distilled few-step regime, EFM at one step reports generative perplexity 98.35, improving on flow-map language models and categorical flow maps across the compared baselines in Table 4 — though the paper candidly notes that single-step sampling exhibits mode collapse, attributing the difficulty to predicting a full-length expansion and applying the flow map over the entire interval in one shot. Even so, the authors describe this as, to their knowledge, the first demonstration of flow maps on variable-length discrete sequences.
Why It Matters Now
Expanding Flow Maps arrive at a moment when the generative modeling community is bifurcating along two axes: speed (few-step distillation, flow maps, consistency models) and flexibility (variable-length diffusion language models, set-based decoding, insertion-based architectures). Tang and Chatterjee's contribution is to show these axes need not be orthogonal.
The expand-transport decomposition offers a unified mathematical handle. Fixed-canvas flow maps, coarse-to-fine molecular pipelines, and variable-length text generators can be read as instances of the same operator algebra — with the expand operator set to identity when dimensionality is known upfront.
Practically, the results suggest that committing to a maximum sequence length or atom count at initialization may be leaving quality on the table. EFlow beats fixed-length FLM on LM1B while solving a strictly more expressive problem. EFM generates valid variable-size graphs in one step where fixed-canvas categorical flow maps collapse. And on GEOM-Drugs, a refiner-augmented EFM at 14 steps outperforms thousand-step diffusion pipelines on all four headline metrics.
Limitations the Authors Name
The paper is explicit about what it has not yet shown. Experiments cap at 29 atoms on GEOM-QM9, 181 on GEOM-Drugs, and 128 tokens on LM1B. Conditioning on per-dimension local time coordinates expands the hypothesis space relative to uniform schedules; the authors expect further tuning will be needed to scale to larger dimensionalities, and note they could not fully explore that regime due to computational cost.
Single-step language generation remains fragile. Sample entropy runs lower than fixed-length baselines, and mode collapse at one step is acknowledged rather than hidden. The conditional noise distribution is not required to be Gaussian — learning insertion means and covariances is flagged as a future extension, as is exploring decreasing-dimensionality operations within the same PDMP formalism.
None of that diminishes the core claim. The paper establishes a principled framework for settings where output size is itself a learned variable, with verified empirical gains across continuous 3D coordinates, discrete graphs, and token sequences.
The Takeaway
Generative modeling spent years optimizing how to transport noise into structure. Expanding Flow Maps ask a prior question: how much structure should exist at each step of that transport?
By factoring each timestep jump into expansion and transport, Tang and Chatterjee give labs a single recipe for growing canvases during inference — not before it. For molecular design pipelines wrestling with unknown graph sizes, for language models that should decide their own stopping points, and for any domain where padding to a maximum length is a lie about the data, that reframing may prove more durable than any single benchmark row.
The paper is on arXiv now. Reproduction code and weights were not listed in the initial submission metadata we reviewed; practitioners should watch for release announcements from the authors' Penn affiliations.
### Sources
- arXiv — Expanding Flow Maps (July 23, 2026)
- University of Pennsylvania (via arXiv PDF) — Expanding Flow Maps (full paper) (July 23, 2026)
- OpenReview — Expanding Flow Maps submission (2026)