The Schedule Layer: LeSAMP Learns When Diffusion Models Should Change Their Sampling Knobs
LeSAMP, posted to arXiv on July 26, uses reinforcement learning to learn prompt-conditioned, timestep-varying diffusion sampling schedules—reporting up to 68% win rates over static baselines on Flux and SD3.5.
Diffusion models ship with a hidden control panel—and almost nobody touches it between prompts.
Text-to-image systems expose classifier-free guidance scales, negative prompts, noise schedules, and timestep-specific knobs. In practice, teams pick one configuration and reuse it everywhere. That default works until it does not: a portrait prompt and a product render rarely want the same sampling trajectory.
On July 26, 2026, Arisrei Lim and Yossi Gandelsman posted LeSAMP (Learning Sampling Parameters for Diffusion Models) to arXiv, proposing a learned policy that emits prompt-conditioned, timestep-varying sampling schedules. The paper treats parameter selection as reinforcement learning: a language model reads the user prompt and outputs schedules; rewards come from human preference models and vision-language judges.
Why static sampling leaves quality on the table
Diffusion quality is not only a weights problem. The denoising path matters. Holding guidance fixed across timesteps assumes every denoising stage benefits equally from the same pressure toward the text condition—a assumption LeSAMP tests and rejects.
The authors evaluate on Flux.1 [dev] and Stable Diffusion 3.5, reporting win rates up to 68.12% against baselines on human preference scores and 73.37% with VLM-as-a-judge. A user study reports wins up to 59.46% over prior baselines.
RL on the inference stack, not the weights
LeSAMP sits above post-training. It does not retrain U-Net or transformer backbones; it learns when to turn inference knobs. That matters for production: teams can bolt a schedule policy onto existing checkpoints without a full fine-tune cycle.
The RL loop uses preference models as reward—oracles, a pattern now familiar from chat RLHF but applied here to per-timestep control signals. The LLM policy maps natural-language intent to numeric schedules, making the system prompt-native in the same way modern image editors wrap diffusion behind chat interfaces.
Benchmarks and limits
Reported gains are model-specific and judge-dependent. VLM-as-judge rewards can diverge from human taste on niche domains (medical imaging, technical diagrams, localized text rendering). The paper positions LeSAMP as complementary to LoRA and distillation—not a replacement.
Still, the result suggests an under-explored axis: inference-time compute allocation. If a policy can learn timestep schedules, similar frameworks may learn when to invoke refiner models, upsamplers, or safety filters—an agenda adjacent to the token-routing debates now surfacing in enterprise AI budgets.
Takeaway for builders
LeSAMP is a reminder that diffusion deployments still run on hand-tuned folklore. Formalizing sampling as a learnable policy could standardize quality across prompts the way adaptive optimizers standardized training.
The arXiv release is open; reproducibility will depend on whether authors publish training code alongside the PDF. For now, the headline is clear: better images without retraining the world—if you teach the scheduler to listen to the prompt.
### Sources
- arXiv — Learning Sampling Parameters for Diffusion Models (July 26, 2026)
- AP News — 'Tokenmaxxing' hits its limits as workplaces look for cheaper AI (July 2026)