Robotics · 2 min read

Future-State-Aware: MIT's VLASH Makes VLAs Real-Time Without Inpainting Overhead

MIT's VLASH framework enables asynchronous VLA inference with up to 11.8× faster reactions — letting π₀.₅ play ping-pong and whack-a-mole without architectural changes.

By Classy AI News · July 31, 2026

Future-State-Aware: MIT's VLASH Makes VLAs Real-Time Without Inpainting Overhead

Vision-Language-Action models can plan impressive manipulation sequences — and still freeze in place while waiting for inference to finish. That synchronous stop-and-go pattern is fatal for ping-pong rallies, whack-a-mole, and any task where the world moves faster than the model thinks. MIT researchers Jiaming Tang, Yufei Sun, and collaborators propose VLASH — future-state-aware asynchronous inference that cuts reaction latency by up to 11.8× without architectural changes or runtime inpainting overhead.

Modern technology workspace with equipment

The prediction-execution gap

In synchronous VLA deployment, the robot generates an action chunk, executes it, then waits for the next inference cycle. During execution, the robot cannot perceive or respond to environmental changes — and between chunks, motion stalls entirely.

Asynchronous inference solves the stall by running inference while executing the current chunk. But it introduces a harder problem: temporal misalignment. The model predicts actions for state sₜ, yet those actions execute at future state sₜ₊Δ after inference delay Δ. Existing async methods either degrade accuracy (naive switching) or add inpainting overhead (Real-time Chunking) that widens the gap further.

VLASH: roll the state forward

VLASH's core insight is simple: condition the policy on the execution-time robot state, not the stale state at inference start.

Because the robot continues executing the previous action chunk during inference, the actions over the delay interval are known. VLASH rolls the robot state forward under those actions — mirroring how humans compensate for reaction delay using proprioception even when visual input is slightly outdated.

At deployment, VLASH modifies only the model input. No new heads, no inpainting pass, no inference-time overhead.

Clean desk setup with computer and accessories

Benchmark results

The paper (arXiv:2512.01031), from MIT, NVIDIA, Tsinghua, UC Berkeley, UCSD, and Caltech, evaluates VLASH across π₀.₅, GR00T N1.6, and SmolVLA:

  • Up to 30.5% accuracy improvement over naive async baselines in simulation.
  • Up to 11.8× reaction speedup over synchronous inference in real-world experiments.
  • With action quantization: 1.5–2.0× task completion speedup with minimal accuracy loss.

Most striking: VLASH enables π₀.₅ to play ping-pong rallies with a human and whack-a-mole — tasks where synchronous control fails entirely.

Training integrates via offset augmentation in standard fine-tuning pipelines — shared observations with multiple state/action offset branches — requiring no architecture changes.

Code is available at github.com/mit-han-lab/vlash.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.