The Relation Graph: GraphVid Moves Video Control From Pixels to Structured Interactions
GraphVid, posted to arXiv on July 23, replaces trajectory drawing with structured interaction graphs for multi-object video control — and reports large quality gains over Motion-I2V with fewer parameters.
Text prompts and pixel trajectories have long been the control language for generative video. They work until a scene contains three objects, two occlusions, and one relationship that cannot be drawn as a line on a frame.
On July 23, 2026, researchers posted GraphVid to arXiv (2607.21580v1), proposing structured interaction graphs that specify how subjects relate before pixels move.
Why trajectories break under complexity
The GraphVid team argues that trajectory-based control scales poorly as scenes grow crowded. GraphVid accepts an image plus an interaction graph encoding relational constraints.
GraphVid-Bench and reported gains
Compared with Motion-I2V, GraphVid reports FID down up to 39.9%, FVD down 37.6%, PSNR from 9.87 to 15.98, and SSIM from 0.38 to 0.61 — with fewer trainable parameters and less training data.
Structured semantics as a control layer
Interaction graphs occupy a middle layer: precise enough for multi-object choreography, abstract enough to survive occlusion.
### Sources
- arXiv — GraphVid: Interactive Graph-Controllable Video Generation (2607.21580v1) (July 23, 2026)
- Pneumetron — GraphVid: Advancing Controllable Video Generation via Structured Interaction Graphs (July 25, 2026)