Generated by Codex with GPT 6 Luna XHigh
High-quality video clips are now easy to generate; keeping a story coherent across many shots is harder. Characters can change appearance, rooms can shift shape, and a flawed early asset can contaminate later stages. Google Research’s latest work treats long-form video as a planning and state-tracking problem around the video models, with separate mechanisms for choosing a story direction, remembering visual details, extending it over time, and correcting artifacts.
The official Google Research Blog published “Automating coherent long-form video generation” on September 24, 2026. It presents four related frameworks that orchestrate generative models: AI video co-director, CANVAS, A²RD, and VQQA. The co-director uses Gemini and Veo, while the other frameworks add continuity memory, long-horizon generation, and visual critique. Together they turn a high-level creative brief into a storyboard, audiovisual segments, and iterative quality checks.
Plan at story scale, remember at shot scale
The AI video co-director frames creative direction as a search problem. Its orchestrator uses a multi-armed bandit to explore combinations of creative strategy, narrative mode, and visual aesthetic. A pre-production agent turns a selected direction into a scene-by-scene storyboard and visual assets; keyframe, video, and audio agents then produce coordinated media. A multimodal judge evaluates the result across those dimensions and feeds a factored reward back to the bandit, so later generations can favor directions that worked.
This arrangement gives the system a global objective instead of relying on a chain of unrelated prompts. The bandit operates over high-level choices, while specialist agents handle concrete production steps. The blog says the orchestration layer can sit over different foundation models; its examples use Gemini and Veo and retain native protections such as SynthID watermarking.
CANVAS addresses continuity inside the story. It keeps structured representations of characters, locations, and object states, along with a persistent visual memory of reference anchors. When a later scene returns to a character or place, the system can retrieve the relevant visual state rather than reconstructing it from a fresh prompt. This matters especially across non-adjacent shots, where a character may have been absent for several scenes before returning.
Extend a narrative without letting it drift
A²RD tackles the longer horizon. It generates a video segment by segment and maintains multimodal video memory about prior segments, characters, and environments. For each new segment, it follows a retrieve, synthesize, refine, and update cycle. The generation mode shifts between extrapolation, which advances the story into new beats, and interpolation, which anchors returning entities and settings to established details. Google demonstrates a continuous ten-minute generation and describes memory lookups as the mechanism for holding identity and layout together across long gaps.
VQQA handles visual defects that remain after generation. It creates questions tailored to the prompt, asks a vision-language model to critique the video, and converts the critique into natural-language guidance for revising the prompt. This acts like a semantic gradient: it points the next generation toward a correction without requiring access to the video model’s internals or editing individual pixels.
The system also guards against a local fix damaging the whole result. It keeps candidate videos from the full refinement trajectory, scores them against the original prompt, and selects the strongest overall candidate instead of assuming the last iteration is best. That global selection step is a practical answer to a common optimization failure: improving one detail can weaken other constraints.
Evaluate the failure modes separately
The post describes three benchmarks designed around long-form production challenges. GenAD-Bench tests exact creative constraints using 400 scenarios built from 50 fictional brands with four products each. HardContinuityBench stresses spatial and environmental consistency across large gaps, costume changes, and complex object states. LVBench-C tests long-horizon temporal dynamics across 120 scenarios, requiring important visual elements to disappear for at least ten segments before returning in a changed but plausible state.
The reported results connect each framework to the kind of failure it targets. The co-director reaches a peak quality score of 81.4 on GenAD-Bench and improves story consistency on ViStoryBench. CANVAS reports continuity gains on ST-Bench and HardContinuityBench; A²RD reduces layout drift on VBench-Long and LVBench-C; and VQQA improves scores on T2V-CompBench, VBench2, and VBench-I2V. The blog is an overview of four research efforts, and it directs readers to their papers for full architectures, training details, and baseline evaluations.
The broader engineering lesson is to make long-horizon quality explicit. A single generator call cannot easily carry every global constraint, visual reference, and temporal change. These systems separate the work into global planning, persistent state, incremental generation, and critique with candidate selection. The same pattern applies to other generative workflows: track what must remain stable, decide what should change, evaluate outputs against the original goal, and retain the best candidate across iterations.