Google Research has introduced an AI video co-director consisting of four agentic frameworks designed to generate coherent, minutes-long videos by addressing identity drift and cascading errors common in multi-shot pipelines. The system operates as a model-agnostic orchestration layer on top of Gemini and Veo, utilizing SynthID watermarking for output.
The suite includes Co-Director, which uses multi-armed bandit search for creative planning; CANVAS, which provides persistent visual memory to maintain character and object consistency; A²RD, a training-free architecture for segment-by-segment long video generation; and VQQA, which employs closed-loop prompt refinement via visual question answering.
Google evaluated these frameworks using new benchmarks including GenAD-Bench, HardContinuityBench, and LVBench-C. Co-Director achieved an 81.4 average on GenAD-Bench, while A²RD demonstrated up to 30% better consistency and 20% better narrative coherence on videos ranging from 1 to 10 minutes.
Co-Director and A²RD code is publicly available on GitHub, with CANVAS code pending release. The full pipeline is not yet a commercial Google product, but the frameworks target the global optimization and world-state tracking problems inherent in long-form AI video creation.