Google has introduced agentic video understanding for its latest Gemini models, enabling dynamic analysis that reduces token usage by up to 88% while improving accuracy by up to 7%. This capability allows the models to actively search and inspect specific video segments rather than processing frames at a fixed rate.
The feature is available via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Key capabilities include sub-second moment retrieval, long-form needle-in-a-haystack search, anomaly detection, and precise counting of actions and objects. Early access partners reported strong performance, with Gemini 3.7 Flash achieving the best accuracy-to-cost efficiency.
Google plans to roll out these improvements to billions of users in the Gemini app and will power YouTube's 'Ask YouTube' feature in the coming months.