Google rolls out 'agentic video understanding' in Gemini Flash to cut token use and costs

GOOGL

·

Google said Tuesday it is rolling out “agentic video understanding” across several Gemini Flash models, a new video-analysis mode that the company says can sharply reduce token use and cost by inspecting only the parts of a video that matter for a given task. Unlike a research preview, the feature is available now through Google’s API and cloud tools, and the company said it will also be used in consumer-facing products including the Gemini app and YouTube’s Ask YouTube.

In a Sept. 1 blog post, Google said the system replaces the usual fixed-rate approach to video processing with what it describes as an “agentic loop.” In standard processing, Google said, models typically sample video at a default rate of 1 frame per second and analyze that stream uniformly. The new mode instead searches across the video, audio and transcript, then loads and inspects only the segments that appear relevant to the prompt. In plain terms, that means the model can skip over large portions of a long video and zoom in on moments that matter.

Google said the launch covers Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. The feature is available now through the Gemini API in Google AI Studio, the company’s developer platform, and through the Gemini Enterprise Agent Platform on Google Cloud. It works with both uploaded video files and YouTube video URLs through the API. Google said agentic video understanding uses standard Gemini API token pricing and does not carry an extra feature fee. Developers can enable it by setting processing to agentic in the API request.

The company framed the main benefit as efficiency for long-form video tasks, where sending an entire file through a model can be expensive and slow. In benchmark results published in the blog post, Google said the new approach can cut token consumption by up to 88%, lower analysis costs by up to 66% and improve accuracy by up to 7%, with the largest gains on Gemini 3.7 Flash. Google said it measured those results on standard video-analysis benchmarks and highlighted LongVideoBench for comparisons involving long videos. Those figures are Google’s own results; the blog summarizes the benchmarks but does not provide a full independent methodology in the announcement.

Google said the system is designed for tasks such as sub-second moment retrieval, so a model can find an exact scene quickly, as well as “needle-in-a-haystack” search across long videos, anomaly detection through dynamic re-sampling, and precise counting of actions or objects. The announcement extends Google’s earlier “agentic vision” work for images, introduced in January, into video. It also fits into the company’s broader push this year to make Gemini more “agentic,” or able to use tools and take multi-step actions to complete a task.

That broader strategy is now moving beyond demos and into shipping products. Google said the same efficiency and quality gains will reach Gemini app users on Flash and Flash-Lite models soon. It also said the system will power YouTube’s Ask YouTube feature on the watch page in the coming months, making this launch a more concrete platform rollout of ideas Google previewed earlier this year.

Tags: #google, #gemini, #ai, #video, #cloud

Stocks: GOOGL