Immediately, we’re launching agentic video understanding throughout our newest fashions: Gemini 3.7 Flash, 3.6 Flash and three.5 Flash-Lite. This new functionality improves accuracy whereas dramatically decreasing token utilization and prices for video evaluation. Much like agentic vision, which mixes code execution with Gemini fashions’ native picture understanding, agentic video understanding makes use of Gemini’s native video instruments to enhance efficiency and unlock new capabilities for video processing like sub-second second retrieval, extra correct anomaly detection, exact counting and extra.
The function is accessible in the present day for video uploads and YouTube movies by way of the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
Benchmarks
Not like present ‘static’ processing, the place the mannequin ingests the video at a set frames-per-second charge (default 1 FPS, adjustable by way of API), agentic video understanding pairs the mannequin’s core reasoning with native video instruments to dynamically search, scan, and examine goal video segments throughout visible frames, audio, and transcripts. Throughout commonplace video evaluation benchmarks, Gemini fashions with agentic video understanding scale back evaluation prices by as much as 66% and token consumption by as much as 88%, whereas bettering accuracy by as much as 7%.
These effectivity beneficial properties are particularly pronounced on long-form video (from 10-minute how-to guides to 90-minute lectures and multi-hour recordings), the place static processing forces builders to decide on between excessive token prices or strategies that drop crucial particulars.
