AIToday
Large Language ModelsGoogle DeepMindPublished: Sep 2, 2026, 04:00 JST2 min read

Gemini launches agentic video understanding, cuts tokens by 88%

Gemini launches agentic video understanding, cuts tokens by 88%

Key takeaway

  • Google DeepMind launched agentic video understanding for Gemini models. It cuts token use by up to 88% and costs by up to 66%.

  • Accuracy improves by up to 7%.

  • The feature is available today via the Gemini API.

3 Key Points

  1. What happened

    Google DeepMind launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. It is available today via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.

  2. Why it matters

    This feature cuts token consumption by up to 88%, reduces costs by up to 66%, and improves accuracy by up to 7% on standard video analysis benchmarks. The efficiency gains are especially pronounced on long-form videos, from 10-minute how-to guides to 90-minute lectures and multi-hour recordings.

  3. What to watch

    Gemini 3.7 Flash with agentic understanding offers the best overall quality and the best quality-cost combination, placing it at the accuracy-to-cost pareto frontier among tested models. The feature will roll out to all users in the Gemini app soon and will power YouTube's 'Ask YouTube' feature in the coming months.

Ask the AI about this article →

Context & Analysis

The launch marks a shift from static video processing, where models ingest video at a fixed frames-per-second rate (default 1 FPS), to a dynamic approach. Agentic video understanding lets Gemini take a goal-directed role in deciding what to watch, at what speed, and through which modality—frames, audio, or transcript—fetching only the moments needed. This reduces development overheads, as developers previously had to manually implement such selective processing.

The efficiency gains are most notable on long-form content, where static processing forces a trade-off between high token costs and dropping critical details. Early access partners saw strong performance, and Google plans to extend the feature to consumer products: it will roll out to all Gemini app users soon, and later power YouTube's 'Ask YouTube' feature on the video watch page. The feature uses standard Gemini API token pricing with no extra fee, which could make sophisticated video analysis more accessible to developers.

FAQ

How do I enable agentic video understanding?
Set processing to 'agentic' in the API configuration. It uses standard Gemini API token pricing with no additional feature fee.
Which models support agentic video understanding?
It is available on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite.
Google DeepMindRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DeepMind chief: frontier AI leadership is all that mattersTHE DECODER · 1h ago
  • John Deere launches AI chatbot for farmersThe Verge AI · 1h ago
  • Google Pics launches with AI image editing for WorkspaceThe Verge AI · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleJamf builds real-time AI cost caps for Amazon Bedrock