AIToday
Large Language ModelsTHE DECODERPublished: Sep 2, 2026, 19:01 JST2 min read

Gemini video analysis cuts token use by 88%

Gemini video analysis cuts token use by 88%

Key takeaway

  • Google's Gemini now analyzes videos by selecting key moments itself.

  • Token use drops by 88% on benchmarks while accuracy edges up.

  • The feature works from 10-minute tutorials to multi-hour recordings.

3 Key Points

  1. What happened

    Google added agent-based video analysis to Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Instead of sampling every frame per second, the model picks relevant sections on its own.

  2. Why it matters

    On Google's 1H-VideoQA and LVBench benchmarks, token usage drops 88% while accuracy improves slightly. This lowers cost for developers and supports tasks like finding anomalies or counting repeated motions in hours of footage.

  3. What to watch

    The feature is live through the Gemini API with no extra fee, and it will roll out to all Gemini app users on Flash and Flash Lite soon. Later, it will power "Ask YouTube" on the playback page.

Ask the AI about this article →

Context & Analysis

This update refines Google's earlier move into "agentic vision," which shipped for Gemini 3 Flash in January. That feature let the model write and run Python to edit images, checking results in a loop. Now the same principle applies to video. The model decides which parts of a video to inspect, through frames, audio, or transcript, and retrieves only what it needs, cutting token use by up to 88% on benchmarks like 1H-VideoQA and LVBench.

The practical payoff is accuracy on small details, such as state changes or cuts shorter than one second, which static per-second sampling could miss. Google also reports that on LongVideoBench, Gemini 3.7 Flash leads in overall quality and cost efficiency, suggesting the agent-based approach is not just cheaper but often more precise.

Google says the feature will reach all Gemini app users on Flash and Flash Lite soon, with "Ask YouTube" integration following over the coming months. Because the API charges standard token rates with no extra fee, the cost savings could make sophisticated video analysis more accessible to developers.

FAQ

How do developers use the new video analysis?
Developers set the processing mode to "agentic" in the API config and pay standard Gemini API token rates with no added fee. It is available for video uploads and YouTube videos in Google AI Studio and on the Gemini Enterprise Agent Platform.
Does it work with long videos?
Yes, the efficiency gains are biggest for long videos, from 10-minute tutorials to 90-minute lectures and multi-hour recordings. Gemini 3.7 Flash scores highest on LongVideoBench, balancing accuracy and cost.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Nvidia invests $3.5B in MediaTek to profit from custom AI chipsYahoo Finance AI · 2h ago
  • Paid actors, AI scripts: Viral anti-Democrat YouTube network exposedSemafor Tech · 2h ago
  • ICRA panel warns of paper floodRobohub · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleChicago Fed's Goolsbee: Watch consumers, not AI hype