AIToday
Large Language ModelsTHE DECODERPublished: Jun 25, 2026, 19:00 JST1 min read

Google's Gemini 3.5 Flash now sees and controls computers

Google's Gemini 3.5 Flash now sees and controls computers

Key takeaway

  • Google has integrated computer-control capability directly into its Gemini 3.5 Flash AI model, enabling it to see and operate screens across computers, browsers, and mobile devices.

  • On benchmark tests, the model ranks highly for this task, making it practical for developers to build automated agents for software testing and office work.

  • The feature includes built-in safeguards and is available now through Google's Gemini API and Enterprise Agent Platform.

3 Key Points

  1. What happened

    Google integrated "Computer Use" into Gemini 3.5 Flash, allowing the model to see, understand, and interact with computers, browsers, and mobile devices on its own. Previously this capability was only available as a separate Gemini 2.5 model. The feature is now available through the Gemini API and the Gemini Enterprise Agent Platform.

  2. Why it matters

    On the OSWorld benchmark, Gemini 3.5 Flash scores 78.4, beating Gemini 3 Flash (65.1) and GPT-5.4 mini (72.1), putting it among the top-performing models for computer interaction tasks. This opens the door for developers to build agents that automate software testing, office tasks, and browser workflows across multiple device types.

  3. What to watch

    Google has built in two optional enterprise safeguards—one requiring user confirmation for sensitive actions, and another automatically stopping tasks when indirect prompt injections are detected. The company also recommends sandboxing, human oversight, and strict access controls to guard against abuse.

Ask the AI about this article →

FAQ

How does Gemini 3.5 Flash's computer control performance compare to other models?
On the OSWorld benchmark, Gemini 3.5 Flash scores 78.4, beating Gemini 3 Flash (65.1) and GPT-5.4 mini (72.1). GPT-5.5 scores 78.7 and Anthropic's Opus 4.8 leads at 83.4.
How is Google protecting against misuse of this screen-control feature?
Google uses adversarial training and offers two optional enterprise safeguards: one requires user confirmation for sensitive or irreversible actions, while the other automatically stops tasks when it detects indirect prompt injections. Google also recommends sandboxing, human oversight, and strict access controls.
Where can developers access this feature?
The feature is available through the Gemini API and the Gemini Enterprise Agent Platform. A Browserbase demo and a GitHub reference implementation are also available.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • AI's math proofs challenge field's core valueJapan Times Tech · 2h ago
  • AI-written citations 38.6% fabricated, study findsHacker News · 2h ago
  • Lam Research breaks ground on AI chip lab in OregonTop Companies AI · 9h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleReddit targets 1 billion users after going public, AI deals