AIToday
Large Language ModelsRoboticsThe Verge AIPublished: Jul 31, 2026, 04:01 JST4 min read

Google DeepMind's Gemini Robotics 2 now controls humanoid robots' entire bodies

Google DeepMind's Gemini Robotics 2 now controls humanoid robots' entire bodies

Key takeaway

  • Google DeepMind has released Gemini Robotics 2, an AI model that enables humanoid robots to control their entire bodies — from walking and crouching to fine-fingered manipulation — rather than just the upper body as before.

  • The update allows robots to perform more complex real-world tasks and work together with other robots to complete multi-step jobs; Google also highlighted improved safety features that detect nearby humans and halt the robot automatically if someone approaches too closely.

3 Key Points

  1. What happened

    Google DeepMind announced Gemini Robotics 2, an updated AI model that can control a humanoid robot's whole body — including feet, hands, and fingertips — whereas the previous version controlled only the upper body. The model now enables robots to walk, crouch, stretch, manipulate objects with five-fingered hands, and perform tasks like picking up items, sealing bags, and unscrewing lightbulbs.

  2. Why it matters

    The upgrade moves humanoid robots closer to completing complex, real-world tasks that require whole-body coordination. Google DeepMind also improved Gemini Robotics ER 2, a vision-language model that helps robots understand their surroundings and process multi-step instructions; this version can now detect nearby humans and trigger safety stops, making it the company's "safest robotics model to date."

  3. What to watch

    Google DeepMind also updated its On-Device Model, which runs locally on robots without internet; it can now adapt faster to robots with "drastically different shapes, sensors and degrees of freedom."

In Depth

Read the full story

Google DeepMind announced an upgrade to its Gemini Robotics AI model line on Thursday, significantly expanding the range of tasks humanoid robots can perform. The flagship update, Gemini Robotics 2, shifts from controlling only a robot's upper body to supporting "whole-body motions" across the entire frame — from feet to fingertips. Videos shared by Google demonstrate the practical impact: Apptronik's Apollo 2 robot can now bend over to pick up a watering can and locate and retrieve specific items from a shelf.

The new model enables a broader set of actions including walking, crouching, stretching, and object manipulation. A key hardware improvement is support for complex five-fingered hands, unlocking dexterity-dependent tasks such as sealing Ziploc bags, tying trash bags, and unscrewing lightbulbs. Google DeepMind notes that while robots "have more to advance in movement speed," the update represents "an important step towards the skills needed to complete more complex, real-world tasks that require whole-body coordination."

Alongside the core model, Google updated Gemini Robotics ER (embodied reasoning), a vision-language model that helps robots analyze their surroundings, process instructions, and execute multi-step sequences. The version 2 update improves performance over extended task durations and adds the ability to "understand when tasks begin and end." The model now also supports collaboration: Google demonstrated a scene in which Apollo 2 instructs a separate dual-arm Google robot to place tools in a bin while cleaning a garage, showing multi-robot coordination in action. Google emphasizes that Gemini Robotics ER 2 is its "safest robotics model to date," equipped to "better detect when humans are nearby, trigger safety tool calls and bring the robot to a safe stop if someone approaches too closely."

Google also refined its Gemini Robotics On-Device Model, which runs locally on a robot without requiring an internet connection. This variant now adapts to new robot embodiments more quickly, including those with "drastically different shapes, sensors and degrees of freedom," broadening its utility across different hardware platforms.

Context & Analysis

Google DeepMind's progression from controlling a robot's upper body to its entire body represents a significant step in making humanoid robots capable of handling real-world work. The announcement on Thursday describes Gemini Robotics 2 as supporting "whole-body motions" — a shift that opens up tasks requiring coordination across the robot's limbs and fingers. The company acknowledges that robots "have more to advance in movement speed," indicating that performance still lags behind a fully capable system, yet this update positions the technology toward practical deployment in environments where multiple robots or complex multi-step actions are needed.

The companion update to Gemini Robotics ER 2 — the vision-language model that interprets instructions and analyzes surroundings — adds a layer of operational maturity. By understanding "when tasks begin and end" and detecting human proximity for safety, the model moves beyond isolated motion control into a more complete robotic system. Google's emphasis on safety (branding it the "safest robotics model to date") signals that deployment concerns around human-robot interaction are part of the development roadmap. The On-Device Model improvement, allowing faster adaptation to robots with different physical configurations, suggests Google is working toward flexibility across hardware partners like Apptronik, whose Apollo 2 robot appears in the company's demonstration videos.

FAQ

What new tasks can robots perform with Gemini Robotics 2?
Robots can now walk, crouch, stretch, bend over to pick up objects like a watering can, retrieve specific items from shelves, seal Ziploc bags, tie trash bags, and unscrew lightbulbs. The model also enables multiple robots of different types to work together — for example, one robot can instruct another to place tools in a bin while cleaning a garage.
What safety features does Gemini Robotics ER 2 include?
Google DeepMind describes it as the company's "safest robotics model to date." It can better detect when humans are nearby, trigger safety tool calls, and bring the robot to a safe stop if someone approaches too closely.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Related Articles

Next articleLinkedIn adds 'AI slop' report button to combat low-quality posts

The AI news that matters, in one minute each morning.

Sign up free