AIToday
Large Language ModelsRoboticsRobotics & Automation NewsPublished: Jul 31, 2026, 22:00 JST5 min read

Google DeepMind releases Gemini Robotics 2 with whole-body AI control for humanoids

Google DeepMind releases Gemini Robotics 2 with whole-body AI control for humanoids

Key takeaway

  • Google DeepMind has released Gemini Robotics 2, a new suite of AI models designed to give robots greater autonomy through whole-body control and multi-robot collaboration.

  • Demonstrated on Apptronik's Apollo 2 humanoid, the system lets robots perform complex tasks—from walking and object manipulation to reasoning through changing environments—without pre-programming.

  • The platform can adapt to new robot hardware in hours and includes safety features that enable robots to refuse unsafe actions and request human intervention when needed.

3 Key Points

  1. What happened

    Google DeepMind launched Gemini Robotics 2, a suite of three AI models enabling humanoid robots to move beyond pre-programmed tasks. The system is demonstrated on Apptronik's Apollo 2 humanoid, which can now perform full-body autonomous movements—walking, crouching, bending, manipulating objects—while reasoning through complex tasks in real time. Gemini Robotics ER 2 also enables multiple robots to collaborate on workflows.

  2. Why it matters

    The platform adds whole-body control (extending beyond earlier upper-body focus) and can be adapted to new robot hardware in a matter of hours, meaning developers can deploy the system across different robot types faster. For industrial and logistics settings, the on-device model runs directly on robot hardware without cloud dependency, requiring fewer than 200 training examples over just a few hours to adapt to new dual-arm platforms. The release also includes ASIMOV-Agentic, a safety benchmark that evaluates whether robots can refuse unsafe actions and request human help when needed.

  3. What to watch

    Gemini Robotics ER 2 is now available through Google AI Studio and in private preview on the Gemini Enterprise Agent Platform; Gemini Robotics 2 and Gemini Robotics On-Device 2 are available to selected early-access partners. Apptronik's Robot Park facility, launched earlier this month, is collecting data to train next-generation versions of these models.

In Depth

Read the full story

Google DeepMind unveiled Gemini Robotics 2, a comprehensive suite of AI models engineered to grant robots full-body autonomy, sophisticated dexterity, and the ability to collaborate across multiple machines. The system is being demonstrated on Apptronik's Apollo 2 humanoid robot, which can now execute full-body autonomous movements—walking, crouching, bending, and manipulating objects—while simultaneously reasoning through complex tasks in real time.

The release comprises three distinct models. Gemini Robotics 2 itself is a vision-language-action (VLA) model that translates visual and language inputs into robot movements, controlling both full humanoid robots and dual-arm systems. Gemini Robotics ER 2 is an embodied reasoning model capable of planning multi-step tasks, communicating with people, and enabling multiple robots to collaborate on workflows. Gemini Robotics On-Device 2 is an optimized VLA model that runs directly on robot hardware without requiring cloud connectivity, and can be adapted to new robot embodiments with only a few hours of training data.

A key advancement in Gemini Robotics 2 is whole-body control. Previous versions focused primarily on upper-body manipulation; the new model enables humanoids to coordinate movements from "feet to fingertips." In a demonstration, Apollo 2 receives the instruction "put the watering can into the green bin in the bottom shelf," then walks across the room, picks up the object, and places it in the correct location. The system also demonstrates improved dexterity, controlling Apptronik's Apollo 2 equipped with five-fingered SharpaWave hands to perform tasks such as tying a trash bag or unscrewing a light bulb. For industrial settings, the model supports conventional parallel grippers for precision insertion and packing.

Multi-robot collaboration represents another new capability. Gemini Robotics ER 2 allows different robot types to communicate and coordinate on complex workflows that would otherwise require a single machine to perform every task. The reasoning model is designed to manage task execution over several minutes, monitoring progress, recovering from failures, and adapting to changing situations. Google DeepMind describes the overall platform as the "intelligence layer" for robots, enabling machines to move beyond pre-programmed tasks and adapt to changing environments, and says it can be adapted to new hardware in a matter of hours.

Safety is integral to the release. Google DeepMind introduced ASIMOV-Agentic, a benchmark designed to evaluate "agentic safety orchestration and uncertainty resolution," measuring a robot's ability to refuse unsafe actions, determine when a task cannot be completed safely, and request human intervention when appropriate. Gemini Robotics ER 2 also improves human-awareness capabilities, enabling robots to detect nearby people, trigger safety functions, and stop safely when someone enters their working area. The announcement builds on the companies' ongoing collaboration; earlier this month, Apptronik launched Robot Park, a facility where Apollo 2 robots collect data used to train next-generation AI models including Gemini Robotics 2. Gemini Robotics ER 2 is now available through Google AI Studio and in private preview on the Gemini Enterprise Agent Platform, while Gemini Robotics 2 and Gemini Robotics On-Device 2 are available to selected early-access partners.

Context & Analysis

Google DeepMind's release of Gemini Robotics 2 marks a shift from task-specific robotic control to whole-body autonomous reasoning. The announcement builds on an established partnership with Apptronik, which launched Robot Park earlier this month—a data collection facility designed to train next-generation AI models including this suite. The three-model architecture (a vision-language-action model, an embodied reasoning module, and an on-device optimized version) addresses different deployment scenarios: cloud-connected settings, multi-robot coordination, and edge-constrained industrial environments where connectivity is limited.

The capability improvements reported—whole-body movement coordination ("feet to fingertips"), dexterity with five-fingered hands for tasks like tying trash bags, and cross-platform adaptability in hours—signal that the barrier to deploying generalist robot control is dropping. For industrial deployments, the on-device model's ability to adapt with fewer than 200 training examples over a few hours suggests a meaningful reduction in setup friction. The inclusion of ASIMOV-Agentic, a safety benchmark for agentic reasoning, indicates Google DeepMind is embedding safety evaluation into the product rather than treating it as a downstream concern—a practical necessity for real-world robotics deployment.

FAQ

What three models make up Gemini Robotics 2?
Gemini Robotics 2 is a vision-language-action model that converts visual and language inputs into robot movements; Gemini Robotics ER 2 is an embodied reasoning model that plans multi-step tasks, communicates with people, and enables multiple robots to work together; and Gemini Robotics On-Device 2 is an optimized model that runs locally on robotic hardware and can be adapted to new robot embodiments using only a few hours of training data.
How quickly can Gemini Robotics adapt to new robot hardware?
According to Google DeepMind, the platform can be adapted to new hardware in a matter of hours. Specifically, Gemini Robotics On-Device 2 can be adapted to new dual-arm robot platforms with fewer than 200 training examples collected over just a few hours.
What safety features does Gemini Robotics 2 include?
Google DeepMind has introduced ASIMOV-Agentic, a benchmark designed to evaluate a robot's ability to refuse unsafe actions, determine when a task cannot be completed safely, and request human intervention when appropriate. Gemini Robotics ER 2 also improves human-awareness capabilities, enabling robots to detect nearby people, trigger safety functions, and stop safely when someone enters their working area.
Robotics & Automation NewsRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Related Articles

Next articleBausch & Lomb CEO: AI hype echoes past tech cycles — what matters is people

The AI news that matters, in one minute each morning.

Sign up free