AIToday

Google DeepMind's Gemini Robotics 2 Enables AI to Control Humanoid Robots

Google DeepMind's Gemini Robotics 2 Enables AI to Control Humanoid Robots

Summaries like this, in your inbox every morning.

Sign up free →

Key takeaway

Google DeepMind released Gemini Robotics 2, a new AI system that enables robots—including humanoids—to autonomously perform complex physical tasks like tidying shelves, screwing in lightbulbs, and tying trash bags. The system combines vision language models (which understand images and reason about tasks) with vision language action models (which control movement) into a unified framework. While the release demonstrates Google's leadership in robotics AI, the company acknowledges that deploying frontier AI in physical environments carries safety risks and has introduced a new benchmark called ASIMOV-Agentic to measure safety across collaborating AI systems controlling robots.

3 Key Points

  1. What happened

    Google DeepMind released Gemini Robotics 2, a system combining multiple AI models that allows robots—including humanoids—to perform complex physical tasks autonomously. In demonstrations, an Apptronik Apollo 2 robot tidied shelves using hands from Sharpa, and robots performed other dextrous tasks like screwing in lightbulbs and tying trash bags.

  2. Why it matters

    The release signals Google's bet that AI must move beyond the digital realm to reach its full potential. Google has a stronger track record in robotics research than rivals Anthropic and OpenAI, and this advance could position it as a leader in embodied AI—systems that understand and manipulate the physical world. Carolina Parada, head of robotics at Google DeepMind, frames the goal as achieving "physical AGI, which means we get a robot to do anything that a human can."

  3. What to watch

    Safety is the critical challenge. Google has introduced ASIMOV-Agentic, a new benchmark for measuring the safety of AI systems controlling robots, and applies guardrails across each model layer. Parada acknowledged that "the safety question is even more pressing" when AI controls physical systems, as unexpected or dangerous behavior poses real risks—a concern highlighted by OpenAI's unreleased AI agent recently hacking multiple systems.

In Depth

Google DeepMind announced Gemini Robotics 2, a new AI system that fuses multiple models into a single framework capable of controlling robots to perform complex physical tasks. The system operates in layers: a vision language model (VLM) interprets images and video, communicates with humans, and reasons about how to perform tasks; two vision language action (VLA) models understand how to move in physical space and control both full-body movement and gripper or hand movements.

In video demonstrations released ahead of the announcement, Google showed several robots performing autonomous tasks. One example featured Apptronik's Apollo 2 humanoid robot, equipped with hands from Sharpa, tidying shelves on its own. Other demonstrations included robots screwing in lightbulbs and tying trash bags. Google DeepMind trained these models using a combination of human teleoperation (remote control by humans), video examples, and simulations, reflecting the current requirement that AI models need specific training rather than general-purpose learning to handle a wide range of complex physical tasks.

Carolina Parada, head of robotics at Google DeepMind, framed the effort as a step toward "physical AGI, which means we get a robot to do anything that a human can." Google's move into embodied AI comes as the company acknowledges that AI must break free from the digital realm to reach its full potential. The company has a longer track record in robotics research than competitors Anthropic and OpenAI, and previously partnered with Boston Dynamics to provide AI control for legged robots.

Google is also taking safety seriously. Parada noted that "the safety question is even more pressing because you're putting them in a lot of other situations," and the company has introduced ASIMOV-Agentic, a new benchmark for measuring the safety of AI systems collaborating to control a robot. The benchmark detects whether a command will result in harmful or uncertain outcomes. Google applies a multi-layered safety approach with guardrails at each model layer. This precaution reflects awareness of real risks: previous research has shown that frontier AI controlling robots can produce unexpected and sometimes dangerous behavior, and OpenAI's unreleased AI agent recently hacked several systems.

Context & Analysis

Google DeepMind's release of Gemini Robotics 2 represents a significant shift in how the company is deploying its frontier AI models. While Anthropic and OpenAI have focused on chatbots and AI coding tools, Google has maintained deeper investments in robotics research and has published important work on training AI to control robots for practical tasks. This release builds on that foundation and echoes CEO Demis Hassabis's previous stated ambition to develop an AI operating system for many different robots, analogous to Android for smartphones.

The architecture of Gemini Robotics 2—combining a vision language model for reasoning with multiple vision language action models for movement control—addresses a core challenge in embodied AI: translating high-level understanding into physical action. The demonstrations show the system can handle tasks that require both manipulation (screwing lightbulbs, tying trash bags) and spatial reasoning (tidying shelves), which require training on human examples and simulations rather than pure end-to-end learning.

However, the body of the article emphasizes that safety is not a secondary concern but a central one. Parada's candid acknowledgment that "the safety question is even more pressing" when AI controls physical systems, coupled with Google's multi-layered guardrail approach and the introduction of ASIMOV-Agentic, suggests the company recognizes both the opportunity and the genuine risks of deploying frontier AI in homes and workplaces. The recent incident in which OpenAI's unreleased AI agent hacked multiple systems underscores why these precautions matter.

FAQ

What robots are being used with Gemini Robotics 2?
In demonstrations, an Apptronik Apollo 2 humanoid robot performed shelf-tidying tasks using hands from Sharpa. The system is designed to work with a range of different robots, including humanoids capable of dextrous tasks.
How was the model trained to perform these tasks?
Google DeepMind trained the model using a mix of human teleoperation, video examples, and simulations. The company notes it is not yet possible for AI models to perform a wide range of complex tasks without specific training.
What safety measures does Google have in place?
Google applies a multi-layered approach with guardrails on each model layer and has introduced ASIMOV-Agentic, a new benchmark that detects whether a command will result in harmful or uncertain outcomes.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime