
Google DeepMind unveiled Gemini Robotics 2, a system that gives robots intelligent whole-body control to perform complex, multi-step real-world tasks—from walking and crouching to manipulating delicate objects like tying knots.
The model runs locally on robot hardware and adapts to new robot types in just a few hours, addressing a key need for autonomous systems in homes and workplaces without relying on constant internet connectivity.
What happened
Google DeepMind released Gemini Robotics 2, a vision-language-action model that enables robots to control entire bodies—including whole-body movement and dexterous manipulation—rather than just upper-body tasks. The system can control a five-fingered, 22 degree-of-freedom hand and allows multiple robots to collaborate on complex tasks lasting several minutes.
Why it matters
Robots can now execute real-world tasks requiring whole-body coordination, such as walking to retrieve an object, bending to a shelf, and placing it precisely. The model runs locally on devices and can adapt to entirely new robot types in just a few hours, making physical AI systems practical for homes and workplaces without requiring constant network connectivity.
What to watch
Gemini Robotics ER 2 (the reasoning model) is available now on Google AI Studio and in private preview on Gemini Enterprise Agent Platform; the VLA and On-Device models are available to early-access partners. DeepMind also introduced ASIMOV-Agentic, a new safety benchmark for agentic robots operating near humans.
Google DeepMind unveiled Gemini Robotics 2 last week, a significant upgrade to its vision-language-action (VLA) models that shifts the scope of robot control from isolated upper-body tasks to intelligent whole-body manipulation. The system combines advanced perception, reasoning, and motor control to enable robots to understand natural-language instructions and execute them through coordinated full-body movement.
In practical terms, Gemini Robotics 2 can instruct a humanoid robot to walk across a room, locate a watering can on a table, pick it up, navigate to a shelf unit, and place the can precisely in a green bin on the bottom shelf—all from a single user request. When controlling Apptronik's Apollo 2 humanoid, the model manages the robot's five-fingered, 22 degree-of-freedom SharpaWave hand, enabling delicate actions such as tying knots or sealing a ziplock bag. The system also works with simpler two-fingered grippers on the Franka Duo platform for tasks like tight packing. DeepMind notes that while robots still need to improve movement speed, whole-body coordination is essential for real-world utility in homes and workplaces.
DeepMind released three interconnected components. Gemini Robotics ER 2 is a separate embodied reasoning (reasoning model) that serves as the robot's planning layer, processing user instructions, observing the environment, reasoning about task steps, coordinating with the VLA to execute actions, and tracking progress until completion. This model now understands when tasks begin and end and can pinpoint key events, allowing robots to execute complex multi-step sequences lasting several minutes and involving hundreds of decisions. It can also self-correct if a step fails and generalize to novel situations. For the first time, DeepMind is enabling multi-robot collaboration, where different robot types communicate and work together on workflows that a single robot could not accomplish alone. Gemini Robotics On-Device 2 is optimized to run locally on robotic hardware without requiring network connectivity, addressing applications where latency or internet access is constrained. It inherits motion-transfer techniques from Gemini Robotics 1.5 and can adapt to new robot embodiments—even those with drastically different shapes, sensors, and degrees of freedom—in just a few hours, typically using fewer than 200 examples.
DeepMind emphasized safety as foundational to the release. The company introduced ASIMOV-Agentic, a new benchmark for agentic safety orchestration and uncertainty resolution, which measures the reasoning agent's ability to refuse unsafe tool calls and to predict whether a task is possible, proactively requesting human intervention when uncertain. Gemini Robotics ER 2 is described as DeepMind's safest robotics model to date in safety constraint following and human proximity benchmarks; it can detect nearby humans, trigger safety tool calls, and bring the robot to a safe stop if someone approaches too closely. Gemini Robotics ER 2 is now available on Google AI Studio and in private preview on Gemini Enterprise Agent Platform, while the VLA and On-Device models are available to early-access partners.
DeepMind's earlier Gemini Robotics models focused on upper-body manipulation for tabletop tasks—a significant but narrow application. Gemini Robotics 2 extends that capability to the entire robot body, enabling tasks that require coordinated walking, crouching, reaching, and precise manipulation. This shift is important because real-world household and workplace tasks rarely involve only stationary, arm-only actions; they require movement through space and dexterous interaction with objects in variable layouts.
The introduction of Gemini Robotics ER 2 as a separate reasoning layer adds another layer of capability. Rather than controlling moment-to-moment actions, this model acts as a "high-level brain" that understands multi-step task sequences lasting several minutes and involving hundreds of decisions. It communicates with humans, observes context, self-corrects when steps fail, and—for the first time in DeepMind's robotics line—enables multiple robots to work together. The ability to track when tasks begin and end represents a measurable step forward in how robots understand progress.
The on-device variant addresses a practical constraint: many robotic applications cannot rely on network connectivity or tolerate latency. By running locally and adapting to new embodiments in hours rather than days or weeks, Gemini Robotics On-Device 2 lowers the barrier to deploying physical AI in settings where internet access is unreliable. DeepMind's introduction of ASIMOV-Agentic, a safety benchmark for agentic orchestration, signals that the company is building guardrails alongside capability—measuring both refusal of unsafe actions and proactive requests for human intervention when uncertainty arises.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
OpenAI is previewing Ultrafast mode, powered by Cerebras infrastructure from a ten-billion-dollar partnership…

Snowflake's Observe announced general availability of a redesigned MCP server and new CLI tool that give AI ag…

Tim O'Reilly, the publisher and tech pioneer, is promoting open-source AI—not just open-weight models, but the…

Kog, a French startup founded by Gaël Delalleau, is using deep GPU-level software optimization to accelerate A…

Meta released Glimmer, an open-weight AI model anyone can download and run on their own hardware, this week

AWS published a technical guide demonstrating how to build multi-agent systems that combine models hosted on A…

The AI news that matters, in one minute each morning.
Sign up free