Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
read at source ↗ deepmind.google
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Source: DeepMind Date: 2026-07-30 URL: https://deepmind.google/blog/gemini-robotics-er-2-powering-robotics-with-video-understanding-task-orchestration-and-multi-robot-collaboration/
Summary
DeepMind’s technical writeup for Gemini Robotics-ER 2, positioned as a “high-level brain” giving robots real-time spatial reasoning: video understanding with task-progress tracking, multi-step planning that orchestrates lower-level robot controllers, multi-robot collaboration, and API/function tool-calling at sub-second latency. Against predecessor ER 1.6 it improves progress classification (57.4%), moment-finding precision (91.3%), and human-proximity safety benchmarks. Available now via the Gemini API and AI Studio, private preview on the Enterprise Agent Platform.
Implications
A genuine capability step in the physical-AI/robotics lane — orchestration-plus-multi-robot-collaboration is new ground versus prior ER releases’ single-robot reasoning, and it lands alongside NVIDIA Cosmos and other robot-data efforts tracked separately, reinforcing physical AI as a persistent (not one-off) thread. This is the technical companion to the same-day general-audience post (“Introducing Gemini Robotics-ER 2”); read together they’re one launch split across two blogs, not two events.