Google reveals Gemini Robotics 2.0, promising improved dexterity and safety
Google DeepMind introduces Gemini Robotics 2, a major upgrade enabling humanoid robots to perform complex tasks with improved dexterity and collaboration capabilities. The new Gemini Robotics ER 2 model processes live video feeds with near 60% accuracy for frame completeness and identifies key task moments with ~90% precision, allowing real-time failure recovery during multi-step operations. A trio of sub-models includes the publicly available ER 2 (vision-language reasoning), Gemini Robotics 2
Analysis
TL;DR
- Google DeepMind introduces Gemini Robotics 2, a major upgrade enabling humanoid robots to perform complex tasks with improved dexterity and collaboration capabilities.
- The new Gemini Robotics ER 2 model processes live video feeds with near 60% accuracy for frame completeness and identifies key task moments with ~90% precision, allowing real-time failure recovery during multi-step operations.
- A trio of sub-models includes the publicly available ER 2 (vision-language reasoning), Gemini Robotics 2 (action generation), and Gemini Robotics On-Device 2 (low-latency offline adaptation), with the latter requiring only ~200 examples to adapt to new robot designs.
- Safety is enhanced through the ASIMOV-Agentic benchmark, which evaluates refusal of unsafe tool calls, task feasibility assessment, and human-assistance triggering, making ER 2 the safest version yet.
Why It Matters
This release marks a significant step toward general-purpose physical AI by enabling robots to understand dynamic environments, recover from errors autonomously, and collaborate safely—critical for deployment in unstructured real-world settings like homes or factories. The public availability of ER 2 lowers barriers for developers to build on embodied reasoning, while the safety benchmarks address critical risks of physical AI interacting with humans.
Technical Details
- Gemini Robotics ER 2: An upgraded vision language model (VLM) that ingests live camera feeds to track task progress, achieving ~60% accuracy in classifying video frame completeness and ~90% in identifying critical task moments (e.g., stopping coffee pour). Enables real-time error correction without restarting entire sequences.
- Collaboration Architecture: ER 2’s improved world understanding allows multiple robots (e.g., Apptronik Apollo 2 and Franka F3 Duo) to coordinate actions without interference, reducing hesitation observed in prior versions.
- Action Models: Gemini Robotics 2 generates precise motor commands from high-level instructions via generative modeling; its on-device variant (On-Device 2) adapts to novel robot hardware using minimal data (~200 motion examples) with low latency.
- Safety Framework: Integrates traditional physical safeguards with the ASIMOV-Agentic benchmark, which tests refusal of unsafe VLA tool calls, assesses task safety feasibility, and triggers human intervention when uncertainty exceeds thresholds.
Industry Insight
The modular design of Gemini Robotics 2—separating perception (ER 2), action generation (Robotics 2), and edge adaptation (On-Device 2)—provides a blueprint for scalable robot development where components can be updated independently. Public access to ER 2 will likely accelerate third-party innovation in embodied AI, while the ASIMOV-Agentic standard may become an industry benchmark for evaluating physical AI safety, pushing competitors to prioritize similar rigorous testing protocols before deployment.
Disclaimer: The above content is generated by AI and is for reference only.