Google DeepMind has released Gemini Robotics 2, an intelligence layer comprising three models designed to enable whole-body control, five-finger dexterity, and multi-robot collaboration. The release includes a Vision-Language-Action (VLA) model, an embodied reasoning VLM, and an on-device VLA, moving beyond previous table-top manipulation capabilities.

  • Gemini Robotics 2 VLA drives full humanoids like Apptronik Apollo 2 with SharpaWave hands, achieving success rates ranging from 32% to 92% across dexterity tasks.
  • Gemini Robotics ER 2, based on Gemini 3.5 Flash, handles high-level planning and tool orchestration with a 128k context window and bidirectional streaming via the Gemini Live API.
  • Gemini Robotics On-Device 2 allows local execution and rapid adaptation to new robot bodies, such as the SO101 platform, where performance jumped from 6.7% to 53.3% success.
  • The suite supports multi-robot collaboration, demonstrated by pairing Apollo 2 with a Franka F3 Duo for shared semantic task handoffs.

The models aim to address the limitations of pre-programmed robots by enabling adaptation to unpredictable environments and skill transfer across different robot bodies.