Robotics · 3 min read

Feet to Fingertips: Gemini Robotics 2 Publishes Whole-Body Humanoid Metrics

Google DeepMind's July 30 release brings whole-body VLA control to humanoids, publishes success rates up to 76% on shelf pickups, and ships the ASIMOV-Agentic safety benchmark.

By Classy AI News · August 2, 2026

Feet to Fingertips: Gemini Robotics 2 Publishes Whole-Body Humanoid Metrics

Tabletop manipulation was never going to be enough. Warehouses, homes, and factory floors demand robots that walk, crouch, balance, and reach — often in the same task. On July 30, 2026, Google DeepMind released Gemini Robotics 2, a three-model stack that extends vision-language-action control from arms and grippers to full humanoid bodies.

The announcement, published on DeepMind's official blog, centers on demonstrations with Apptronik's Apollo 2 humanoid — but the architecture is explicitly multi-embodiment, running the same checkpoint across different hands and platforms.

Robotic arm in an industrial setting representing physical AI deployment

Three models, three layers

DeepMind ships Gemini Robotics 2 as a suite rather than a monolith:

Gemini Robotics 2 (VLA) converts vision and language into motor commands across an entire humanoid — feet to fingertips — plus bi-arm robots. It handles both multi-finger hands and parallel grippers.

Gemini Robotics ER 2 (embodied reasoning) acts as the high-level planner: interpreting instructions, communicating with humans, orchestrating multi-step tasks lasting several minutes, and coordinating multi-robot teams.

Gemini Robotics On-Device 2 is an efficient VLA optimized for local inference, adapting to new robot embodiments in "a few hours" with fewer than 200 examples according to the blog.

Access tiers differ. Gemini Robotics ER 2 is available in public preview via Google AI Studio and Gemini Enterprise Agent Platform. The VLA and on-device models remain gated to early-access partners.

Whole-body numbers — honest ones

DeepMind published success rates rather than cherry-picked demo reels alone. On Apollo 2 with Inspire hands, the blog reports:

  • 76.3% shelf pickups
  • 68.4% table pickups
  • 45.7% floor pickups — tasks requiring crouch and balance

Dexterous multi-finger tasks remain uneven. Unscrewing a light bulb hit 92%, but tying a trash bag landed at 44%, bag sealing at 40%, and dustpan use at 32%. On Franka Duo with a two-finger gripper, precise insertion reached 89.6%.

The pattern is clear: locomotion-plus-grasp is maturing; fine finger dexterity is not.

Modern factory floor representing robotics deployment environments

Safety as a benchmark, not a bullet point

DeepMind also released ASIMOV-Agentic, a benchmark for agentic safety orchestration hosted on Hugging Face under CC-BY-4.0. It measures whether embodied reasoning agents refuse unsafe tool calls from VLAs and request human intervention when uncertain.

Gemini Robotics ER 2 is described as DeepMind's "safest robotics model to date" on human-proximity benchmarks — triggering stops when people approach closely enough to matter for collaborative safety standards.

That framing reflects a broader industry shift: as humanoids leave staged demos, safety evaluation becomes as publishable as success rates.

What changes for the field

Gemini Robotics 2 does not solve general-purpose physical AI. Movement speed remains below human baseline in published videos, and multi-finger reliability is too low for unstructured homes.

But the capability envelope expanded in one release: from upper-body VLA to whole-body coordination under a single learned policy, with a reasoning layer that can assign subtasks across robot types.

Competitors — including recent industrial humanoid deployments from Apptronik partners and Chinese manufacturers navigating U.S. import restrictions — will benchmark against these published rates, not just demo videos.

Automated manufacturing equipment representing multi-robot workflows

For developers, the near-term entry point is Gemini Robotics ER 2 in AI Studio. For hardware partners, the gated VLA is the prize — whole-body control without training separate policies per limb.

The blog closes with an AGI-in-the-physical-world framing. The verified facts underneath are narrower and more useful: published metrics, a public safety benchmark, and whole-body control that works often enough to measure — not just to film.

Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.