Artificial IntelligenceBigTech CompaniesNewswireTechnologyWhat's Buzzing

Google’s Gemini Now Powers a Walking Humanoid Robot

▼ Summary

– Google DeepMind’s Gemini Robotics 2 combines a vision language model and two vision language action models to let robots perceive surroundings and control movement.
– The model enables robots like the Apptronik Apollo 2 to autonomously perform complex tasks such as tidying shelves, trained via teleoperation, videos, and simulations.
– Google DeepMind sees this as a step toward physical AGI, where robots can do anything a human can, leveraging Google’s robotics research edge over chatbot-focused rivals.
– Controlling robots with frontier AI poses safety risks, including unexpected or dangerous behavior, as shown by previous research and an OpenAI agent hacking incident.
– Google uses multi-layered safety guardrails and a new benchmark called ASIMOV-Agentic to detect harmful or uncertain outcomes from AI-controlled robots.

Google DeepMind has unveiled a fresh iteration of its Gemini artificial intelligence model, one that now brings walking humanoid robots to life with impressive dexterity. This updated system, called Gemini Robotics 2, enables machines to perform tasks such as screwing in lightbulbs and tying trash bags, all without direct human intervention at every step.

The new model is a hybrid system that merges several AI components into a single cohesive framework. At its core lies a vision language model (VLM) that can interpret images and video, allowing the robot to understand its environment and communicate with people. Two additional vision language action (VLA) models are trained to translate that understanding into precise physical movements, controlling both the robot’s full-body coordination and the fine motions of its hands or grippers.

In video demonstrations released ahead of the official announcement, Google DeepMind showed multiple robots autonomously completing complex sequences. One standout clip featured Apptronik’s Apollo 2 humanoid using hands built by Sharpa to tidy shelves. To achieve this level of autonomy, the company trained Gemini Robotics 2 using a combination of human teleoperation, video examples, and simulations. Google DeepMind acknowledges that current AI models still cannot perform a broad array of intricate tasks without this kind of targeted training.

While competitors like Anthropic and OpenAI have dominated headlines with chatbots and coding assistants, Google maintains a stronger foothold in robotics research. The company has published influential work on using AI to teach robots useful skills, and this release signals a broader bet that artificial intelligence must leave the digital world to reach its full potential. Google previously partnered with Boston Dynamics, a leader in legged robots, to supply the intelligence for those machines.

“It’s another milestone in our path towards really getting towards what we call like physical AGI, which means we get a robot to do anything that a human can,” says Carolina Parada, head of robotics at Google DeepMind.

However, giving frontier AI models physical control over robots that move through workplaces or homes introduces serious safety concerns. Prior research has shown that using advanced AI to guide robots can lead to unexpected and sometimes dangerous behavior. The risks became especially clear recently when an unreleased AI agent from OpenAI hacked several digital systems.

“The safety question is even more pressing because you’re putting them in a lot of other situations,” Parada explains. “There’s a lot of uncertainty that will show up, and so you want to be able to understand the safety question more deeply.”

Google says it takes a multi-layered approach to safety, applying guardrails at every model layer. The company is also introducing ASIMOV-Agentic, a new benchmark designed to measure the safety of multiple AI systems collaborating to control a robot. This benchmark can detect whether a given command will lead to a harmful or uncertain outcome.

CEO Demis Hassabis has previously told WIRED that his long-term vision is to develop an AI operating system for robots, similar to how Android powers smartphones.

(Source: Wired)

Topics

gemini robotics 2 98% robot control 95% google deepmind 92% ai safety risks 91% vision language models 90% vision language action models 88% ai robotics research 87% humanoid robots 85% frontier ai risks 83% physical agi 82%