Google's Gemini Robotics 2 brings reasoning to humanoid robots

Google DeepMind introduced Gemini Robotics 2, an AI model designed to give humanoid robots the ability to reason about their surroundings, plan multi-step tasks and control their full bodies rather than executing only pre-programmed instructions. Unlike its predecessor, which focused on upper-body movement, the new model handles walking, bending, balancing, crouching and coordinated hand use. Google also unveiled a companion system, Gemini Robotics ER 2, that acts as a high-level "brain" for understanding spoken instructions, breaking them into sub-tasks, tracking completion and reassigning work to lower-level vision-language-action models when plans need to change. Five-finger robotic hands powered by the system can tie knots and seal zip-lock bags, while simpler two-finger grippers are tuned for warehouse work such as tightly packing boxes. DeepMind says Gemini Robotics ER 2 is also built to coordinate multiple robots working on the same long-horizon job.
Naver builds no-code tools for its ARC robot platform
Naver Labs Europe developed ARCBRAIN Flow, a no-code and low-code solution that lets non-developers configure robot services on top of Naver's existing ARC robot operating system. The platform allows users to combine functions such as delivery routes and completion notifications through a drag-and-drop interface, with a Naver official comparing the ease of use to "moving shapes with a mouse in PowerPoint." Naver filed a trademark application for ARCBRAIN Flow with the Korea Intellectual Property Office in July and is pursuing additional patents in Europe. The tool is intended to extend Naver's robot demonstrations at its robot-friendly 1784 headquarters in Seongnam into ordinary buildings, and the company is evaluating its use at Tokyo Midtown Yaesu, a large Tokyo complex that was not originally designed for robots. ARCBRAIN Flow joins existing ARC components ARC Brain (cloud-based robot control), ARC Eye (vision-based positioning) and ARC Mind (web-based robot OS).
MIT researchers speed up robot decision-making with VLASH
A team led by MIT associate professor Song Han, with collaborators from Nvidia, UC Berkeley, UC San Diego, Caltech and Tsinghua University, developed VLASH, a method that lets robots plan the next movement while finishing the current one. By predicting a robot's future state and feeding that estimate into a vision-language-action model, the system cut reaction delays by more than 30-fold and roughly doubled task speed in a cube-sorting test while maintaining 90% accuracy. The method also reorganized existing training data to reduce fine-tuning time by a factor of five without added compute. The researchers said future work will extend VLASH with world models for more dynamic environments, and the paper is scheduled to be presented at the Intelligent Robots and Systems Conference. Funding came from the MIT-IBM Computing Research Lab, Amazon, the National Science Foundation and Nvidia.
Different bets on the robot software stack
The three developments illustrate distinct strategies for owning the software layer beneath tomorrow's robots. Google is betting that a general-purpose reasoning model can replace scripted control for humanoid hardware, Naver is positioning ARC as an open "Android of robots" platform that any manufacturer can plug into, and MIT's VLASH is a research technique aimed at making any vision-language-action model faster regardless of robot type. Analysts quoted in the Naver coverage argue that whichever company standardizes robot software first could capture the same ecosystem leverage Google secured in mobile. It remains unclear how quickly these platforms will reach commercial deployment, and whether Gemini Robotics 2's full-body control and ARCBRAIN Flow's no-code approach will overlap or compete in shared industrial deployments.
Share this article







