Researchers from Stanford and Caltech connected the GPT-6 Astra model directly to the Unitree G1 robot in the HomeBody system
The HomeBody system from Stanford and Caltech connects the GPT-6 Astra model directly to the Unitree G1 robot without a traditional control layer. Thanks to spatial memory and a skill library, the robot autonomously tidies up an unfamiliar kitchen.
Researchers from Stanford and Caltech have created a system called HomeBody that allows the humanoid robot Unitree G1 to move autonomously in an unfamiliar kitchen, tidy up, and take objects out of drawers. According to the description, the system skips the usual trained control layer between the language model and the robot — an interchangeable vision-language model, in this case GPT-6 Astra, calls directly into an extensible skill library for grasping objects, navigation, or opening drawers.
The robot first explores the space, creates its digital twin in Nvidia Isaac Sim, and records the position of objects in spatial memory, which allows it to find them even after they disappear from its field of view. For tasks like \"clean the kitchen,\" the model plans individual steps and corrects itself in case of an error.
The authors also state limitations: the latency of the GPT-6 Astra model, overheating of the servomotors in the robot's fingers, and high computing costs. The system's code is published on GitHub. According to the source, earlier benchmarks showed significantly better spatial reasoning for the Astra model, but another test pointed out safety issues when Astra controls the robot. According to the article, OpenAI has already announced plans to return to robotics, including use for personal purposes.
Why it matters
The experiment shows that a language model can control a robot directly, without an expensive trained control layer specific to the given machine — which could simplify the development of more general home robots. At the same time, it is still a research prototype with real limitations: latency, hardware overheating, high compute costs, and documented safety risks when a model controls the robot, so deployment outside the lab is not yet on the table.
Relevant practical impact
What this means
For a business
For companies tracking robotics and automation, this is evidence of a research direction in which a vision-language model replaces an expensive custom robot control layer. However, the source explicitly states high compute costs, latency, and safety issues, so the approach does not yet represent a deployable commercial solution.
Development More business impacts →Check the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.