StationeryBench benchmark: GPT-6 Astra model significantly outperformed MolmoAct2 in spatial reasoning for robots
The StationeryBench benchmark compared the GPT-6 Astra model from OpenAI with the MolmoAct2 model from Ai2 on robotic tasks involving objects on a table. Astra achieved a median of 46/100 and completed 7 out of 100 tasks, while MolmoAct2 achieved a median of 12/100 without completing a single task.
The new robotics benchmark StationeryBench compared the GPT-6 Astra model from OpenAI with the MolmoAct2 model from Ai2 on five tasks involving objects on a table—for example, unscrewing the cap of a marker, pouring out paper clips, or passing a ruler between two robotic arms. Both models controlled the same dual-arm YAM robots across 200 trials. The GPT-6 Astra model fully completed 7 out of 100 tasks, while the MolmoAct2 model completed none. The median progress score was 46 out of 100 points for GPT-6 Astra and 12 out of 100 for MolmoAct2. The results, videos, and benchmark code are published on GitHub.
Yoav Artzi, a researcher at Cornell and Google DeepMind, described the result as a “step change” in spatial reasoning. According to him, the GPT-6 Astra model achieves near-human accuracy on the as-yet-unpublished REMAP benchmark, although he added that “even Astra does not achieve what humans can do in other scenarios”. Artzi believes the model was probably trained on a large amount of 3D data, such as scenes from Blender, which would be consistent with its significant improvement on tasks requiring spatial reasoning—but this is his estimate, not information confirmed by OpenAI.
According to available information, OpenAI has long-term plans to develop its own consumer robots.
Why it matters
The result shows a concrete, measurable advance in the ability of a general-purpose AI model to control a real robotic arm during manipulation tasks, which is relevant to researchers and companies developing robotics or embodied AI. At the same time, this is a benchmark with a low overall success rate (only 7 out of 100 tasks fully completed) and unpublished methodological details (REMAP), so the results should be treated as an early indication of the direction of development, not as evidence of readiness for practical deployment.
Relevant practical impact
What this means
For a business
For companies developing robotics or embodied AI applications, the result suggests that general-purpose models are increasingly capable of handling spatial/manipulation tasks, but this is an early benchmark with a low success rate (7 % of tasks fully solved), not a solution ready for deployment.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.