Skip to content
worth noting New models

StationeryBench benchmark: GPT-6 Astra model significantly outperformed MolmoAct2 in spatial reasoning for robots

only one source so far

The StationeryBench benchmark compared the GPT-6 Astra model from OpenAI with the MolmoAct2 model from Ai2 on robotic tasks involving objects on a table. Astra achieved a median of 46/100 and completed 7 out of 100 tasks, while MolmoAct2 achieved a median of 12/100 without completing a single task.

The new robotics benchmark StationeryBench compared the GPT-6 Astra model from OpenAI with the MolmoAct2 model from Ai2 on five tasks involving objects on a table—for example, unscrewing the cap of a marker, pouring out paper clips, or passing a ruler between two robotic arms. Both models controlled the same dual-arm YAM robots across 200 trials. The GPT-6 Astra model fully completed 7 out of 100 tasks, while the MolmoAct2 model completed none. The median progress score was 46 out of 100 points for GPT-6 Astra and 12 out of 100 for MolmoAct2. The results, videos, and benchmark code are published on GitHub.

Yoav Artzi, a researcher at Cornell and Google DeepMind, described the result as a “step change” in spatial reasoning. According to him, the GPT-6 Astra model achieves near-human accuracy on the as-yet-unpublished REMAP benchmark, although he added that “even Astra does not achieve what humans can do in other scenarios”. Artzi believes the model was probably trained on a large amount of 3D data, such as scenes from Blender, which would be consistent with its significant improvement on tasks requiring spatial reasoning—but this is his estimate, not information confirmed by OpenAI.

According to available information, OpenAI has long-term plans to develop its own consumer robots.

What changed

Why it matters

The result shows a concrete, measurable advance in the ability of a general-purpose AI model to control a real robotic arm during manipulation tasks, which is relevant to researchers and companies developing robotics or embodied AI. At the same time, this is a benchmark with a low overall success rate (only 7 out of 100 tasks fully completed) and unpublished methodological details (REMAP), so the results should be treated as an early indication of the direction of development, not as evidence of readiness for practical deployment.

Relevant practical impact

What this means

01

For a business

For companies developing robotics or embodied AI applications, the result suggests that general-purpose models are increasingly capable of handling spatial/manipulation tasks, but this is an early benchmark with a low success rate (7 % of tasks fully solved), not a solution ready for deployment.

Development
What to decide Monitor developments in spatial reasoning benchmarks for robotic AI models (e.g. StationeryBench, REMAP) before deciding whether to deploy similar models for manipulation tasks.
More business impacts →
benchmark GPT-6 Astra OpenAI spatial reasoning robotics

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
The Decoder (daily AI news) independent context · first detected GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks