Skip to content
worth noting New models verified update

OpenAI claims that GPT-5.6 Sol outperformed Opus 5 on ARC-AGI-3 – but only with its own testing environment

only one source so far updated July 30, 2026

OpenAI states that GPT-5.6 Sol achieved a score of 38.3 % on ARC-AGI-3 compared with 30.2 % for Opus 5 from Anthropic. However, this holds only with the custom API configuration used by OpenAI; in the official test, GPT-5.6 Sol scores only 7.8 %.

OpenAI responded to the earlier result achieved by Opus 5 on the ARC-AGI-3 benchmark by claiming that GPT-5.6 Sol achieves a score of 38.3 %, higher than Opus 5 at 30.2 %. However, OpenAI did not use the official benchmark testing environment, but instead used its own Responses API with two features: “Retained Reasoning”, which preserves the chain of reasoning of the model between steps, and “Compaction”, which summarizes older context instead of deleting it. In the official testing environment, where the reasoning of the model is discarded after each step, GPT-5.6 Sol achieved only 7.8 %. According to OpenAI, no benchmark measures just the model itself; it also measures the technical setup around it. The Decoder points out that ARC-AGI-3 is deliberately designed to test raw model performance without external aids, and that Opus 5 achieved its 30.2 % under the same constraints – meaning without enhancements like those now used by OpenAI.

The original report concerned Opus 5 from Anthropic achieving a score of 30.2 % on ARC-AGI-3, almost four times the previous record of 7.8 % held by GPT-5.6 Sol (Max) from OpenAI. According to the ARC Prize team, which develops the benchmark, this lead is indeed due to stronger logical reasoning that enables more independent exploration, planning, and task execution in unfamiliar environments. During testing, Opus 5, among other things, converted tasks into algebraic notation and independently formulated reflection equations, something no model had previously done, and solved five previously unsolved environments.

What changed

Why it matters

The difference between 7.8 % and 38.3 % for the same GPT-5.6 Sol model shows how strongly benchmark scores are influenced by the technical configuration of the test, rather than just the capabilities of the model itself. When comparing AI models using published benchmarks, you therefore need to verify whether a result comes from the official testing environment or from a custom configuration modified by the manufacturer, because the two figures are not comparable.

What was added since the original report

Verified updates

  1. New verified information

    OpenAI achieves 38.3 % on ARC-AGI-3 with its own API setup (Retained Reasoning + Compaction); In the official harness without these features, GPT-5.6 Sol scores only 7.8 %; ARC-AGI-3 was deliberately designed without external aids to measure raw model performance; Retained Reasoning and Compaction are not part of the official testing environment; Benchmark results reflect both the model and the technical setup around it

    • OpenAI achieves 38.3 % on ARC-AGI-3 with its own API setup (Retained Reasoning + Compaction)
    • In the official harness without these features, GPT-5.6 Sol scores only 7.8 %
    • ARC-AGI-3 was deliberately designed without external aids to measure raw model performance
    • Retained Reasoning and Compaction are not part of the official testing environment
    • Benchmark results reflect both the model and the technical setup around it

Two audiences, two different impacts

What this means

01

For individuals

When reading benchmark rankings, you need to verify that the figures being compared come from the same testing environment – a difference of 7.8 % versus 38.3 % for the same model shows that a score alone, without methodological context, can be misleading.

What to do When comparing AI models using benchmarks, verify that the figures come from the same testing environment and settings.
More practical updates →
02

For a business

Companies selecting an AI model for production deployment should not rely on marketing benchmark figures, because the same model can show significantly different performance depending on which additional API features (e.g. preserving context between steps) the provider used during testing.

Development
What to decide Before selecting an AI model for production, test its performance on your own tasks instead of relying on published benchmark figures.
More business impacts →
AI models ARC-AGI-3 benchmarking Claude Opus 5 GPT-5.6 Sol logical reasoning methodology Opus 5 reasoning

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

2
The Decoder (daily AI news) independent context · first detected Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence The Decoder (daily AI news) independent context OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness