OpenAI claims that GPT-5.6 Sol outperformed Opus 5 on ARC-AGI-3 – but only with its own testing environment
OpenAI states that GPT-5.6 Sol achieved a score of 38.3 % on ARC-AGI-3 compared with 30.2 % for Opus 5 from Anthropic. However, this holds only with the custom API configuration used by OpenAI; in the official test, GPT-5.6 Sol scores only 7.8 %.
OpenAI responded to the earlier result achieved by Opus 5 on the ARC-AGI-3 benchmark by claiming that GPT-5.6 Sol achieves a score of 38.3 %, higher than Opus 5 at 30.2 %. However, OpenAI did not use the official benchmark testing environment, but instead used its own Responses API with two features: “Retained Reasoning”, which preserves the chain of reasoning of the model between steps, and “Compaction”, which summarizes older context instead of deleting it. In the official testing environment, where the reasoning of the model is discarded after each step, GPT-5.6 Sol achieved only 7.8 %. According to OpenAI, no benchmark measures just the model itself; it also measures the technical setup around it. The Decoder points out that ARC-AGI-3 is deliberately designed to test raw model performance without external aids, and that Opus 5 achieved its 30.2 % under the same constraints – meaning without enhancements like those now used by OpenAI.
The original report concerned Opus 5 from Anthropic achieving a score of 30.2 % on ARC-AGI-3, almost four times the previous record of 7.8 % held by GPT-5.6 Sol (Max) from OpenAI. According to the ARC Prize team, which develops the benchmark, this lead is indeed due to stronger logical reasoning that enables more independent exploration, planning, and task execution in unfamiliar environments. During testing, Opus 5, among other things, converted tasks into algebraic notation and independently formulated reflection equations, something no model had previously done, and solved five previously unsolved environments.
Why it matters
The difference between 7.8 % and 38.3 % for the same GPT-5.6 Sol model shows how strongly benchmark scores are influenced by the technical configuration of the test, rather than just the capabilities of the model itself. When comparing AI models using published benchmarks, you therefore need to verify whether a result comes from the official testing environment or from a custom configuration modified by the manufacturer, because the two figures are not comparable.
What was added since the original report
Verified updates
-
OpenAI achieves 38.3 % on ARC-AGI-3 with its own API setup (Retained Reasoning + Compaction); In the official harness without these features, GPT-5.6 Sol scores only 7.8 %; ARC-AGI-3 was deliberately designed without external aids to measure raw model performance; Retained Reasoning and Compaction are not part of the official testing environment; Benchmark results reflect both the model and the technical setup around it
- OpenAI achieves 38.3 % on ARC-AGI-3 with its own API setup (Retained Reasoning + Compaction)
- In the official harness without these features, GPT-5.6 Sol scores only 7.8 %
- ARC-AGI-3 was deliberately designed without external aids to measure raw model performance
- Retained Reasoning and Compaction are not part of the official testing environment
- Benchmark results reflect both the model and the technical setup around it
Two audiences, two different impacts
What this means
For individuals
When reading benchmark rankings, you need to verify that the figures being compared come from the same testing environment – a difference of 7.8 % versus 38.3 % for the same model shows that a score alone, without methodological context, can be misleading.
For a business
Companies selecting an AI model for production deployment should not rely on marketing benchmark figures, because the same model can show significantly different performance depending on which additional API features (e.g. preserving context between steps) the provider used during testing.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.