OpenAI says GPT-5.6 Sol achieved 38.3% on the ARC-AGI-3 benchmark, surpassing Claude Opus 5’s previously reported 30.2% score. According to the company, the improvement was achieved by evaluating the model through its Responses API, which supports advanced runtime features designed to preserve reasoning across multi-step tasks.
The company emphasized that the higher score reflects how modern reasoning models perform in real-world deployments rather than under restrictive benchmark conditions.
However, the evaluation was not conducted using the official ARC-AGI-3 testing harness, making direct comparisons with official leaderboard scores a topic of discussion.
Two API Features Made the Difference
OpenAI attributed the performance increase to two optional Responses API capabilities:
Retained Reasoning
This feature allows GPT-5.6 Sol to preserve its reasoning process between successive actions instead of restarting from scratch after every step. By maintaining intermediate reasoning, the model can solve complex tasks more effectively.
Compaction
Compaction intelligently summarizes older context rather than discarding it when conversations become longer. This enables the model to retain important information while staying within context limits.
According to OpenAI, combining these two features significantly improves the model’s ability to complete difficult reasoning challenges included in ARC-AGI-3.
Official Benchmark Produced a Much Lower Score
When GPT-5.6 Sol was evaluated using the official ARC-AGI-3 benchmark environment, OpenAI reported that the model achieved only 7.8%.
The primary reason, according to the company, is that the official testing harness removes intermediate reasoning after each action, preventing the model from carrying forward its thought process throughout a task.
OpenAI argues that this evaluation method does not accurately represent how reasoning models are used in production environments where context preservation is available.
Benchmark Methodology Sparks Debate
The announcement has renewed industry discussion about how AI benchmarks should measure reasoning performance.
OpenAI believes benchmark results should account for the complete inference environment, including runtime features that are available to developers through production APIs.
Supporters of standardized testing argue that benchmarks should isolate model capability from implementation-specific optimizations, ensuring every provider is evaluated under identical conditions.
The debate reflects a broader challenge facing AI benchmarking as modern models increasingly depend on sophisticated inference systems rather than model weights alone.
ARC Prize Responds to OpenAI’s Results
Following OpenAI’s announcement, ARC Prize co-founder François Chollet responded by clarifying the project’s evaluation policy.
Chollet explained that evaluation harnesses specifically designed to improve ARC-AGI-3 performance—or those containing benchmark-specific knowledge—are not considered valid for official leaderboard comparisons.
However, he also stated that general-purpose API capabilities available to all developers are acceptable, provided they were not created specifically for ARC-AGI-3 and that benchmark settings and inference costs are reported transparently.
He acknowledged ongoing discussions between ARC Prize and OpenAI regarding evaluation methodology, particularly around context compaction, and welcomed continued efforts to develop testing approaches that better reflect practical AI deployments.
The Race for Better AI Reasoning Continues
The latest benchmark results demonstrate how competition among leading AI companies is shifting beyond model architecture to include inference infrastructure and runtime capabilities.
As models become increasingly capable of multi-step reasoning, benchmark organizations face growing pressure to ensure evaluation methods remain both fair and representative of real-world usage.
Whether future leaderboards adopt advanced runtime features as part of standard testing remains uncertain. What is clear, however, is that benchmarking AI systems is becoming as much about how models are deployed as it is about how they are trained.
For developers and researchers, OpenAI’s latest results highlight the evolving nature of AI evaluation and the importance of transparent, standardized methodologies when comparing the world’s most advanced reasoning models.


