OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings
Published · Jul 30 · Thu Source · The Decoder

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings

OpenAI claims GPT-5.6 Sol reached 38.3% on ARC-AGI-3 using custom API settings, outperforming Anthropic's Opus 5. Yet, the model scored 7.8% in the official test environment, sparking debate over benchmark neutrality.

KeywordsOpenAIAnthropicGPTAPIGPT-5.6SolOpusARC-AGI-3

OpenAI asserts its GPT-5.6 Sol model outperformed Anthropic's Opus 5 on the ARC-AGI-3 benchmark. The reported score of 38.3% relies on specific API configurations rather than the standardized testing protocol.

Under the official ARC Prize test setup, GPT-5.6 Sol reportedly achieved only 7.8%. This significant gap highlights potential inconsistencies in how different providers interact with the evaluation framework.

The ARC Prize maintains its environment is provider-neutral, though critics suggest the testing infrastructure might rely on outdated API versions. This situation underscores ongoing challenges in establishing standardized metrics for frontier model capabilities.

Such disputes complicate comparisons between leading AI systems. Industry observers will likely scrutinize future benchmark submissions to ensure reproducibility and fairness across competing large language models.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.