Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
The Decoder · 2026/7/30 09:03:11

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings
AI 中文解读
OpenAI在最新一轮AI逻辑推理测试中放出大招,声称自家的GPT-5.6 Sol模型通过两项特殊设置,得分反超了竞争对手Claude Opus 5。不过,这次“胜利”全靠OpenAI自己定制的测试环境,而非官方标准,引发了公平性争议。
通俗来说,ARC-AGI-3就像一场要求AI“从头到尾”独立解谜的考试。官方测试规定,AI每次推理后都清空内存,不能保留之前的思考过程。而OpenAI偷偷开了两个“外挂”:一个让模型能连接之前的推理步骤,另一个把旧信息压缩成摘要而不是直接删除。结果在官方标准下,GPT-5.6 Sol只拿到7.8分,而用上这两个“小技巧”后,分数飙升到38.3%,直接超过了对手。
这场风波对普通人意味着两件事:一是AI的能力并非铁板一块,同样的模型用不同方法调用,效果可能天差地别;二是今后看各种AI排名得多个心眼——就像跑鞋的性能可能受跑道影响,模型得分也严重依赖测试环境。对开发者和普通用户来说,未来选择AI工具时,除了看基准分数,更值得关注它在你实际使用场景中的真实表现。
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings
Matthias Bastian
View the LinkedIn Profile of Matthias Bastian
Jul 30, 2026
Nano Banana Pro prompted by THE DECODER
Ask about this article…
Search
Update – Jul 30, 2026
Added ARC Prize statements
Update:
ARC Prize co-founder François Chollet responded to OpenAI's results by distinguishing between two kinds of test setups. Harnesses "custom-made to solve the benchmark or that contain knowledge about the benchmark format" are off limits, he said. General-purpose API settings "that were not developed for ARC-AGI-3 and that are available to all API users" are fair game. In effect, Chollet is conceding that ARC Prize's own GPT-5.6 Sol score put OpenAI at a disadvantage.
He noted that ARC Prize has had "a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction," and welcomed the company "starting to figure out the answer." Different providers using different settings does create "a potential parity issue," Chollet said, but he considers that acceptable "as long as the settings and the cost are clearly reported."Ad
Original article:AdDEC_D_Incontent-1
OpenAI says it can keep up on ARC-AGI-3. After Anthropic's Claude Opus 5 quadrupled the record score on the logic benchmark, OpenAI is now showing that GPT-5.6 Sol hits 38.3 percent with two API settings, beating Opus 5's 30.2 percent.
GPT-5.6 Sol's ARC-AGI-3 scores jump dramatically when using OpenAI's custom harness with retained reasoning and compaction, compared to the official test harness. | Image: OpenAI
OpenAI isn't using the official test environment, though. Instead, it runs GPT-5.6 Sol through its own Responses API with "Retained Reasoning," which keeps the model's chain of thought between steps, and "Compaction," which summarizes old context instead of truncating it. In the official harness, GPT-5.6 Sol scored just 7.8 percent because the model's reasoning gets discarded after each action.Ad
OpenAI argues that benchmarks never measure just the model but also the technical setup around it. That's true, and ARC-AGI-3 is designed to test pure model performance. The official ARC scores use a standardized approach without provider-specific settings to ensure fair comparisons, ARC Prize said in response to OpenAI's results. The sticking point is whether ARC Prize used an older "OpenAI-style completions API" that lacked features the Claude API already offered, which would make the comparison unfair to OpenAI.
AdDEC_D_Incontent-2
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now
Source: OpenAI | via X
分享
阅读原文 ↗