Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
The Decoder · 2026/7/26 09:43:02

Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
AI 中文解读
Anthropic的Claude Opus 5在最新智能测试中一骑绝尘,得分是此前最强模型的近四倍,展现出AI从未有过的新能力——比如自己把问题翻译成数学公式,这标志着AI的“逻辑思考”迈出了关键一步。
这个测试叫ARC-AGI-3,可以理解成一个专门考“机灵劲儿”的闯关游戏。每次遇到的都是全新关卡,人类一眼就知道怎么玩,但AI很难。它必须自己观察规则、制定策略,然后一步步执行,而不是靠死记硬背。Opus 5不仅拿下了30.2%的高分(之前最高才7.8%),还突破了五个从未被解开的关卡,其中四个甚至达到了人类水平。更令人惊讶的是,它开始像科学家一样,把难题转换成代数符号,自己推导解题公式——这种主动“抽象思考”的行为在AI中是头一回出现。
对普通人来说,这意味着AI正在从“背答案的学霸”变成“能自己解新题的选手”。未来,你可能会遇到更聪明的生活助手——比如面对一个从没见过的家电故障,它不会翻说明书,而是自己推理出修理步骤;或者在学习新技能时,它能像老师一样拆解逻辑,帮你找到完全陌生的解题思路。这样的AI将不再是“知识仓库”,而更像一个会动脑子的伙伴,在日常工作、教育甚至创意领域带来更靠谱的帮助。
Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
Matthias Bastian
View the LinkedIn Profile of Matthias Bastian
Jul 26, 2026
Nano Banana Pro prompted by THE DECODER
Key Points
Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max).
The ARC Prize team attributes the lead to genuinely stronger logical reasoning that enables more autonomous exploration and planning in unfamiliar environments.
During testing, Opus 5 displayed behavior not previously seen from an AI model, including translating tasks into algebraic notation and independently formulating reflection equations, while also solving five previously unsolved environments.
Ask about this article…
Search
The creators of the ARC-AGI benchmark say Anthropic's Claude Opus 5 owes its massive lead on ARC-AGI-3 to genuinely better reasoning.
The model scored 30.2 percent on ARC-AGI-3, making it the new leader. The previous record was 7.8 percent, set by OpenAI's GPT-5.6 Sol (Max). Opus 5 solved five previously unsolved environments, four of them at or above human level. That also puts it ahead of Anthropic's "Fable-class" models, which hit around 20 percent according to ARC Prize.
ARC Prize's analysis credits the lead to stronger logical reasoning, "which enables more autonomous exploration, planning, and execution across unfamiliar environments." During testing, Opus 5 also showed behavior that researchers hadn't seen from a model before. It translated tasks into algebraic notation and independently formulated reflection equations for the first time.Ad
Claude Opus 5 leads the ARC-AGI-3 leaderboard by a wide margin, far ahead of GPT-5.6 Sol (Max) and other competitors. | Image: ARC Prize
Six of the 25 public demo environments have now been solved. The full results, replays, and benchmarking code are publicly available. On the older ARC-AGI-2 benchmark, Opus 5 scores 90.4 percent, and it reaches 97.5 percent on ARC-AGI-1. Both results match previous top scores, though at slightly higher costs, according to ARC Prize.AdDEC_D_Incontent-1
ARC-AGI-3 measures how well AI models solve new tasks they didn't encounter during training, including ones humans can usually handle with ease. The current version works like a game. The model must infer the rules of an interactive environment, plan its actions, and carry them out step by step. This tests general reasoning rather than stored knowledge.
Some AI systems may have already passed the benchmark, but they rely on extra software known as a harness. Official scores count only the language model's own performance. ARC Prize argues that future AGI systems shouldn't need outside help to solve new tasks. Opus 5 would likely score even higher if used within Claude Code.Ad
Independent tests suggest narrower gains
Anthropic hasn't explained the gain, but targeted data labeling and reinforcement learning are plausible factors. Unlike earlier models, Opus 5 was developed after ARC-AGI-3 and its format became public. That may have let Anthropic target the benchmark's skills and puzzle formats, though it doesn't show the company trained on the exact tasks. Annotators could have labeled reasoning traces, useful actions, failed attempts, and recovery steps from similar puzzles. Reinforcement learning could then reward exploration, planning, rule discovery, and self-correction.
Tests on Witness, Guanghan Ning's private benchmark for interactive puzzle games, point to narrower gains. Opus 5 scored 43.4, statistically tying Kimi K3 and Fable 5 while improving far less over Opus 4.8 than it did on ARC-AGI-3. On a puzzle
分享
阅读原文 ↗