Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
arXiv Machine Learning · 2026/8/4 12:30:13

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

AI 中文解读
核心亮点:这项研究用F1赛车设计和《万智牌》卡牌对战当“考场”,给AI科学家做了一次别开生面的能力测试,结果发现它们最缺的不是“想点子”,而是“挑点子”。 通俗解读:过去测试AI的科研能力,要么用虚拟题目,要么拿老问题“考古”,容易作弊。这次研究者换了个思路:让AI去预测F1 2026赛季的真实赛车设计,以及《万智牌》新卡池的职业比赛牌组。结果AI确实能提出不少像模像样的方案,比如GPT-5.2在F1里能对上40项真实创新中的10项,Gemini 3 Flash在万智牌里也选中了职业选手用的热门新卡。但整体看,AI提出的想法大多“跑偏”,真正能跟专家方案吻合的很少。 实际影响:这说明AI在科研辅助上,目前更像一个“创意喷泉”,能提供海量思路,但还缺乏专业研究员那种“沙里淘金”的判断力。未来如果AI能补上筛选和聚焦这一环,或许能真正帮科学家加速研发,比如更快找到新材料配方或药物组合。对普通人来说,这意味着AI离“独立搞科研”还有距离,但作为灵感工具已经能派上用场了。
Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in evaluating scientific reasoning and research replication, but often rely on synthetic tasks or retrospective targets, which may be confounded by prior exposure. We hypothesize that complex, adversarial, fast-moving real-world domains where expert practitioners independently generate observable outputs can provide a practical solution to fill this gap and evaluate the capabilities needed for AI scientists, including reasoning, novelty, and hypothesis formulation. We instantiate this framework in two structurally different domains, Formula 1 (F1), where models ideate around car design concepts for the 2026 season, and real pre-season innovations provide a ground truth, and Magic: The Gathering (MTG), where models propose decks from a recently updated card pool and are evaluated against 19 Pro Tour (PT) decklists. Across both domains, models produce plausible outputs, but few align with real-world expert solutions. In F1, the best model, GPT-5.2 matched 10 of 40 real innovations with 166 ideas proposed across runs. In MTG, the best deck from Gemini 3 Flash recovered 5 of 7 new-set cards from the third-place PT deck, and across all 108 decks, the cards models selected most often were also the cards most widely adopted by PT decks (Spearman $ρ= 0.74$, $p = 0.0003$). These results suggest that a key capability gap for AI scientists is not idea generation, but filtering, prioritization, and coherent novelty.
分享
阅读原文