Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
arXiv AI · 2026/8/2 04:15:12

Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable

AI 中文解读
能否用可信的方法证明AI真的变强了?这篇论文给审计AI科学发现提供了一把“双刃剑”,一边判定负面结果是铁的数学事实,另一边避免了自卖自夸。 简单说,过去AI说自己发现了新东西,靠的是分数提升或统计显著性,但这些可能只是“刷题刷出来的”,换套规则漏洞百出。研究者设计了一套审计方式:先用数学上不可能出错的标准戳穿假进步,再让外部裁判独立验证真成果。结果发现,不少AI在自选裁判下表现惊艳,换裁判就露馅;而少数真正有效的方法,确实能用更少的算力完成更多任务。这就像给学生考试,以前老师出题又自己判卷,现在换成外部评委,还能提前证明某些答案必然错误——这样才算实打实地进步。 对普通人而言,这项研究让AI做科研、写代码或诊断病症时,结论更可靠。以后医学AI说发现新药,或教育AI说能提分,我们不用全凭“系统自证”,而是有更硬核的第三方核查机制。虽然眼下离落地还有距离,但它为“AI自我吹嘘”时代加上了一道安全阀,防止我们被夸大宣传蒙蔽。
When a self-improving AI-for-science system claims a new capability, the evidence is usually a benchmark delta, a description-length gate, or a p-value. None separates a real gain from extra search, from a changed verifier, or from adaptation to a fallible oracle. We build a two-sided audit whose negative side is a formal fact: a pseudoknot-free oracle provably cannot represent a crossing base pair, so the prior verifier's range is bounded exactly, offline, before any run. "New" is relative to the agent's prior self, never to the base model. First, how far a single fallible oracle can inflate a capability claim. An invented, solver-free operator solves 43/60 crossing RNA targets under the predictor it optimizes, above a context-free floor of 0/60; under three predictors, 1/60 survives. Paired on the same 43 targets, a predictor the operator never saw confirms 2 of its designs against 26 for a minimum-free-energy solver (p = 8e-7). No statistic computed from the system and its own oracle sees that gap. Second, agent-written procedures can beat a human-written one under a judge no objective can flatter, at a fraction of the compute. Of six frontier models, the two whose operators ran without timeouts carry over at 0.293 against our 0.095 (n = 951 paired units, target-clustered [+0.108, +0.297], p = 5e-5) while spending 4.6-10x fewer oracle calls. Three rungs: difference under an outside adjudicator (reached), not bought with compute (reached, both directions), mechanism identified and transferable (not reached; seven candidates tested, none moves the statistic). The ceiling is the panel itself: its three predictors share nearest-neighbour thermodynamic parameters, two agreeing at kappa = 0.673. The audit is as unsparing about our own system: matched undirected search is an exact zero, and a search-free probe puts 84% of our headline effect on targets a random sequence already solves.
分享
阅读原文