Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
arXiv AI · 2026/7/29 04:00:00

LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation

AI 中文解读
大模型也能“互相考试”了?一项新研究提出了LivingArena评估框架,让AI模型轮流出题考对方,谁能把对手难住谁就更强。这就像学霸之间互相出难题,既能发现对方的知识盲区,又能避免传统考试题被提前泄露或刷分的问题。 具体来说,LivingArena让模型扮演“考官”和“考生”:考官要绞尽脑汁想出对方答不出的问题,答对了得分,答错了扣分。为了保证问题有标准答案,还有一组强大的模型当裁判,防止考官乱出题。测试了10个前沿大模型后,这个框架生成了稳定的实力排行榜。有趣的是,模型们确实能精准抓住对方的弱点,比如在某个领域反复出题直到对方崩溃。 对普通人来说,这意味着AI评测会更真实可靠。以往我们看到的大模型排行榜可能已经过时或存在水分,而LivingArena这种动态考试能持续反映模型的最新能力。未来开发者可以更精准地发现AI的短板,用户也能更放心地选择真正可靠的AI助手。这种评估方式成本低、可扩展,就像给AI行业装上了一面永不过时的“照妖镜”。
arXiv:2607.24780v1 Announce Type: new Abstract: Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective. In this paper, our question is: \emph{Do LLMs know what other LLMs don't? And can we leverage this dynamic for evaluation?} We present \textbf{LivingArena}, an automated, contamination-resistant evaluation framework. In this framework, models take turns proposing questions, aiming to pose items that opponents cannot answer correctly. Questioners are encouraged to actively identify and exploit opponents' knowledge boundaries, receiving rewards when the answerer fails, while the answerer is rewarded otherwise. To ensure questions contain objectively verifiable answers, a judge panel of strong models validates them, penalizing questioners if validation fails. Evaluating ten frontier LLMs, LivingArena yields a stable Elo leaderboard. Our behavioral analyses show that models identify and exploit their peers' cognitive boundaries: self-play and tournament logs indicate that they localize and double down on opponents' weak dimensions. Beyond static knowledge recall, peer probing measures factual rigor and the higher-order ability to probe an opponent's weaknesses, correlating only weakly with human preference and offering a scalable, low-cost approach to continuous evaluation.
分享
阅读原文