Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
arXiv AI · 2026/8/2 05:57:46
Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking
AI 中文解读
核心亮点:人类写的“错误答案”也能用来测试AI的“幻觉”问题了,而且效果不输AI自己生成的,这让AI评测变得更客观、更持久。
通俗解读:AI经常“一本正经地胡说八道”,比如看着一张猫的图片,却硬说成狗,这叫“幻觉”。以前测试AI有没有这个毛病,得靠AI自己生成一些胡说八道的例子,但不同AI说的瞎话不一样,标准很难统一。现在研究人员干脆让真人来写这些错误描述——用了1600条真人写的样本,涵盖中英法意四种语言,还对比了五款主流AI生成的18400条样本。结果发现,真人写的错误描述和AI自己犯的错在类型上非常接近,而且真人写的更规范、更容易控制,用来当“考卷”反而更靠谱。
实际影响:以后AI公司发布新模型时,我们不用再担心它是不是只会在自家考题上表演“标准答案”了。有了这种真人编写的测评集,AI的“胡说八道”会更容易被发现,你用AI查资料、认图片时,看到的结果会更可信。尤其对需要多语言支持的用户来说,中文、英文、法文、意大利文都能测,覆盖面更广,用起来更放心。
In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make benchmarking detection independent of particular models. To this end, we construct a dataset of 1,600 human-written samples, spanning four languages (Chinese, English, French, Italian), and 18,400 samples from five vision-and-language models, all annotated for hallucinations using a fine-grained span-level labeling scheme. We find that human-written samples result in higher agreement and allow greater control of dataset contents, while remaining distributionally similar to samples derived from vision-and-language samples and providing a reasonable portrayal of detection capabilities - suggesting that human data is a viable substitute for model-based hallucination benchmarks.
分享
阅读原文 ↗