Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
arXiv Machine Learning · 2026/8/3 16:20:11
Intention Inference Under Execution Noise: Separating Aleatoric and Epistemic Uncertainty in Social Dilemmas
AI 中文解读
科研人员最近破解了一个AI博弈中的“读心术”难题。在现实合作中,对方的背叛可能是故意,也可能只是手滑失误。传统AI模型把所有行动都当成真实意图,导致“宁可错杀一千”的过度报复。这项新研究让AI学会区分“故意使坏”和“操作失误”,就像在玩剪刀石头布时,不仅看你出了什么,还猜你本来想出什么。研究用“反复囚徒困境”游戏做实验,发现带“读心能力”的AI在面对善意但偶尔犯错的对手时,能显著提升合作成功率,避免恶性循环。不过有趣的是,如果双方都过于纠结猜测对方心思,在噪音太大时反而会因猜疑一起崩盘,说明“读心”并非万能。这项技术未来可应用于自动驾驶(判断旁车是恶意别车还是技术差)、智能客服(识别用户是真生气还是表达歧义)等场景,让机器在复杂人际互动中少一点误判,多一些通情达理。
In noisy social dilemmas, intended actions are stochastically corrupted before execution, so an observed defection may reflect hostile intent or action error. Standard Markov Decision Process (MDP) formulations treat executed actions as states, structurally precluding this distinction and causing systematic over-retaliation. We introduce a Partially Observable MDP (POMDP) formulation encoding opponent intentions as latent states and executed actions as noisy observations, solved within the active inference (AIF) framework with a cost function that decomposes into epistemic and pragmatic components that jointly address inferring current intent and learning how intent evolves. In the Iterated Prisoner's Dilemma with symmetric noise, we derive a critical noise threshold governing cooperation collapse, connecting it to a fixed-point condition on learned priors. Experiments reveal that the value of intention inference is context-dependent: the POMDP provides consistent advantages against conditionally cooperative opponents, but mutual intention inference under sufficient noise produces correlated belief-driven collapse. The advantage is specific to games where intent attribution is decision-relevant.
分享
阅读原文 ↗