Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
arXiv AI · 2026/8/4 16:25:38

Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition

AI 中文解读
核心亮点:科学家发现,AI的“耐心”竟能被精准操控——通过给大模型装上“时间偏好调节器”,能让它在“眼前小利”和“长远大利”之间自由切换。 通俗解读:想象一下,你问AI“该先吃蛋糕还是先减肥”,它原本可能只考虑眼前。现在研究者找到了一种方法,像调节收音机旋钮一样,在AI的“神经回路”里找到一条专门管时间感知的“线路”。轻轻一拨,AI就会变得更“短视”或更“远见”。他们用Qwen3模型做了实验,发现这个“旋钮”不仅对简单的二选一有效,还能改变AI在真实金钱决策(比如“现在拿50块还是三个月后拿100块”)中的偏好,甚至能提升它规划旅行路线的能力。 实际影响:这意味着未来AI助手将不再是“死板的建议机器”。当它帮你做理财规划、健康管理或职业选择时,能根据你的需求调整“思考视角”——如果你容易冲动,它可以更强调长期收益;如果你需要应急,它也能优先考虑眼前需求。但这也带来警示:如果这种“偏好调节”被滥用,AI可能在不知不觉中操控用户的决策倾向,因此对AI的“价值观旋钮”必须建立严格的监管机制。
We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-forced temporal-choice answers to find a short-term versus long-term direction in the model's residual stream, and evaluate contrastive activation-addition steering on a held-out binary temporal-choice task, an out-of-distribution monetary intertemporal-choice task, and a TravelPlanner capability benchmark. The central result is that temporal-horizon directions can be identified with simple contrastive linear probes and then used for steering to induce large, bidirectional preference changes. On an out-of-distribution monetary choice task that varies reward size and delay, steering strongly shifts the model's indifference threshold between smaller-sooner and larger-later rewards in both directions. We further show improvements on a planning-related capability metric under moderate temporal steering. These results suggest that model intertemporal preferences are measurable and steerable, which is relevant for AI systems that give advice involving delayed costs and benefits, and for safety questions about long-horizon planning.
分享
阅读原文