Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
arXiv AI · 2026/8/3 16:20:17

Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce

AI 中文解读
现在买东西也能“动嘴不动手”了?这篇论文带来的“氛围商务”新概念,就是让AI代理替你完成买卖交易。但它最吸引人的地方在于,搭建了一个能全程监控AI行为的“沙盘”,让AI在替你买东西时每个动作都有记录、可查证,靠谱程度大大提升。简单说,就像你雇了个私人助理去帮你砍价下单,但这个助理身上装了记录仪,事后能回放它到底做了什么,不用担心被“坑”。研究团队还准备了近八百个真实商品的测试库,检验十个主流AI模型的“办事能力”,结果显示它们差距不小,而且研究发现,光看最终结果不够,必须盯着过程才能发现AI有没有出错。这意味着未来网购可能彻底变样:你只需说“帮我找个两千元内适合打游戏的笔记本”,AI就能自动比价、谈判、下单,全程透明可追溯。对于嫌挑选麻烦、怕被坑的消费者来说,这无疑是个福音,AI购物助理有望从“玩具”变成真正让人放心的“管家”。
In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to agents. Commerce, however, requires independently controlled Buyer and Merchant agents to interact in a shared market while preserving their private objectives and distinct authority. We introduce Agentic Commerce World (ACWorld), an environment for evaluating such agents across ongoing transactions. Through its Vibe Commerce Protocol (VCP), ACWorld validates agent actions before updating shared transaction state and records the resulting interactions, making agent behavior auditable and evaluation reproducible. The ACWorld Benchmark contains a 200-task capability-coverage track and a 60-task large-catalog track that searches 785,022 transactable listings. Across ten models, mean scores range from 65.9% to 85.6% and from 56.1% to 91.4%, respectively. Our analysis shows that process-level evidence is necessary: final state alone can miss evaluated errors, incomplete trajectories still retain useful process signals, and large-catalog tasks expose bottlenecks across stages.
分享
阅读原文