Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
arXiv AI · 2026/7/31 04:00:00

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

AI 中文解读
核心亮点:AI办公终于有了“性价比”评分表——不光比谁能干,还比谁干得值。 通俗解读:现在AI帮人干办公活儿已不稀奇,但之前没人算过这笔账:AI干活到底比人省多少?质量又差多少?这个名为OmegaUse-OfficeVal的测试系统,专门给AI出了一百道真实办公难题,每道题都标了人类做这事要花多久、市场价多少。结果几家顶尖AI表现不错:成本几乎可以忽略不计,速度也飞快,但交出来的“作业”质量还明显不如熟练员工。简单说,AI现在像个手脚麻利但手艺生疏的实习生——便宜够快,但活儿还不够精细。 实际影响:如果你常处理表格、写文档这类繁琐工作,这消息算是分水岭。它说明AI当帮手已经具备经济价值,用来初筛资料、起草内容完全够用,老板们以后算账时能直观看到“AI干十分钟成本几分钱,人力干两小时要花几百”。但指望它交出的成品直接能用,还为时过早。这项评测最实在的意义,是让企业和普通用户都看清该把AI用在哪儿,别盲目期待,也别轻易错过。
arXiv:2607.27155v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, and more information is available on our project website: https://omegause-officeval.github.io.
分享
阅读原文