Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
The Decoder · 2026/7/27 12:28:06
METR introduces a new metric to calculate exactly when AI agents become more expensive than humans

METR introduces a new metric to calculate exactly when AI agents become more expensive than humans

AI 中文解读
AI到底什么时候比人类更划算?研究机构METR提出了一个叫“支出地平线”的新指标,用美元来衡量AI代理在解决问题上的性价比。简单说,就是把AI跑任务花的钱和人类劳动花的钱放到同一个秤上比:当预算低于某个点时,AI更便宜;一旦超过,人类反而更划算。他们拿一个公开的AI训练比赛做测试,发现人类每让模型提速1%,平均要花16小时和约2500美元的成本。而AI一开始能快速搞定简单活,但任务变复杂后就越来越吃力。这个指标对普通人和企业的影响很直接:以后评估用AI还是雇人,不再靠感觉,而是有了清晰的成本线。比如在写代码、做翻译等重复性工作上,小预算用AI更省钱;但遇到需要深度创意或复杂决策的长期项目,人类依然有价格优势。随着新一代AI模型崛起,这个临界点可能还会变,但至少现在,我们有了一个靠谱的计算工具。
METR introduces a new metric to calculate exactly when AI agents become more expensive than humans Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Jul 27, 2026 Nano Banana Pro prompted by THE DECODER METR's new metric, the "expenditure horizon," puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming, the metric has blind spots, and the newest generation of models could change the picture. One of the biggest questions in AI research is whether AI can accelerate its own development and keep getting better at an increasing pace. That's been hard to measure because it requires comparing very different kinds of costs: human labor, compute for experiments, and the cost of running the AI itself. Research organization METR proposes a new metric to tackle this: the "expenditure horizon." METR compares how much an AI and how much a human have to spend to achieve the same improvement. The expenditure horizon is the point where both cost the same. Below that budget, the AI is the better deal. Above it, the human works cheaper. The green line represents returns from human labor, the red curve represents an AI agent's returns. Where they cross is the expenditure horizon. | Image: METR The idea builds on a pattern METR has seen in previous tests: AI agents often solve simple, low-cost tasks faster than humans. But as budgets grow and tasks get harder, they fall behind. Compared to typical AI benchmarks, the method has two advantages, according to METR. First, it doesn't just give a pass-or-fail verdict. Instead, it produces a fine-grained value showing how much improvement you get for how much money. Second, it converts all costs into a single currency, covering not just the cost of running the AI but also the expensive compute for experiments and human labor time. Humans spend about $2,500 for each one-percent speedup METR chose the NanoGPT speedrun as its testing ground. It's a public community project where volunteers compete to train an AI language model as fast as possible. The task stays the same; only the training approach can change. Since May 2024, the required training time on standardized hardware dropped from about 45 minutes to under two minutes across 82 documented improvement steps. The NanoGPT speedrun timeline: 82 improvement steps delivered a cumulative 33x speedup. | Image: METR To figure out the cost of human work, METR interviewed two of the project's most active contributors and also had an AI model (Opus-4.6) estimate the effort behind each improvement. Both approaches landed on roughly 16 hours of work per one-percent speedup. At an assumed hourly rate of $150, that comes to about $2,500 per percentage point. METR stresses that this number is very uncertain. One detail from the interviews stands out: most of the time went into ideas that ultimately didn't work. AI agents have only made small contributions so far For the comparison, METR had six AI models work on the same task independently. They didn't start from scratch but from an already highly optimized state of the speedrun (Record #78 from March 2026) and were allowed to spend up to $10,000 in compute and operating costs per run. The result: estimated expenditure horizons between $0 and $3,300. AI model curves compared to human labor (intersection at roughly $2,500 per percent). Only GPT-5.5 and Opus-4.8 reach meaningful expenditure horizons. | Image: METR The differences between models were stark. GPT-5 and Opus-4.1 produced no real progress after careful verification. Their apparent gains turned out to be random noise. GPT-5.5 and Opus-4.8, on the other hand, delivered real improvements of about 1 and 1.5 percent, respectively. The quality of AI-generated ideas was mixed. Th
分享
阅读原文