Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
VentureBeat ML · 2026/8/3 23:50:58
Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use

Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use

AI 中文解读
阿里这次放了个大招!昨晚发布的Qwen3.8-Max号称在“AI自己操作电脑干活”这项能力上,直接超过了GPT-5.6 Sol Max和Fable 5这两个顶尖模型。要知道,这可是中国团队第一次在Agentic Computing这个最前沿的赛道上跑到了最前面。 简单来说,以前的AI更多是聊天工具,你问一句它答一句;而Qwen3.8-Max更像一个“数字员工”,你给它一个长期任务,比如“把这个软件项目做完”,它能自己规划步骤、写代码、查资料、改错,甚至连续干十几天不用人管。这背后靠的是2.4万亿参数的混合专家模型,相当于把一群各有所长的“小AI专家”组合成一个超强团队。 对普通人来说,最直接的影响是未来AI助手可能从“回答问题”升级成“替你完成工作”——比如自动整理一个月的数据报表,或者按你的要求设计个简单网页。而对企业和开发者来说,最大的悬念在于阿里承诺下周就会开放模型权重。如果采用宽松许可,就意味着企业可以把这头“AI猛兽”装进自己的服务器,不用再受制于云端API费用。不过目前许可协议还没公布,万一跟之前月之暗面那样搞个限制性条款,开源的意义就打了折扣。
Chinese e-commerce and cloud giant Alibaba's famed Qwen team of AI researchers last night unveiled Qwen3.8-Max, a new flagship 2.4-trillion-parameter mixture-of-experts (MoE) multimodal large language model (LLM) that targets one of the most competitive corners of the frontier AI market: autonomous software engineering and long-horizon enterprise work. If the company's published benchmarks hold up under broader independent testing, Qwen3.8-Max doesn't merely compete with today's leading proprietary models — it surpasses several of them on some key benchmarks in agentic computing.Most notably, Qwen reports that Qwen3.8-Max scores 86.1 on the OSWorld-Verified benchmark measuring how well ahead of GPT-5.6 Sol Max (83.2) and Fable 5 (85.0), while also posting the highest reported score on PaperBench and leading or remaining highly competitive across software engineering, research reproduction, multimodal reasoning, and visual web development benchmarks.The release also signals a potentially significant strategic shift for Alibaba: the company says open weights for Qwen3.8-Max will be released next week, alongside Qwen3.8-27B. If that happens under a permissive license, it would represent the first time a Max-class Qwen model becomes available for self-hosted deployment—a move that could substantially reshape enterprise adoption. One important caveat remains, however: Alibaba has not yet disclosed the licensing terms, leaving open the possibility that the release could use a more restrictive custom license, as we saw recently with Chinese rival Moonshot's open Kimi K3 frontier model, rather than a broadly permissive one such as Apache 2.0.A different definition of 'frontier'Over the past year, the competitive landscape for foundation models has become increasingly specialized.OpenAI has largely focused its GPT series on general reasoning, multimodal interaction and enterprise productivity.Anthropic's Claude series has emphasized coding and dependable long-context reasoning. Google continues to push Gemini toward multimodal productivity and web-native workflows. Moonshot AI's Kimi K3 recently entered the conversation by pairing frontier-class performance with an open-weight release.Qwen3.8-Max attempts to combine many of these strengths into a single model aimed squarely at enterprise automation.Rather than emphasizing conversational intelligence, Alibaba is positioning the model as an autonomous coworker capable of executing projects that span days rather than minutes. According to the company, Qwen3.8-Max can autonomously complete software projects lasting more than 10 days, reproduce research papers involving thousands of lines of code, perform iterative chip-design optimization, and continuously revise plans using multimodal feedback loops.Those demonstrations remain company-produced and have not yet been broadly replicated by independent evaluators. Nevertheless, they illustrate a growing industry trend: frontier models are increasingly competing on their ability to finish entire workflows rather than answer individual prompts.Benchmarks increasingly reward autonomous executionThe benchmark suite released alongside Qwen3.8-Max reflects this shift.Instead of focusing solely on traditional reasoning exams or coding puzzles, many of the highlighted evaluations measure long-horizon execution.On OSWorld-Verified, which evaluates computer-use agents interacting with desktop environments, Qwen3.8-Max posts 86.1, ahead of GPT-5.6 Sol Max's 83.2, Fable 5's 85.0, and Gemini 3.1 Pro's 76.2.The model also leads:PaperBench: 93.0TerminalBench 2.1: 86.6Vision2Web: 69.0LVBench: 81.8ERQA: 77.8Elsewhere, it remains competitive with proprietary leaders while trailing in several categories. On the professional software engineering benchmark SWE-Pro, for example, OpenAI's model posts the highest reported score, while Opus 4.8 continues to lead on certain software engineering evaluati
分享
阅读原文