Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
SiliconANGLE AI · 2026/8/3 20:14:59
Alibaba debuts Qwen3.8-Max model with 2.4T parameters

Alibaba debuts Qwen3.8-Max model with 2.4T parameters

AI 中文解读
【Alibaba debuts Qwen3.8-Max model with 2.4T parameters】UPDATED 16:14 EDT / AUGUST 03 2026 AI Alibaba debuts Qw...
UPDATED 16:14 EDT / AUGUST 03 2026 AI Alibaba debuts Qwen3.8-Max model with 2.4T parameters by Maria Deutscher Alibaba Group Holding Ltd. today debuted a new addition to its Qwen series of open-source large language models. Qwen3.8-Max is the Chinese e-commerce giant’s most capable LLM to date. It features 2.4 trillion parameters, about seven times more than the Qwen3.5 model that Alibaba released in February. The LLM activates 95 billion of its parameters when answering a query. Qwen3.8-Max supports prompts with up to 1 million tokens worth of data. According to Alibaba, that enables the model to analyze more than 200 pages of text or about 100 hours of footage per request. Prompt responses contain up to 131,000 tokens. Alibaba hasn’t yet shared a detailed overview of Qwen3.8-Max’s architecture. Qwen3.5 and Qwen3.6, the previous two entries in the model lineup, shared certain technical properties including a so-called Gated DeltaNet attention mechanism. It’s possible Alibaba’s newest LLM also uses the technology. An LLM’s attention mechanism is a software module it uses to interpret prompts. The module ingests the inputted text, identifies the most important parts and prioritizes them during processing. Gated DeltaNet is a variation of the technology that was introduced by Nvidia Corp. researchers last year. Linear scale A standard attention mechanism scales quadratically, which means that doubling the size of a prompt quadruples the amount of memory and computing power needed to process it. A Gated DeltaNet module, in contrast, scales linearly. Its processor and memory requirements don’t increase as fast as prompt size grows, which reduces LLM hardware usage. Gated DeltaNet is not the only implementation of attention that offers linear scaling. What sets it apart is the use of two data processing techniques known as the gated update rule and the delta rule. According to Nvidia, these techniques make LLMs with linear-scaling attention mechanisms better at processing long prompts. Alibaba evaluated Qwen3.8-Max’s capabilities across multiple tests ahead of its debut today. According to the company, the model successfully completed a 16-day coding project without human input. In another test, Qwen3.8-Max performed a chip design optimization task that comprised over 500 steps. The model scored 1,668 points on a benchmark called Frontend Code Arena that measures LLMs’ interface development skills. That puts it just 37 points behind the most advanced configuration of Claude Opus 5, Anthropic PBC’s latest LLM. Qwen3.8-Max outperformed more than a dozen other frontier models including Meta Platforms Inc.’s flagship Muse Spark 1.1. The model is available via Alibaba’s cloud platform on launch. The company plans to open-source it next week along with a smaller, more hardware-efficient version called Qwen3.8-27B.  Image: Qwen A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities. 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network. Are you AWS customer?  Support SiliconANGLE Financially by buying your AWS services from our Marketplace portal page and links.   https://siliconangle.com/aws-marketplace/ About SiliconANGLE Media SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and
分享
阅读原文