Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
Latent Space · 2026/8/1 01:38:09
![[AINews] not much happened today](https://substackcdn.com/image/fetch/$s_!1adH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHOiugSLbQAAck-3.jpg)
[AINews] not much happened today
AI 中文解读
DeepSeek低调放了个大招:新模型V4-Flash 0731一夜之间性能暴涨,几乎追上GPT-5.6,价格却压到最低。这次更新不靠堆硬件、不换架构,纯粹靠“后训练”把智商拔高了一大截,反而比大动干戈的升级更让人惊艳。
通俗点说,这就像同一台电脑没换芯片、没加内存,只升级了一下系统和算法,跑分却大幅飙升——写代码、操作电脑这类“动手能力”直接提升了近50%。尤其缓存命中后价格便宜到几乎可以忽略不计,开发者用起来几乎没有成本压力,相当于用白菜价买到了顶级AI干活。
对普通人来说,最直接的感受是:以后用AI花钱少、办事更靠谱了。无论是让它帮你写代码、处理文件还是当“自动办公助理”,完成复杂任务的成功率都会明显更高。DeepSeek沉寂一年多后重回牌桌,还刚融了700亿准备上市,加上价格优势,意味着未来市面上的AI工具会更好用、更便宜,最终受益的还是我们这些普通用户。
It might seem strange that we aren’t giving title story to a noteworthy DeepSeek open weights model update that still bumps up the Pareto Frontier that GPT 5.6 pushed out only yesterday:@teortaxesTex Temporary error with cache hit rate calculation - it was rectified a couple of minutes after your screenshot! \n\n0731 is only marginally higher Cost per Task than the earlier version, and via the DeepSeek API with ~99% cache hit discount it is most certainly on our Pareto frontier ","username":"ArtificialAnlys","name":"Artificial Analysis","profile_image_url":"https://pbs.substack.com/profile_images/2042402069320290304/A8C1lP07_normal.jpg","date":"2026-07-31T08:26:16.000Z","photos":[{"img_url":"https://pbs.substack.com/media/HOiugSLbQAAck-3.jpg","link_url":"https://t.co/ia4rJOhf1V"}],"quoted_tweet":{},"reply_count":10,"retweet_count":32,"like_count":418,"impression_count":75216,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":false}" data-component-name="Twitter2ToDOM">But because it is a post-train only update with no further details, there’s really not all that much to report, apart from noting that DeepSeek is finally relevant again after over a year of comparative obscurity (with V4 Pro this April as an exception) after becoming way too prominent, well timed after their $70B pre-IPO fundraise.AI News for 7/30/2026-7/31/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!AI Twitter RecapDeepSeek V4-Flash 0731: post-training leap, API launch, and immediate open-weights releaseDeepSeek’s biggest story of the day was the official public-beta launch of DeepSeek-V4-Flash API, with DeepSeek stating that its upgraded agent capabilities now surpass V4-Pro-Preview and that the API now supports the Responses API format and is “fully adapted for Codex” (@deepseek_ai). In a follow-up, DeepSeek clarified that the improvement applies only to the Flash API, while V4-Pro API/App/Web remain unchanged for now; V4-Pro official is still pending (@deepseek_ai). Community observers quickly highlighted the magnitude of the jump: @cline called out Terminal-Bench 82.7, up +25.8 from the April preview’s 56.9.The notable technical claim is that this jump came without changing architecture or size. Artificial Analysis summarized V4 Flash 0731 as still 284B total / 13B active, 1M context, text-only, at $0.14 / $0.28 per 1M input/output tokens with an unusually aggressive 98% cache-hit discount to $0.0028 / 1M cached tokens (@ArtificialAnlys). On their index, the model rose from 40 → 50, landing 1 point behind GPT-5.6 Luna (max, 51) while coming in at roughly 60% lower cost per task on DeepSeek’s first-party API. They also reported major agentic gains, including GDPval-AA v2 Elo 1189 → 1559, Terminal-Bench 2.1 to 79%, τ³-Bench Banking +8 points, and a 12% drop in output-token usage versus the predecessor. Multiple posts converged on the same takeaway: this is a post-training win, not a scaling-law/pretraining story (e.g. @kimmonismus, @EMostaque, @Yuchenj_UW).Open-weights followed almost immediately. The official weights landed on Hugging Face and were widely amplified by @MiaAI_lab, @_akhaliq, and others. The release is under MIT, and @vllm_project highlighted serving details: 256 routed experts, 6 active per token, 1M context, three reasoning-effort levels, and an included DSpark speculative decoding module that can be enabled via a single flag. Local/quantized deployment followed immediately: @UnslothAI published runnable quants requiring roughly 168GB RAM for lossless 4-bit and 110GB for 3-bit, while @danielhanchen later shared additional UD quants.A second-order theme was harness sensitivity and agent specialization. A number of posts argued that Flash’s gains are best understood in
分享
阅读原文 ↗