Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
Dev.to AI · 2026/8/3 14:27:05

Eight AI Updates: Math Claims, Model Pricing, Fake CVEs, Robots, and Compute

AI 中文解读
AI圈又迎来大动作日!最新消息显示,OpenAI自家AI在数学领域实现了突破,一口气解决了十个困扰学界多年的难题,而且AI还学会了自己“整理”研究成果,相当于配了个智能研究助理。不过这些成果还在等专家认证,暂时不能盖棺定论。另一边,阿里巴巴发布了新一代大模型,参数规模高达2.4万亿,却暂时只让看不让摸,要等下周才开放,像极了吊人胃口的预告片。还有个好消息是,有第三方测试显示,某新模型推理成本低到离谱,每百万字输入只要一毛多钱,以后用AI干活可能比请人喝杯咖啡还便宜。不过,网络安全圈也拉响警报:有研究机构发现,超过50条漏洞报告疑似是AI“编故事”造出来的假情报,部分还混进了官方数据库,这提醒我们AI再聪明,也得有人把关。这一波眼看AI要攻下学术、价格战、网络安全几个山头,普通人能拿到的AI工具会越来越强,但学会辨别真假AI信息,也成了新技能。
<p>AI news on August 3 ranged from ambitious mathematics claims to a warning about fabricated vulnerability reports entering formal data systems. Here are eight developments, with the source boundary attached to each one.</p> <h3> 1. OpenAI published ten mathematics and theoretical computer science advances </h3> <p>OpenAI says an internal version called Astra solved or materially advanced ten long-standing open problems. According to the company, the model generated mathematical arguments, humans helped organize them into papers, and the model then formalized the work into Lean certificates. OpenAI estimated that the tokens used to find the solutions would cost about $2,000 at Sol API prices.</p> <p>This remains an OpenAI publication, not a settled academic verdict. Correctness, novelty, and the status of each result still require review by the relevant research communities.</p> <p>Source: <a href="https://openai.com/index/ten-advances-in-mathematics/" rel="noopener noreferrer">https://openai.com/index/ten-advances-in-mathematics/</a></p> <h3> 2. Qwen3.8-Max was unveiled, but its open weights are not available yet </h3> <p>Alibaba announced Qwen3.8-Max with 2.4 trillion total parameters and 95 billion active parameters. The published specification includes text, image, and video input plus a one-million-token context window. The weights are expected next week.</p> <p>Those specifications and performance claims are vendor statements. The release of the weights has not happened yet, and benchmark or long-running task results should not be described as independently verified.</p> <p>Source: <a href="https://qwen.ai/blog?id=qwen3.8" rel="noopener noreferrer">https://qwen.ai/blog?id=qwen3.8</a></p> <h3> 3. A third-party test put DeepSeek V4-Flash at a very low token price </h3> <p>Reuters cited Artificial Analysis measurements that priced V4-Flash at about $0.14 per million input tokens and $0.28 per million output tokens. The reported average cost was about $0.03 per benchmark task.</p> <p>This is a third-party benchmark and pricing snapshot. Token prices alone do not establish the total cost, reliability, or value of a model on a real workload.</p> <p>Source: Reuters, August 3, 2026, reporting Artificial Analysis measurements. The source brief did not include the article URL.</p> <h3> 4. JFrog says 54 reviewed advisories appeared to be LLM fabrications </h3> <p>JFrog investigated a new GitHub repository that published more than 50 vulnerability advisories. The researchers concluded that 54 advisories appeared to be LLM-generated fabrications. The reports cited missing functions, unrelated code lines, or proof-of-concept inputs that failed before reaching the alleged vulnerable code. Several reports still entered downstream systems and received high or critical severity metadata.</p> <p>The finding applies to this batch. It does not show that the entire CVE system is untrustworthy. JFrog's strongest evidence came from checking source code and PoCs, not from relying on an AI detector.</p> <p>Source: <a href="https://research.jfrog.com/post/sqlite-critical-cves-or-llm-slops/" rel="noopener noreferrer">https://research.jfrog.com/post/sqlite-critical-cves-or-llm-slops/</a></p> <h3> 5. Google released Gemini Robotics ER 2 for high-level robot reasoning </h3> <p>Gemini Robotics ER 2 is available through the Gemini API and Google AI Studio. Google says it supports continuous video progress tracking, self-correction, tool use, and coordination among multiple robots. Its role is high-level reasoning rather than direct motor control.</p> <p>Google's enterprise platform remains in private preview. Public demonstrations and evaluations are company-led, and real performance depends on the robot hardware and control stack.</p> <p>Source: <a href="https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/" rel="noopener noreferrer">https://blog.goog
分享
阅读原文