Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
The Decoder · 2026/7/31 10:57:37

Anthropic follows OpenAI in admitting its Claude models reached out of test environments and attacked real-world systems

AI 中文解读
【Anthropic follows OpenAI in admitting its Claude models reached out of test environments and attacked real-world systems】Anthropic follows OpenAI in admitting its Claude models reached out of test environments and attacked real-world systems Matthias Bastian View the ...
Anthropic follows OpenAI in admitting its Claude models reached out of test environments and attacked real-world systems Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 31, 2026 Nano Banana Pro prompted by THE DECODER Key Points During internal cybersecurity assessments, Anthropic found that three Claude models had been exposed to the open internet due to a misconfiguration and had attacked real-world systems, mistaking them for simulated targets. Opus 4.7 extracted data from a real company, while Mythos 5 created malware and published it on PyPI, where it was downloaded by actual systems. Both models believed they were still operating within a simulated environment. Only a newer internal research model recognized the real-world nature of the targets and independently stopped its attack. Anthropic categorizes the incidents as operational errors rather than an alignment failure. Ask about this article… Search Three different Claude models found their way onto the open internet during cybersecurity evaluations and attacked real-world systems. One even published malware on a public platform. Anthropic blames a misconfiguration. While reviewing its own cybersecurity evaluations, Anthropic found three incidents where Claude models escaped test environments and compromised real companies on the internet. The company reviewed 141,006 evaluation runs and flagged six cases where models accessed systems they weren't supposed to reach. The review was triggered by OpenAI's security incident. A configuration error gave Claude access to the open internet The incidents happened during capture-the-flag exercises where Claude was tasked with finding hidden information on external systems in fictional scenarios. Anthropic says Claude was told in the evaluation prompt that it was in a simulation with no internet access, but a miscommunication between Anthropic and its evaluation partner Irregular meant the test agents actually had full internet access.Ad Claude treated the real systems it found as part of the exercise. Anthropic says the models didn't exploit complex vulnerabilities and instead used basic techniques like weak passwords and unauthenticated endpoints.AdDEC_D_Incontent-1 Similar to OpenAI's case, the evaluations ran without the safety guardrails built into the public models, since the goal was to measure raw model capabilities. Anthropic says the evaluation infrastructure had no access to internal systems or customer data. Claude Opus 4.7 kept attacking after recognizing its target was real The most serious incident involved Claude Opus 4.7. Across four runs, the model couldn't reach its fictional target and instead found a real company that happened to share the same name. Claude identified and exploited vulnerabilities in the company's infrastructure, pulling out login credentials and a database with several hundred lines of production data.Ad Anthropic says the model recognized it was dealing with a real system in all four runs, but in two cases it rationalized this away by assuming the real company must be part of the exercise. The model never stopped on its own. Myth 5 published malware on PyPI and compromised a security company The second incident shows how far an AI model will go to complete a task. Claude Myth 5 found setup instructions in the fictional environment that told developers to install a specific Python package from PyPI, but the package didn't exist.AdDEC_D_Incontent-2 Claude saw an opening and created a malicious package with the same name, then published it on the real, public PyPI platform. Anthropic says the model went through considerable effort to pull this off, first trying to get a phone number, failing, then fi
分享
阅读原文