Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
Carnegie Mellon AI · 2026/7/30 16:00:00
Healthcare Blind Spots: AI Models Prone To Fabricating Diagnoses

Healthcare Blind Spots: AI Models Prone To Fabricating Diagnoses

AI 中文解读
核心亮点:AI聊天机器人竟会“凭空编造”医疗诊断,在缺少关键图片时,有近两成概率根据患者年龄、性别和种族瞎猜病情。 通俗解读:卡内基梅隆大学一位学生做了个实验:他故意不给AI提供患者照片或X光片,只告诉AI患者的年龄、性别和种族,然后问“这个人皮肤上的痣怎么了?”或者“这张胸片有什么问题?”结果发现,像Claude、GPT-5这样的顶级AI,竟有18%的几率会“脑补”出一个诊断。比如,对65岁的白人男性,AI几乎每次都说是黑色素瘤;对年轻的黑人患者,AI则倾向于诊断为结节病。而实际上这两种病都非常罕见。说白了,AI并没有它看起来那么聪明,它会在信息不全时自己“编故事”,而不是承认自己不知道。 实际影响:如果你将来用AI咨询健康问题,或者医院系统里用了AI辅助诊断,就要格外小心了。AI可能因为你填的年龄、性别、种族不同,就给出截然不同甚至完全错误的答案。这项研究提醒医疗行业:在真正把AI用进医院之前,必须严格测试它会不会对不同人群产生偏见。对我们普通人来说,千万别把AI的医疗建议当成权威结论——它连“没图”都不告诉你,直接瞎猜,这可不靠谱。
CMU Study Shows LLMs May Invent Information When Missing Key Data 07/20/2026    Amanda Sapio The Breakdown In 18% of cases, Claude, GPT-5 and Gemini fabricated medical diagnoses, despite supporting images being omitted. The AI models used demographic-based clinical assumptions to invent these diagnoses. CMU research emphasizes the need for demographic sensitivity testing and verification before deploying AI tools in medical systems. * * * Siddharth Vohra is in the Master of Science in Computer Vision program at the Robotics Institute A Carnegie Mellon University School of Computer Science student recently showed that large language models (LLMs) intermittently invent false information when responding to medical questions.  For his study, Robotics Institute master’s student Siddharth Vohra asked popular LLMs to describe a medical image that was intentionally omitted from the query. Rather than requesting the missing image, the models fabricated a diagnosis based on the user’s age, gender and race 18% of the time.  “A 65-year-old white man asking Claude about a skin mole receives melanoma in nearly every response,” said Vohra, the study’s sole author. “When chest X-ray questions are presented, OpenAI’s GPT-5 names sarcoidosis for roughly 77% of young Black patients.” In reality, fewer than one in 10,000 moles will become melanoma, according to the Memorial Sloan Kettering Cancer Center. The prevalence of sarcoidosis varies throughout the world, but the Cleveland Clinic reports that there are typically fewer than 200,000 cases of sarcoidosis at any given time in the U.S., making it quite rare. Vohra was raised by two physician parents who exposed him to the healthcare industry at an early age. After reading a research paper detailing how visual language systems often provide incorrect responses, he decided to investigate how demographic information influences the way models respond to medical questions.   “I analyzed close to 11,700 model responses across Claude, GPT-5 and Gemini,” Vohra said. “About 82% of the time, the models refused to provide a response because no image was attached. However, in the remaining 18%, the models invented a diagnosis instead of asking for the missing image.” When conducting the study, Vohra focused on chest X-ray, brain MRI and dermatology images in 12 simulated patient profiles. The findings reinforced a critical concern Vohra sees across the AI industry: users often assume that models understand more than they actually do. “There’s a general perception that AI models are smart because they perform well on certain tasks,” Vohra said. “But being good at one thing doesn’t mean they’re good at everything. AI is used for a host of tasks, from coding to healthcare advice, but different domains require different levels of reliability.” Because small changes in the words used to describe a patient can drastically influence a model’s conclusions, Vohra stressed that clinical pipelines should audit and test for demographic sensitivity before deploying AI models in healthcare settings. “These models are incredibly useful and they’re improving quickly,” Vohra said. “As labs continue advancing the models, stringent testing and verification checks are necessary, especially when relying on the models for medical decisions. Strong performance on one task doesn’t guarantee safety in another.”  Vohra’s current focus includes running the study at a much larger scale, with the goal of surfacing these failure patterns so the labs building these systems can identify and address them. He presented his research at the TrustVLM workshop at the ACM International Conference on Multimedia Retrieval in Amsterdam this past June. He was also recently selected for the Google Gemini Academic Program Award 2026, which promotes trustworthy AI models and mission-critical AI deployments. For more on Vohra’s work, visit his website. For More Information: Aaron Aupperlee | 412-268-9068 | aaupperlee@cmu.edu
分享
阅读原文