Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
arXiv AI · 2026/7/30 16:24:36

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

AI 中文解读
核心亮点:一项名为EndoCLIP的新模型,能从28万份肠镜报告中自动“找回”每张检查图像对应的描述,让AI学会像经验丰富的医生一样看懂肠道影像。 通俗解读:以前做肠镜检查,医生会写一份总的报告,但报告和具体某张图片对不上号,AI很难学会“看片”。这次研究人员像拼图一样,把海量旧报告里的文字描述和对应的病变图片重新匹配起来,凑成12万多个“图配文”训练样本,相当于给AI请了一位私人老师。结果这个AI不仅能在没见过的新数据上准确识别息肉和肿瘤,在判断良性还是恶性时,水平已经接近有经验的医生。 实际影响:这意味着以后做肠镜检查,AI可以实时辅助医生识别可疑病灶,减少漏诊。患者拿到的报告也可能更清晰,比如直接标出“哪个位置、多大、像什么性质”。更重要的是,这套方法不需要额外人工标注,能直接利用医院积累的普通报告来训练AI,未来推广到胃镜、病理等其他检查,让更多医院低成本用上AI诊断工具。
Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and six multi-centre clinical classification tasks, EndoCLIP outperforms general-purpose and biomedical vision-language encoders in both zero-shot and linear-probe settings. On benign-versus-malignant classification, its linear probe approaches the performance of expert readers in a blinded study involving 12 endoscopists. These results suggest that recovering finding-to-frame correspondence can transform routine documentation into scalable supervision, enabling clinical targets to be specified in language rather than separately annotated for each task.
分享
阅读原文