Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
Nature Machine Intelligence · 2026/7/24 00:00:00
Thinking and rethinking data AI readiness
AI 中文解读
数据质量决定了AI模型的成败,但现在的挑战是:不仅要让数据“规范”,还要让数据“懂AI”。最近《自然·机器智能》的一篇社论指出,过去生物学家们遵循FAIR原则(数据要容易找到、能打开、能互相配合、能重复用)已经不够用了。因为AI发展太快,数据集本身会随着时间变化,比如新的实验数据加入、标注标准调整,可模型还是用旧数据训练出来的,两者就会“对不上号”。这样一来,研究速度再快,结果也可能不靠谱。为此,美国国立卫生研究院在2022年启动了Bridge2AI项目,专门把生物医学数据改造成适合AI直接使用的格式。这对普通人的生活意味着什么?未来AI辅助诊断、药物研发会更可靠——因为底层数据是“活”的,会跟着医学知识一起更新,而不是一批数据用到底。简单说,AI不是学会一本“死书”,而是能读懂一本“随时修订的活教材”。
Download PDF
Editorial
Published: 24 July 2026
Thinking and rethinking data AI readiness
Nature Machine Intelligence
volume 8, page 1013 (2026) Cite this article
Save article
View saved research
2045 Accesses
2 Altmetric
Metrics details
Training machine learning models on high-quality biological datasets can quickly produce abundant results. But as datasets often evolve over time, further work is required to update models and maintain alignment between models and datasets.
Artificial intelligence (AI) applications in biological sciences thrive on high-quality datasets, and on sustained community efforts in collecting, curating, annotating, formatting and other data management tasks. The FAIR guidelines were developed to encourage good practices in ensuring that datasets are open, findable, interoperable and reusable (M. Wilkinson Sci. Data 3, 160018; 2016). However, with the fast development of AI tools, FAIR alone is no longer sufficient, and in recent years there has been a strong focus on the ‘AI readiness’ of datasets, which requires going beyond the FAIR principles. For example, Bridge2AI was launched by the National Institutes of Health (NIH) in 2022 with a focus on preparing biomedical datasets for the widespread adoption of AI tools. The Open Data Institute, a UK-based non-profit organization, recently published a framework that outlines requirements for dataset AI readiness in four categories: dataset properties, metadata, surrounding infrastructure and governance. Furthermore, ‘AI-ready’ extensions to FAIR have been proposed, such as FAIR-R, which focuses on whether annotations, provenance, responsible use and documentation are sufficiently developed for machine learning.Existing efforts primarily describe the characteristics of AI-ready datasets but it is an underexplored question how AI-readiness should be sustained as biological knowledge changes over time. In a Comment in this issue, Amarda Shehu and Ruth Nussinov highlight this challenge. They point out that AI readiness cannot be an intrinsic property of biological datasets, as datasets are continuously evolving in response to discoveries, updated classifications, new acquisition strategies with higher capacity, and changes in clinica
分享
阅读原文 ↗