Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
arXiv Machine Learning · 2026/8/2 04:35:45

Beyond Gene Reconstruction: Learning Cell Representations through Complementary Transcriptomic Views

AI 中文解读
核心亮点:单细胞AI模型不再只靠“猜基因”训练,而是通过对比不同视角的细胞信息,让AI更懂细胞整体状态。 通俗解读:以前训练AI分析细胞,主要让它“补全”被遮住的基因数据,就像做填空题,能学会基因之间如何搭配,但不懂细胞整体是什么样的。这次科学家换了个思路,把每个细胞的基因按功能分成两组,让AI对比这两组信息,同时故意制造一些“相似但错误”的样本,逼它分辨真假,从而学会更全面的细胞特征。这就像让AI同时看一个房间的正面和侧面照片,最后它脑子里形成的不是两张图,而是一个立体房间模型。 实际影响:这项技术能帮医生更准确地区分不同细胞类型,比如识别癌细胞和正常细胞,也能更精准地推断基因调控网络,对研究疾病机制和开发靶向药物很有价值。虽然普通人暂时感受不到直接变化,但它能加速生物医学研究,未来可能推动更个性化的疾病诊断和治疗方案。AI理解细胞的能力,正从“背单词”升级到“读懂整篇文章”。
The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked expression values. This objective encourages these models to learn gene dependencies but does not directly optimize whole-cell representations, which are essential for many downstream tasks. To bridge this gap, we propose a contrastive pretraining framework that learns cell representations through complementary transcriptomic views. Since standard contrastive learning is not readily applicable to single-cell pretraining, we introduce specific adaptations along three dimensions --- co-expression-guided gene partitioning, expression-aware contrast-set construction, and competence-gated contrastive onset. Specifically, we first construct two complementary views of each cell by partitioning its genes according to their co-expression structure. Then, to prevent the model from using gene-set identity as a shortcut, we construct hard negatives by permuting expression values while keeping gene identities unchanged. Finally, we introduce a competence-aware controller to determine how the contrastive objective is applied. Experiments on cell-type annotation and gene regulatory network inference demonstrate competitive transfer under the evaluated protocols. In the six-network GRN evaluation, our method records the highest mean AUROC and AUPRC point estimates among the compared variants, while the highest-scoring variant differs across individual networks. These results establish complementary-view contrastive learning as an effective direction for single-cell pretraining beyond gene reconstruction.
分享
阅读原文