Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
arXiv Machine Learning · 2026/7/31 14:22:33
Analytical and Bootstrap Confidence Intervals of Double Machine Learning: Simulation studies and an application to rural-urban difference in obesity prevalence
AI 中文解读
双机器学习置信区间研究:算法选择是关键!这项研究发现,用不同AI算法来治疗效应,结果差异很大。更意外的是,数据量越大,统计结果的可靠性反而下降,颠覆了传统认知。研究还用美国县级数据证实,越偏远的地区肥胖率越高。
通俗说,科学家用AI评估某项政策或干预措施的效果时,需要让AI辅助排除干扰因素。过去大家默认只要数据够多,结果就可靠,但新研究发现,AI采用的算法不同,结论的“置信度”差异很大,甚至数据越多,误差反而增加。这就像问不同师傅估分,数据越多他们答案反而越分散。
这项研究提醒我们,用AI做医疗、经济决策时,不能只看结果,还要看它用的是什么算法。未来AI分析报告可能会附上“算法可靠性说明”,帮我们判断该不该信。同时,研究揭示乡村肥胖问题更严重,也提醒政策制定者多关注偏远地区的健康资源投入。
Double Machine Learning (DML) is a popular approach for treatment effect estimation in various settings, which allows a wide range of flexible machine learning methods to be used for nuisance parameter estimation while preserving valid inference. In practice, however, applied researchers must choose among many machine learning algorithms for nuisance models, and the impact of this choice on the variance estimation of DML is not well characterized. We conduct a comprehensive simulation study to compare the coverage probability of DML confidence intervals across different machine learning algorithms. In this study, we compare (1) analytical confidence intervals derived by DML theory versus (2) bootstrap confidence interval. We use a set of learners including ordinary least squares, LASSO, Random Forest, LightGBM, and Neural Networks under different data generation settings. We evaluate the performance across difference settings by bias, confidence interval width, and most importantly, coverage probability. Our results show substantial variability in coverage performance across analytical and bootstrap confidence intervals, highlighting that learner choice plays a critical role in reliable DML inference. Surprisingly, we find that in many settings, when sample size increases, the coverage probability of both DML analytical and bootstrap confidence interval decreases. We further investigate coverage probabilities using a real dataset on rural urban differences among U.S. counties. The real data analysis discovers that (1) the model performance still varies by the learner choices and (2) greater rurality has a statistically significant increasing effect on county level obesity prevalence.
分享
阅读原文 ↗