面向基层医疗场景的机器学习模型在冠心病二级预防中的风险分层研究
DOI:
作者:
作者单位:

(新疆医科大学第五附属医院全科,新疆乌鲁木齐市830000)

作者简介:

林琪雄,硕士研究生,研究方向为冠心病的基础与临床,E-mail:18819797625@163.com。通信作者史树银,主任医师,硕士研究生导师,研究方向为心血管内科、全科医学科、医院管理,E-mail:shy@xjmu.edu.com。

通讯作者:

基金项目:

国家自然科学基金项目(81960073);新疆维吾尔自治区自然科学基金项目(2023D01C56);“天山英才”医药卫生高层次人才培养计划(TSYC202301B015)


Research on risk stratification of machine learning models in secondary prevention of coronary heart disease for primary care settings
Author:
Affiliation:

Department of General Practice, the Fifth Affiliated Hospital of Xinjiang Medical University, Urumqi, Xinjiang 830000, China)

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
    摘要:

    目的]开发并验证一种基于易获取临床指标的轻量化机器学习模型,以实现冠心病(CHD)的精准风险预测。 [方法]回顾性纳入1 376例经冠状动脉造影检查的患者数据。通过单因素Logistic回归(LR)、Boruta算法和LASSO回归筛选出13个核心预测变量。利用训练集(n=963)构建了包括Logistic回归、支持向量机(SVM)、随机森林(RF)等7种机器学习模型,并在测试集(n=413)上通过受试者工作特征曲线下面积(AUC)、召回率、精确率、F1值及Brier评分等指标评估其性能。采用SHAP方法对最优模型进行可解释性分析。 [结果]单因素Logistic回归显示,男性(OR=2.445,95%CI:1.870~3.202)、高血压(OR=1.697,95%CI:1.305~2.210)和糖尿病(OR=1.932,95%CI:1.413~2.665)与CHD风险升高相关,高密度脂蛋白与CHD风险降低相关(OR=0.351,95%CI:0.221~0.551;均P<0.001)。测试集中,SVM模型的AUC为0.775(95%CI:0.730~0.821),召回率为0.790,F1值为0.784,Brier评分为0.183 7(95%CI:0.168 3~0.202 0)。DeLong检验显示,SVM与Logistic回归、RF、极端梯度提升(XGBoost)及轻量梯度提升机(LGBM)模型的AUC差异无统计学意义(均P>0.05)。决策曲线分析(DCA)显示,在预设的10%~80%阈值概率范围内,依据SVM预测结果决定是否实施进一步检查或转诊的决策策略,其净获益高于“全部实施进一步检查或转诊”和“全部不实施进一步检查或转诊”两种默认参照策略。SHAP分析显示,男性、高血压、糖尿病及炎症指标升高对模型预测呈正向影响,而高密度脂蛋白、左心室射血分数、白蛋白水平升高呈负向影响。 [结论]本研究依托基层易得指标,初步构建适配基层医疗场景、用于CHD二级预防风险分层的轻量化SVM风险预测模型,但后续仍需多中心外部数据验证方可推广应用。

    Abstract:

    Aim To develop and validate a lightweight machine learning model based on readily available clinical indicators to achieve accurate coronary heart disease (CHD) risk prediction. Methods A total of 1 376 patients who underwent coronary angiography were retrospectively enrolled. Thirteen core predictive variables were identified through univariate Logistic regression, Boruta algorithm, and LASSO regression. Using the training set (n=963), seven machine-learning models, including Logistic regression, support vector machine (SVM), and random forest (RF), were constructed. Their performance was evaluated on the test set (n=413) using metrics such as the area under the curve (AUC), recall, precision, F1-score, and Brier score. The SHAP method was employed to analyze the interpretability of the optimal model. Results Univariate Logistic regression analysis identified male sex (OR=2.5,5%CI:1.870~3.202), hypertension (OR=1.7,5%CI:1.305~2.210), and diabetes mellitus (OR=1.2,5%CI:1.413~2.665) as significant correlates of elevated CHD risk, whereas high density lipoprotein (HDL) was associated with a reduced likelihood of CHD (OR=0.1,5%CI:0.221~0.551; all P<0.001). In the independent test cohort, the SVM model achieved an AUC of 0.775 (95%CI:0.730~0.821), with a recall of 0.790, an F1-score of 0.784, and a Brier score of 0.183 7 (95%CI:0.168 3~0.202 0). DeLong's test revealed no statistically significant differences in AUC between the SVM model and the comparative models, including Logistic regression, RF, extreme gradient boosting (XGBoost), and light gradient boosting machine (LGBM) (all P>0.05). Decision curve analysis (DCA) further demonstrated that across the predefined threshold probability range of 10% to 80%, the decision strategy of guiding further diagnostic workup or referral based on SVM-derived predictions yielded superior net clinical benefit compared with the two default reference strategies:“performing universal further examinations or referrals for all participants” and “withholding further examinations or referrals from all participants”. Shapley additive explanations (SHAP) analysis revealed that male sex, hypertension, diabetes mellitus, and elevated inflammatory biomarkers exerted positive contributions to the model's CHD risk prediction, while higher levels of HDL, left ventricular ejection fraction (LVEF), and albumin were associated with negative predictive effects. Conclusion Based on routine indicators accessible at primary medical institutions, this study preliminarily constructed a lightweight SVM risk prediction model suitable for primary care settings to conduct risk stratification in the secondary prevention of CHD. The model exhibited acceptable discrimination performance and interpretability within this single-center cohort, which can provide a reference framework for early identification and risk stratification of CHD in primary care. Further multicenter external data validation is required prior to its clinical promotion and application.

    参考文献
    相似文献
    引证文献
引用本文

林琪雄,张柳,李霞,史树银.面向基层医疗场景的机器学习模型在冠心病二级预防中的风险分层研究[J].中国动脉硬化杂志,2026,34(8):743~751.

复制
文章指标
  • 点击次数:
  • 下载次数:
历史
  • 收稿日期:2026-03-30
  • 最后修改日期:2026-07-30
  • 录用日期:
  • 在线发布日期: 2026-09-24