Abstract:Aim To develop and validate a lightweight machine learning model based on readily available clinical indicators to achieve accurate coronary heart disease (CHD) risk prediction. Methods A total of 1 376 patients who underwent coronary angiography were retrospectively enrolled. Thirteen core predictive variables were identified through univariate Logistic regression, Boruta algorithm, and LASSO regression. Using the training set (n=963), seven machine-learning models, including Logistic regression, support vector machine (SVM), and random forest (RF), were constructed. Their performance was evaluated on the test set (n=413) using metrics such as the area under the curve (AUC), recall, precision, F1-score, and Brier score. The SHAP method was employed to analyze the interpretability of the optimal model. Results Univariate Logistic regression analysis identified male sex (OR=2.5,5%CI:1.870~3.202), hypertension (OR=1.7,5%CI:1.305~2.210), and diabetes mellitus (OR=1.2,5%CI:1.413~2.665) as significant correlates of elevated CHD risk, whereas high density lipoprotein (HDL) was associated with a reduced likelihood of CHD (OR=0.1,5%CI:0.221~0.551; all P<0.001). In the independent test cohort, the SVM model achieved an AUC of 0.775 (95%CI:0.730~0.821), with a recall of 0.790, an F1-score of 0.784, and a Brier score of 0.183 7 (95%CI:0.168 3~0.202 0). DeLong's test revealed no statistically significant differences in AUC between the SVM model and the comparative models, including Logistic regression, RF, extreme gradient boosting (XGBoost), and light gradient boosting machine (LGBM) (all P>0.05). Decision curve analysis (DCA) further demonstrated that across the predefined threshold probability range of 10% to 80%, the decision strategy of guiding further diagnostic workup or referral based on SVM-derived predictions yielded superior net clinical benefit compared with the two default reference strategies:“performing universal further examinations or referrals for all participants” and “withholding further examinations or referrals from all participants”. Shapley additive explanations (SHAP) analysis revealed that male sex, hypertension, diabetes mellitus, and elevated inflammatory biomarkers exerted positive contributions to the model's CHD risk prediction, while higher levels of HDL, left ventricular ejection fraction (LVEF), and albumin were associated with negative predictive effects. Conclusion Based on routine indicators accessible at primary medical institutions, this study preliminarily constructed a lightweight SVM risk prediction model suitable for primary care settings to conduct risk stratification in the secondary prevention of CHD. The model exhibited acceptable discrimination performance and interpretability within this single-center cohort, which can provide a reference framework for early identification and risk stratification of CHD in primary care. Further multicenter external data validation is required prior to its clinical promotion and application.