Comparative Analysis of Machine Learning Algorithms for Hypertension Risk Prediction: Gradient Boosting, Random Forest, XGBoost, Logistic Regression, and Support Vector Machine

Authors

  • theo vhaldino Universitas Aisyiyah Palembang
  • Arpa Pauziah Universitas Aisyiyah Palembang
  • Khoirin Universitas Aisyiyah Palembang
  • Rudiansyah Universitas Aisyiyah Palembang
  • Debbi Susita Yuzzaki Universitas Aisyiyah Palembang
  • Ahmad Syarif Universitas Aisyiyah Palembang

DOI:

https://doi.org/10.52523/jhast.v4i2.128

Keywords:

Pembelajaran Mesin; Gradient Boosting; XGBoost; Prediksi Risiko.

Abstract

Hypertension remains a major contributor to cardiovascular morbidity, yet community screening often relies on late or opportunistic measurements. This study compares Logistic Regression, Random Forest, Gradient Boosting, XGBoost, and Support Vector Machine for hypertension risk prediction using a public dataset of 1,985 patient records and ten demographic, clinical, and lifestyle predictors. Numerical features were standardized, categorical features were one-hot encoded, and missing medication values were imputed using the mode. Models were evaluated using an 80:20 stratified train-test split and 5-fold cross-validation. The best model was optimized with RandomizedSearchCV and examined through feature importance and ablation study. Tuned Gradient Boosting achieved the highest cross-validation F1-score of 0.9994 and ROC-AUC of 1.0000. On 397 test records, it obtained 0.9975 accuracy, 1.0000 precision, 0.9951 recall, 0.9976 F1-score, and 1.0000 ROC-AUC, with only one false negative. Ablation showed that removing blood pressure history reduced test F1-score to 0.7681, confirming its dominant contribution. The model is promising for screening support, but external clinical validation is required before diagnostic use.

References

W. H. Organization, Global report on hypertension: The race against a silent killer. World Health Organization, 2023.

B. Zhou et al., “Worldwide trends in hypertension prevalence and progress in treatment and control from 1990 to 2019: a pooled analysis of 1201 population-representative studies with 104 million participants,” Lancet, vol. 398, no. 10304, pp. 957–980, 2021.

T. Riskesdas, “Laporan nasional RISKESDAS 2018,” Jakarta Lemb. Penerbit Badan Penelit. dan Pengemb. Kesehat., 2019.

K. T. Mills et al., “Global disparities of hypertension prevalence and control: a systematic analysis of population-based studies from 90 countries,” Circulation, vol. 134, no. 6, pp. 441–450, 2016.

G. A. Mensah et al., “Global burden of cardiovascular diseases and risks, 1990-2022,” J. Am. Coll. Cardiol., vol. 82, no. 25, pp. 2350–2473, 2023.

M. M. Ahsan, S. A. Luna, and Z. Siddique, “Machine-learning-based disease diagnosis: A comprehensive review,” in Healthcare, 2022, vol. 10, no. 3, p. 541.

S. M. S. Islam et al., “Machine learning approaches for predicting hypertension and its associated factors using population-level data from three South Asian countries,” Front. Cardiovasc. Med., vol. 9, p. 839379, 2022.

H. Zhao et al., “Predicting the risk of hypertension based on several easy-to-collect risk factors: a machine learning method,” Front. public Heal., vol. 9, p. 429, 2021.

Y. Peng, J. Xu, L. Ma, and J. Wang, “Prediction of hypertension risks with feature selection and XGBoost,” J. Mech. Med. Biol., vol. 21, no. 05, p. 2140028, 2021.

J. Du et al., “Developing a hypertension visualization risk prediction system utilizing machine learning and health check-up data,” Sci. Rep., vol. 13, no. 1, p. 18953, 2023.

S. H. Hwang et al., “Machine learning–based prediction for incident hypertension based on regular health checkup data: derivation and validation in 2 independent nationwide cohorts in South Korea and Japan,” J. Med. Internet Res., vol. 26, p. e52794, 2024.

F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011.

T. Chen and C. Guestrin, “Xgboost: A scalable tree boosting system,” in Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, 2016, pp. 785–794.

G. S. Collins et al., “TRIPOD+ AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods,” bmj, vol. 385, 2024.

R. F. Wolff et al., “PROBAST: a tool to assess the risk of bias and applicability of prediction model studies,” Ann. Intern. Med., vol. 170, no. 1, pp. 51–58, 2019.

J. C. Stoltzfus, “Logistic regression: a brief primer,” Acad. Emerg. Med., vol. 18, no. 10, pp. 1099–1104, 2011.

L. Breiman, “Random forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001.

J. H. Friedman, “Greedy function approximation: a gradient boosting machine,” Ann. Stat., pp. 1189–1232, 2001.

C. Cortes and V. Vapnik, “Support-vector networks,” Mach. Learn., vol. 20, no. 3, pp. 273–297, 1995.

T. Fawcett, “An introduction to ROC analysis,” Pattern Recognit. Lett., vol. 27, no. 8, pp. 861–874, 2006.

Downloads

Published

2026-09-30

How to Cite

vhaldino, theo, Arpa Pauziah, Khoirin, Rudiansyah, Debbi Susita Yuzzaki, & Ahmad Syarif. (2026). Comparative Analysis of Machine Learning Algorithms for Hypertension Risk Prediction: Gradient Boosting, Random Forest, XGBoost, Logistic Regression, and Support Vector Machine. Journal Health Applied Science and Technology, 4(2), 100–111. https://doi.org/10.52523/jhast.v4i2.128