Analisis Komparatif Algoritma Pembelajaran Mesin untuk Prediksi Risiko Hipertensi: Gradient Boosting, Random Forest, XGBoost, Regresi Logistik, dan Support Vector Machine

Penulis

  • theo vhaldino Universitas Aisyiyah Palembang
  • Arpa Pauziah Universitas Aisyiyah Palembang
  • Khoirin Universitas Aisyiyah Palembang
  • Rudiansyah Universitas Aisyiyah Palembang
  • Debbi Susita Yuzzaki Universitas Aisyiyah Palembang
  • Ahmad Syarif Universitas 'Aisyiyah Palembang

DOI:

https://doi.org/10.52523/jhast.v4i2.128

Kata Kunci:

Pembelajaran Mesin; Gradient Boosting; XGBoost; Prediksi Risiko.

Abstrak

Hipertensi tetap menjadi salah satu penyebab utama morbiditas kardiovaskular, sementara skrining di masyarakat sering kali masih bergantung pada pengukuran yang dilakukan secara terlambat atau secara oportunistik. Penelitian ini membandingkan Logistic Regression, Random Forest, Gradient Boosting, XGBoost, dan Support Vector Machine untuk memprediksi risiko hipertensi menggunakan dataset publik yang terdiri dari 1.985 data pasien dan sepuluh prediktor demografis, klinis, serta gaya hidup. Fitur numerik distandarisasi, fitur kategorikal dikonversi menggunakan one-hot encoding, sedangkan nilai obat yang hilang diimputasi menggunakan modus. Model dievaluasi menggunakan pembagian data 80:20 secara stratifikasi serta 5-fold cross-validation. Model terbaik kemudian dioptimalkan menggunakan RandomizedSearchCV dan dianalisis melalui feature importance serta ablation study. Gradient Boosting yang telah dituning menghasilkan performa terbaik dengan F1-score cross-validation sebesar 0,9994 dan ROC-AUC sebesar 1,0000. Pada 397 data pengujian, model tersebut memperoleh akurasi 0,9975, precision 1,0000, recall 0,9951, F1-score 0,9976, dan ROC-AUC 1,0000, dengan hanya ditemukan satu false negative. Hasil ablation study menunjukkan bahwa penghilangan fitur riwayat tekanan darah menyebabkan F1-score pada data pengujian menurun menjadi 0,7681, yang mengonfirmasi bahwa fitur tersebut memberikan kontribusi yang dominan terhadap prediksi. Model ini menunjukkan potensi yang baik sebagai pendukung skrining hipertensi, tetapi diperlukan validasi klinis eksternal sebelum model dapat digunakan untuk tujuan diagnostic.

Referensi

W. H. Organization, Global report on hypertension: The race against a silent killer. World Health Organization, 2023.

B. Zhou et al., “Worldwide trends in hypertension prevalence and progress in treatment and control from 1990 to 2019: a pooled analysis of 1201 population-representative studies with 104 million participants,” Lancet, vol. 398, no. 10304, pp. 957–980, 2021.

T. Riskesdas, “Laporan nasional RISKESDAS 2018,” Jakarta Lemb. Penerbit Badan Penelit. dan Pengemb. Kesehat., 2019.

K. T. Mills et al., “Global disparities of hypertension prevalence and control: a systematic analysis of population-based studies from 90 countries,” Circulation, vol. 134, no. 6, pp. 441–450, 2016.

G. A. Mensah et al., “Global burden of cardiovascular diseases and risks, 1990-2022,” J. Am. Coll. Cardiol., vol. 82, no. 25, pp. 2350–2473, 2023.

M. M. Ahsan, S. A. Luna, and Z. Siddique, “Machine-learning-based disease diagnosis: A comprehensive review,” in Healthcare, 2022, vol. 10, no. 3, p. 541.

S. M. S. Islam et al., “Machine learning approaches for predicting hypertension and its associated factors using population-level data from three South Asian countries,” Front. Cardiovasc. Med., vol. 9, p. 839379, 2022.

H. Zhao et al., “Predicting the risk of hypertension based on several easy-to-collect risk factors: a machine learning method,” Front. public Heal., vol. 9, p. 429, 2021.

Y. Peng, J. Xu, L. Ma, and J. Wang, “Prediction of hypertension risks with feature selection and XGBoost,” J. Mech. Med. Biol., vol. 21, no. 05, p. 2140028, 2021.

J. Du et al., “Developing a hypertension visualization risk prediction system utilizing machine learning and health check-up data,” Sci. Rep., vol. 13, no. 1, p. 18953, 2023.

S. H. Hwang et al., “Machine learning–based prediction for incident hypertension based on regular health checkup data: derivation and validation in 2 independent nationwide cohorts in South Korea and Japan,” J. Med. Internet Res., vol. 26, p. e52794, 2024.

F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011.

T. Chen and C. Guestrin, “Xgboost: A scalable tree boosting system,” in Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, 2016, pp. 785–794.

G. S. Collins et al., “TRIPOD+ AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods,” bmj, vol. 385, 2024.

R. F. Wolff et al., “PROBAST: a tool to assess the risk of bias and applicability of prediction model studies,” Ann. Intern. Med., vol. 170, no. 1, pp. 51–58, 2019.

J. C. Stoltzfus, “Logistic regression: a brief primer,” Acad. Emerg. Med., vol. 18, no. 10, pp. 1099–1104, 2011.

L. Breiman, “Random forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001.

J. H. Friedman, “Greedy function approximation: a gradient boosting machine,” Ann. Stat., pp. 1189–1232, 2001.

C. Cortes and V. Vapnik, “Support-vector networks,” Mach. Learn., vol. 20, no. 3, pp. 273–297, 1995.

T. Fawcett, “An introduction to ROC analysis,” Pattern Recognit. Lett., vol. 27, no. 8, pp. 861–874, 2006.

Diterbitkan

2026-09-30

Cara Mengutip

vhaldino, theo, Arpa Pauziah, Khoirin, Rudiansyah, Debbi Susita Yuzzaki, & Ahmad Syarif. (2026). Analisis Komparatif Algoritma Pembelajaran Mesin untuk Prediksi Risiko Hipertensi: Gradient Boosting, Random Forest, XGBoost, Regresi Logistik, dan Support Vector Machine. Jurnal Kesehatan Terapan Sains Dan Teknologi, 4(2), 100–111. https://doi.org/10.52523/jhast.v4i2.128