Komparasi Algoritma Naive Bayes, Random Forest, dan Decision Tree untuk Prediksi Penyakit Stroke Menggunakan Orange Data Mining


  • Wulan Liviana Simbolon * Mail Universitas HKBP Nommensen Pematangsiantar, Pematangsiantar, Indonesia
  • David Ofel Gihon Purba Universitas HKBP Nommensen Pematangsiantar, Pematangsiantar, Indonesia
  • Ningsih Purba Universitas HKBP Nommensen Pematangsiantar, Pematangsiantar, Indonesia
  • Saudurma Sidabutar Universitas HKBP Nommensen Pematangsiantar, Pematangsiantar,, Indonesia
  • Betharya Tampubolon Universitas HKBP Nommensen Pematangsiantar, Pematangsiantar, Indonesia
  • Jaya Tata Hardinata Universitas HKBP Nommensen Pematangsiantar, Pematangsiantar, Indonesia
  • (*) Corresponding Author
Keywords: Machine Learning; Stroke Prediction; Naive Bayes; Random Forest; Decision Tree

Abstract

Stroke is one of the leading causes of death and long-term disability worldwide, making accurate prediction methods essential to support early detection and clinical decision-making. The problem addressed in this study is the lack of evidence regarding which classification algorithm provides the best performance for predicting stroke using the Healthcare Stroke Dataset. This study aims to compare the performance of the Naive Bayes, Random Forest, and Decision Tree algorithms using Orange Data Mining to identify the most effective predictive model. A quantitative approach with a comparative experimental design was employed. The dataset used in this research was the Healthcare Stroke Dataset obtained from Kaggle, consisting of 5,110 records with 12 attributes. The research process included data preprocessing using the Impute widget, feature selection using the Rank widget, classification model development, and model evaluation through 10-fold cross-validation. Performance was assessed using Accuracy, Area Under the Curve (AUC), Precision, Recall, F1-Score, Matthews Correlation Coefficient (MCC), Confusion Matrix, and Receiver Operating Characteristic (ROC) analysis. The results indicate that the Decision Tree algorithm achieved the highest accuracy of 95.1%, followed by Random Forest with 94.8%, while Naive Bayes achieved 92.4%. However, Naive Bayes obtained the highest AUC value of 0.804, demonstrating superior class discrimination capability on an imbalanced dataset. These findings suggest that algorithm selection should not rely solely on accuracy but also consider the model's ability to distinguish between classes consistently. This study contributes to providing recommendations for selecting appropriate classification algorithms to support the development of machine learning-based early stroke prediction systems

References

I. A. I. D. Prayoga and I. G. A. G. A. Kadyanan, “Penggunaan Algoritma C4.5 dan Random Forest guna Meningkatkan Efisiensi Klasifikasi Penyakit Stroke,” Jnatia, vol. 3, no. 3, pp. 553–564, 2025, [Online]. Available: https://www.kaggle.com/datasets/fedesoriano/stroke-prediction-dataset

A. Ristyawan, A. Nugroho, and T. K. Amarya, “Optimasi Preprocessing Model Random Forest untuk Prediksi Stroke,” JATISI (Jurnal Tek. Inform. dan Sist. Informasi), vol. 12, no. 1, 2025, doi: 10.35957/jatisi.v12i1.9587.

H. Hendriyansyah, A. Irma Purnamasari, and T. Suprapti, “Penerapan Algoritma Decision Tree Dalam Klasifikasi Penyakit Stroke Otak,” JATI (Jurnal Mhs. Tek. Inform., vol. 8, no. 3, pp. 3038–3043, 2024, doi: 10.36040/jati.v8i3.9602.

N. Biswas, K. M. M. Uddin, S. T. Rikta, and S. K. Dey, “A comparative analysis of machine learning classifiers for stroke prediction: A predictive analytics approach,” Healthc. Anal., vol. 2, no. October, p. 100116, 2022, doi: 10.1016/j.health.2022.100116.

O. Shobayo, O. Zachariah, M. O. Odusami, and B. Ogunleye, “Prediction of Stroke Disease with Demographic and Behavioural Data Using Random Forest Algorithm,” Analytics, vol. 2, no. 3, pp. 604–617, 2023, doi: 10.3390/analytics2030034.

S. N. Nasokha and R. A. Firmansyah, “Implementasi Algoritma Naïve Bayes untuk Mendiagnosis Kerentanan Warga terhadap Penyakit Stroke,” Technol. Informatics Insight J., vol. 3, no. 2, pp. 114–123, 2024, doi: 10.32639/9hr05d41.

Wartika and A. Nursikuwagus, “Classification for Diagnosing Stroke Using Orange Data Mining,” J. Math. Sci. Informatics, vol. 5, no. 1, pp. 83–94, 2025, doi: 10.46754/jmsi.2025.06.007.

C. Maulana Sidiq, A. Faqih, and G. Dwilestari, “Algoritma Decision Tree C4.5 Digunakan Untuk Mengklasifikasikan Data Stroke,” JATI (Jurnal Mhs. Tek. Inform., vol. 8, no. 2, pp. 1869–1874, 2024, doi: 10.36040/jati.v8i2.8388.

M. F. Banjar, I. Irawati, F. Umar, and L. N. Hayati, “Analysis of Stroke Classification Using Random Forest Method,” Ilk. J. Ilm., vol. 14, no. 3, pp. 186–193, 2022, doi: 10.33096/ilkom.v14i3.1252.186-193.

Abdul Roni, Maria Fransiska Fitriani, Nazwa Aurellia Ainanur, Sumanto, and Ade Surya Budiman, “Comparison of K-Nearest Neighbor Algorithm Performance and Naïve Bayes in Predicting Stroke Disease,” J. Artif. Intell. Eng. Appl., vol. 5, no. 1, pp. 1–6, 2025.

A. F. Riany and G. Testiana, “Penerapan Data Mining untuk Klasifikasi Penyakit Jantung Koroner Menggunakan Algoritma Naïve Bayes,” MDP Student Conf., vol. 2, no. 1, pp. 297–305, 2023, doi: 10.35957/mdp-sc.v2i1.4388.

E. Sabna and O. Dewi, “Prediksi Penyakit Stroke menggunakan Algoritma Decision Tree dan Naïve Bayes,” RIGGS J. Artif. Intell. Digit. Bus., vol. 4, no. 3, pp. 1294–1299, 2025, doi: 10.31004/riggs.v4i3.2132.

Q. Lu, S. Li, H. Xu, H. Shen, and Y. Wei, “A hybrid random forest model for stroke risk prediction,” Front. Neurol., vol. 17, no. April, 2026, doi: 10.3389/fneur.2026.1822304.

R. E. Yunus et al., “Stroke Prognostication in Patients Treated with Thrombolysis Using Random Forest,” Open Neuroimag. J., vol. 17, no. 1, pp. 1–12, 2024, doi: 10.2174/0118744400298093240520070257.

Y. Aulia, A. Andriyansyah, S. Suharjito, and S. W. Nensi, “Analisis Prediksi Stroke dengan Membandingkan Tiga Metode Klasifikasi Decision Tree, Naïve Bayes, dan Random Forest,” J. Ilmu Komput. dan Inform., vol. 3, no. 2, pp. 89–98, 2024, doi: 10.54082/jiki.90.

I. Virgiawan, “Analisis Perbandingan Algoritma Naïve Bayes dan Random Forest Dalam Klasifikasi Penyakit Stroke Pada Puskesmas,” vol. 6, no. 4, pp. 2807–2814, 2025, doi: 10.47065/bits.v6i4.6771.

S. Handayani, “Komparasi Metode Klasifikasi Data Mining untuk Prediksi Penyakit Ginjal,” J. Sist. Inf., vol. 6, no. 1, pp. 34–40, 2017, doi: 10.51998/jsi.v6i1.118.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Komparasi Algoritma Naive Bayes, Random Forest, dan Decision Tree untuk Prediksi Penyakit Stroke Menggunakan Orange Data Mining

Article History
Published: 2026-07-31
Abstract View: 27 times
PDF Download: 16 times
Issue
Vol 4 No 2 (2026): Juli
Pages: 153 - 161
Section
Articles