Analisis Dampak SMOTE terhadap Feature Importance pada Klasifikasi Data Migraine menggunakan Random Forest dan Extra Trees
Abstract
This study analyzes the impact of the Synthetic Minority Over-sampling Technique (SMOTE) on model performance and feature importance in the classification of migraine patients using Random Forest (RF) and Extra Trees (ET) algorithms. Evaluation was conducted based on recall and F1-Score for the minority class, as well as Permutation Importance analysis. The results indicate that ET, especially when combined with SMOTE (ET + SMOTE), delivers the best performance for the minority class. ET + SMOTE achieved an average F1-Score of 0.7000 and an average recall of 0.8041 using 11 optimal features, indicating better feature efficiency. The application of SMOTE significantly affected the ranking of important features. Although SMOTE improved detection for some minority classes, its impact was not always consistent and occasionally reduced performance on other minority classes. This study concludes that SMOTE alters feature contributions and model interpretability, as well as enhances performance on certain minority classes, particularly when combined with ET.
Keywords
SMOTE, Feature Importance, Random Forest, Extra Trees, Permutation Importance.References
- [1] M. Vincent and S. Wang, “The International Classification of Headache Disorders , 3rd edition,” vol. 38, no. 1, pp. 1–211, 2018, doi: 10.1177/0333102417738202.
- [2] F. Bill and M. G. Foundation, “Global , regional , and national burden of migraine and tension-type headache , 1990 – 2016 : a systematic analysis for the Global Burden of Disease Study 2016,” vol. 17, no. November, 2018, doi: 10.1016/S1474-4422(18)30322-3.
- [3] A. X. Wang, V.-T. Le, H. N. Trung, and B. P. Nguyen, “Addressing imbalance in health data: Synthetic minority oversampling using deep learning,” Comput. Biol. Med., vol. 188, p. 109830, 2025, doi: https://doi.org/10.1016/j.compbiomed.2025.109830.
- [4] J. M. Johnson and T. M. Khoshgoftaar, “Survey on deep learning with class imbalance,” J. Big Data, 2019, doi: 10.1186/s40537-019-0192-5.
- [5] M. Salmi, D. Atif, D. Oliva, A. Abraham, and S. Ventura, Handling imbalanced medical datasets : review of a decade of research, vol. 57, no. 10. Springer Netherlands, 2024. doi: 10.1007/s10462-024-10884-2.
- [6] N. Chawla, K. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic Minority Over-sampling Technique,” ArXiv, vol. abs/1106.1813, 2002, [Online]. Available: https://api.semanticscholar.org/CorpusID:1554582
- [7] M. T. Ribeiro and C. Guestrin, “‘ Why Should I Trust You ?’ Explaining the Predictions of Any Classifier,” 2016.
- [8] C. Molnar, “Interpretable Machine Learning A Guide for Making Black Box Models Explainable”.
- [9] F. Doshi-velez and B. Kim, “Towards A Rigorous Science of Interpretable Machine Learning,” no. Ml, pp. 1–13, 2017.
- [10] C. Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,” Nat. Mach. Intell., vol. 1, no. 5, pp. 206–215, 2019, doi: 10.1038/s42256-019-0048-x.
- [11] A. Barredo Arrieta et al., “Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI,” Inf. Fusion, vol. 58, pp. 82–115, 2020, doi: https://doi.org/10.1016/j.inffus.2019.12.012.
- [12] L. Breiman, “Random Forests,” vol. 45, no. 1, 2001, doi: 10.1007/978-3-030-62008-0_35.
- [13] P. Geurts, D. Ernst, and L. Wehenkel, “Extremely randomized trees,” Mach. Learn., vol. 63, no. 1, pp. 3–42, 2006, doi: 10.1007/s10994-006-6226-1.
- [14] E. Scornet, “Trees , forests , and impurity-based variable importance in regression,” pp. 1–40.
- [15] C. Strobl, A. Boulesteix, A. Zeileis, and T. Hothorn, “Bias in random forest variable importance measures : Illustrations , sources and a solution,” vol. 21, 2007, doi: 10.1186/1471-2105-8-25.
Most read articles by the same author(s)
- Henny Leidiyana, Siti Nurajizah, Pendekatan Hibrida Statistik dan Machine Learning untuk Peramalan Jumlah Kunjungan Turis , Jurnal Komtika (Komputasi dan Informatika): Vol. 9 No. 2 (2025)
- Henny Leidiyana, Arya Anugrah, Aplikasi Pengendalian Persediaan Barang Berbasis Android dengan Metode Economic Order Quantity (EOQ) pada Bengkel Dunia Motor , Jurnal Komtika (Komputasi dan Informatika): Vol. 4 No. 2 (2020)
- Henny Leidiyana, Titik Misriati, Riska Aryanti, Klasifikasi Sentimen Terhadap Kebijakan Tapera Menggunakan Komparasi Machine Learning dan SMOTE , Jurnal Komtika (Komputasi dan Informatika): Vol. 8 No. 2 (2024)
- Ali Rachman, Henny Leidiyana, Sistem Informasi Fasilitas di DKI Jakarta berbasis Android dengan Algoritma Floyd Warshall , Jurnal Komtika (Komputasi dan Informatika): Vol. 4 No. 1 (2020)
- Henny Leidiyana, Risvan Dwi Hariyanto, Sistem Pakar untuk Mendiagnosa Penyakit Persendian Menggunakan Metode Certainty Factor , Jurnal Komtika (Komputasi dan Informatika): Vol. 4 No. 1 (2020)
