Jurnal Komtika (Komputasi dan Informatika)

Articles

Analisis Dampak SMOTE terhadap Feature Importance pada Klasifikasi Data Migraine menggunakan Random Forest dan Extra Trees

Henny Leidiyana

Abstract

This study analyzes the impact of the Synthetic Minority Over-sampling Technique (SMOTE) on model performance and feature importance in the classification of migraine patients using Random Forest (RF) and Extra Trees (ET) algorithms. Evaluation was conducted based on recall and F1-Score for the minority class, as well as Permutation Importance analysis. The results indicate that ET, especially when combined with SMOTE (ET + SMOTE), delivers the best performance for the minority class. ET + SMOTE achieved an average F1-Score of 0.7000 and an average recall of 0.8041 using 11 optimal features, indicating better feature efficiency. The application of SMOTE significantly affected the ranking of important features. Although SMOTE improved detection for some minority classes, its impact was not always consistent and occasionally reduced performance on other minority classes. This study concludes that SMOTE alters feature contributions and model interpretability, as well as enhances performance on certain minority classes, particularly when combined with ET.

Keywords

SMOTE, Feature Importance, Random Forest, Extra Trees, Permutation Importance.

References

  1. [1] M. Vincent and S. Wang, “The International Classification of Headache Disorders , 3rd edition,” vol. 38, no. 1, pp. 1–211, 2018, doi: 10.1177/0333102417738202.
  2. [2] F. Bill and M. G. Foundation, “Global , regional , and national burden of migraine and tension-type headache , 1990 – 2016 : a systematic analysis for the Global Burden of Disease Study 2016,” vol. 17, no. November, 2018, doi: 10.1016/S1474-4422(18)30322-3.
  3. [3] A. X. Wang, V.-T. Le, H. N. Trung, and B. P. Nguyen, “Addressing imbalance in health data: Synthetic minority oversampling using deep learning,” Comput. Biol. Med., vol. 188, p. 109830, 2025, doi: https://doi.org/10.1016/j.compbiomed.2025.109830.
  4. [4] J. M. Johnson and T. M. Khoshgoftaar, “Survey on deep learning with class imbalance,” J. Big Data, 2019, doi: 10.1186/s40537-019-0192-5.
  5. [5] M. Salmi, D. Atif, D. Oliva, A. Abraham, and S. Ventura, Handling imbalanced medical datasets : review of a decade of research, vol. 57, no. 10. Springer Netherlands, 2024. doi: 10.1007/s10462-024-10884-2.
  6. [6] N. Chawla, K. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic Minority Over-sampling Technique,” ArXiv, vol. abs/1106.1813, 2002, [Online]. Available: https://api.semanticscholar.org/CorpusID:1554582
  7. [7] M. T. Ribeiro and C. Guestrin, “‘ Why Should I Trust You ?’ Explaining the Predictions of Any Classifier,” 2016.
  8. [8] C. Molnar, “Interpretable Machine Learning A Guide for Making Black Box Models Explainable”.
  9. [9] F. Doshi-velez and B. Kim, “Towards A Rigorous Science of Interpretable Machine Learning,” no. Ml, pp. 1–13, 2017.
  10. [10] C. Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,” Nat. Mach. Intell., vol. 1, no. 5, pp. 206–215, 2019, doi: 10.1038/s42256-019-0048-x.
  11. [11] A. Barredo Arrieta et al., “Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI,” Inf. Fusion, vol. 58, pp. 82–115, 2020, doi: https://doi.org/10.1016/j.inffus.2019.12.012.
  12. [12] L. Breiman, “Random Forests,” vol. 45, no. 1, 2001, doi: 10.1007/978-3-030-62008-0_35.
  13. [13] P. Geurts, D. Ernst, and L. Wehenkel, “Extremely randomized trees,” Mach. Learn., vol. 63, no. 1, pp. 3–42, 2006, doi: 10.1007/s10994-006-6226-1.
  14. [14] E. Scornet, “Trees , forests , and impurity-based variable importance in regression,” pp. 1–40.
  15. [15] C. Strobl, A. Boulesteix, A. Zeileis, and T. Hothorn, “Bias in random forest variable importance measures : Illustrations , sources and a solution,” vol. 21, 2007, doi: 10.1186/1471-2105-8-25.