Fig 1.
Taxonomy of approaches to addressing imbalanced data issue.
Fig 2.
Concepts of oversampling and undersampling techniques.
Fig 3.
An example of using DBSCAN to categorize data points of a dataset for fixed radius (ε) and the minimum number of samples = 5.
Fig 4.
The stages of RN-SMOTE.
Fig 5.
The methodology for dividing the samples of datasets for evaluation in classification methods: (a) splitting the datasets into training and test sets for datasets with sufficient samples; (b) using k-fold cross-validation for evaluating datasets with fewer samples.
Table 1.
The used datasets with their details.
Table 2.
The definition of a confusion matrix.
Table 3.
Imbalanced classification metric.
Table 4.
The average values of evaluation metrics on ILDP, QSAR, Blood and Health risk imbalanced datasets using SVM classifiers and 10-fold cross validation methodology.
Table 5.
The average values of evaluation metrics on ILDP, QSAR, Blood and Health risk imbalanced datasets using ADA classifiers and 10-fold cross validation methodology.
Table 6.
The average values of evaluation metrics on ILDP, QSAR, Blood and Health risk imbalanced datasets using RF classifiers and 10-fold cross validation methodology.
Fig 6.
Comparison of CRN_SMOTE results with RN_SMOTE in ILPD dataset for KSMOTE = 5.
Fig 7.
Comparison of CRN_SMOTE results with RN_SMOTE on QSAR dataset for KSMOTE = 5.
Fig 8.
Comparison of CRN_SMOTE results with RN_SMOTE on Blood dataset for KSMOTE = 5.
Fig 9.
Comparison of CRN_SMOTE results with RN_SMOTE on Health risk dataset for KSMOTE = 5.
Table 7.
A comparison of the RN-SMOTE, SMOTE-Tomek Link, SMOTE-ENN, and the proposed 1CRN-SMOTE methods on the ILPD and QSAR datasets is presented, based on various classification metrics using the Random Forest classifier.
Table 8.
A comparison of the RN-SMOTE, SMOTE-Tomek Link, SMOTE-ENN, and the proposed 1CRN-SMOTE methods on the Blood and Health-risk datasets is presented, based on various classification metrics using the Random Forest classifier.
Fig 10.
Clusters generated using DBSCAN: (a) Clusters consisting of samples from category 1; (b) Clusters consisting of samples from category 2.
Table 9.
A comparison of the CRN-SMOTE and RN-SMOTE methods on the health risk dataset based on different classification metrics using the Random Forest classifier.