Fig 1.
Schematic diagram of the workflow for evaluating the effectiveness of sampling for imbalanced classification.
The process is repeated five times (i = 1, 2, 3, 4, 5) with repeated random division of an imbalanced classification dataset.
Fig 2.
Heatmap of the difference in the area under the precision-recall curve between classification with and without sampling on the 31 imbalanced datasets.
Combinations of the seven sampling methods [i.e., random oversampling (O_Random), synthetic minority oversampling technique (O_SMOTE), borderline synthetic minority oversampling technique (O_Border), random undersampling (U_Random), condensed nearest neighbors undersampling (U_Condensed), NearMiss2 (U_NearMiss), and SMOTETomek] and eight machine learning methods [i.e., adaptive boosting (AdaBoost), extreme gradient boosting (XGBoost), random forests (RFs), support vector machines (SVMs), the linear discriminant analysis (LDA), lasso, ridge, and elastic net] were compared using the 31 datasets.
Fig 3.
Heatmap of the difference in the area under the receiver operating characteristics curve between classification with and without sampling on the 31 imbalanced datasets.
Combinations of the seven sampling methods [i.e., random oversampling (O_Random), synthetic minority oversampling technique (O_SMOTE), borderline synthetic minority oversampling technique (O_Border), random undersampling (U_Random), condensed nearest neighbors undersampling (U_Condensed), NearMiss2 (U_NearMiss), and SMOTETomek] and eight machine learning methods [i.e., adaptive boosting (AdaBoost), extreme gradient boosting (XGBoost), random forests (RFs), support vector machines (SVMs), the linear discriminant analysis (LDA), lasso, ridge, and elastic net] were compared using the 31 datasets.
Fig 4.
Comparison of the effectiveness of the seven sampling methods.
The number of cases in which a sampling method enhanced (blue) or reduced (red) the performance in (A) the area under the precision-recall curve (AUPRC) and (B) the area under the receiver operating characteristics curve (AUROC) is shown. Seven sampling methods—random oversampling (O_Random), synthetic minority oversampling technique (O_SMOTE), borderline synthetic minority oversampling technique (O_Border), random undersampling (U_Random), condensed nearest neighbors undersampling (U_Condensed), NearMiss2 (U_NearMiss), and SMOTETomek—were compared.
Fig 5.
Comparison of machine learning methods by the effectiveness of sampling.
The number of cases in which (A) the area under the precision-recall curve (AUPRC) and (B) the area under the receiver operating characteristics curve (AUROC) of a machine learning method were improved (blue) or reduced (red) by sampling. Eight machine learning methods—adaptive boosting (AdaBoost), extreme gradient boosting (XGBoost), random forests (RFs), support vector machines (SVMs), the linear discriminant analysis (LDA), lasso, ridge, and elastic net—were compared.
Table 1.
Number of datasets on which a combination of machine learning and sampling methods performed the best in terms of the area under the precision-recall curve.
Table 2.
Number of datasets on which a combination of machine learning and sampling methods performed the best in terms of the area under the receiver operating characteristics curve.
Table 3.
Performance of linear discriminant analysis (LDA) on the Letter_a dataset with and without the four sampling methods.
Table 4.
Performance of linear discriminant analysis (LDA) on the Fraud_Detection dataset with and without the four sampling methods.
Fig 6.
(A) Precision-recall (PR) and (B) receiver operating characteristics (ROC) curves of linear discriminant analysis with and without the four sampling methods on the Letter_a dataset. The PR and ROC curves on the test dataset of the first fold of the first iteration of the 5x2 cross-validation run are shown. Four sampling methods were compared: random oversampling (O_Random), synthetic minority oversampling technique (O_SMOTE), random undersampling (U_Random), and SMOTETomek. AUC indicates the area under the PR or ROC curve.
Fig 7.
(A) Precision-recall (PR) and (B) receiver operating characteristics (ROC) curves of linear discriminant analysis with and without the four sampling methods on the Fraud_Detection dataset. The PR and ROC curves on the test dataset of the first fold of the first iteration of the 5x2 cross-validation run are shown. Four sampling methods were compared: random oversampling (O_Random), synthetic minority oversampling technique (O_SMOTE), random undersampling (U_Random), and SMOTETomek. AUC means the area under the PR or ROC curve.