Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Fig 1.

Schematic diagram of the workflow for evaluating the effectiveness of sampling for imbalanced classification.

The process is repeated five times (i = 1, 2, 3, 4, 5) with repeated random division of an imbalanced classification dataset.

More »

Fig 1 Expand

Fig 2.

Heatmap of the difference in the area under the precision-recall curve between classification with and without sampling on the 31 imbalanced datasets.

Combinations of the seven sampling methods [i.e., random oversampling (O_Random), synthetic minority oversampling technique (O_SMOTE), borderline synthetic minority oversampling technique (O_Border), random undersampling (U_Random), condensed nearest neighbors undersampling (U_Condensed), NearMiss2 (U_NearMiss), and SMOTETomek] and eight machine learning methods [i.e., adaptive boosting (AdaBoost), extreme gradient boosting (XGBoost), random forests (RFs), support vector machines (SVMs), the linear discriminant analysis (LDA), lasso, ridge, and elastic net] were compared using the 31 datasets.

More »

Fig 2 Expand

Fig 3.

Heatmap of the difference in the area under the receiver operating characteristics curve between classification with and without sampling on the 31 imbalanced datasets.

Combinations of the seven sampling methods [i.e., random oversampling (O_Random), synthetic minority oversampling technique (O_SMOTE), borderline synthetic minority oversampling technique (O_Border), random undersampling (U_Random), condensed nearest neighbors undersampling (U_Condensed), NearMiss2 (U_NearMiss), and SMOTETomek] and eight machine learning methods [i.e., adaptive boosting (AdaBoost), extreme gradient boosting (XGBoost), random forests (RFs), support vector machines (SVMs), the linear discriminant analysis (LDA), lasso, ridge, and elastic net] were compared using the 31 datasets.

More »

Fig 3 Expand

Fig 4.

Comparison of the effectiveness of the seven sampling methods.

The number of cases in which a sampling method enhanced (blue) or reduced (red) the performance in (A) the area under the precision-recall curve (AUPRC) and (B) the area under the receiver operating characteristics curve (AUROC) is shown. Seven sampling methods—random oversampling (O_Random), synthetic minority oversampling technique (O_SMOTE), borderline synthetic minority oversampling technique (O_Border), random undersampling (U_Random), condensed nearest neighbors undersampling (U_Condensed), NearMiss2 (U_NearMiss), and SMOTETomek—were compared.

More »

Fig 4 Expand

Fig 5.

Comparison of machine learning methods by the effectiveness of sampling.

The number of cases in which (A) the area under the precision-recall curve (AUPRC) and (B) the area under the receiver operating characteristics curve (AUROC) of a machine learning method were improved (blue) or reduced (red) by sampling. Eight machine learning methods—adaptive boosting (AdaBoost), extreme gradient boosting (XGBoost), random forests (RFs), support vector machines (SVMs), the linear discriminant analysis (LDA), lasso, ridge, and elastic net—were compared.

More »

Fig 5 Expand

Table 1.

Number of datasets on which a combination of machine learning and sampling methods performed the best in terms of the area under the precision-recall curve.

More »

Table 1 Expand

Table 2.

Number of datasets on which a combination of machine learning and sampling methods performed the best in terms of the area under the receiver operating characteristics curve.

More »

Table 2 Expand

Table 3.

Performance of linear discriminant analysis (LDA) on the Letter_a dataset with and without the four sampling methods.

More »

Table 3 Expand

Table 4.

Performance of linear discriminant analysis (LDA) on the Fraud_Detection dataset with and without the four sampling methods.

More »

Table 4 Expand

Fig 6.

(A) Precision-recall (PR) and (B) receiver operating characteristics (ROC) curves of linear discriminant analysis with and without the four sampling methods on the Letter_a dataset. The PR and ROC curves on the test dataset of the first fold of the first iteration of the 5x2 cross-validation run are shown. Four sampling methods were compared: random oversampling (O_Random), synthetic minority oversampling technique (O_SMOTE), random undersampling (U_Random), and SMOTETomek. AUC indicates the area under the PR or ROC curve.

More »

Fig 6 Expand

Fig 7.

(A) Precision-recall (PR) and (B) receiver operating characteristics (ROC) curves of linear discriminant analysis with and without the four sampling methods on the Fraud_Detection dataset. The PR and ROC curves on the test dataset of the first fold of the first iteration of the 5x2 cross-validation run are shown. Four sampling methods were compared: random oversampling (O_Random), synthetic minority oversampling technique (O_SMOTE), random undersampling (U_Random), and SMOTETomek. AUC means the area under the PR or ROC curve.

More »

Fig 7 Expand