Table 1.
Statistics of the datasets.
Fig 1.
Haematological features used in different datasets.
Fig 2.
Point biserial correlation coefficients between SARS-COV-2 results and individual features for a) dataset 1 b) dataset 2 and c) dataset 3.
Parameters with p values < 0.05 are shown in blue, remaining values are in red.
Fig 3.
Description of data sources used for training, testing and external validation of different ML-models based on haematological features for COVID-19 characterization.
Fig 4.
Receiver Operating Characteristics Curves (ROC) across different ML models for four sub-datasets.
Table 2.
Internal evaluation of the XGBoost model on different datasets and comparison with published datasets.
Interval computed within 95% confidence limit.
Fig 5.
Comparative performances of different sub-datasets trained on XGBoost model.
The datasets with published AUC scores 21were compared.
Table 3.
External evaluation of XGBoost algorithm based on a) 4- hematological features and b) 14-hematological features trained and tested across different datasets.
Interval computed within 95% confidence limit.
Fig 6.
The external performance diagram generated using the online tool (https://qualiml.pythonanywhere.com), depicted the results from external validation studies on COVID-19 diagnosis trained and tested on i) same population–Brazil-Brazil (Br) and ii) different population–Italy-Brazil (It).
The Minimum sample size, depicted by the hue brightness. The width of the ellipse equals to the width of 95% confidence interval with respect to the given performance metrics.
Table 4.
XGBoost model performance metrics were shown from one hundred iterations on external validation datasets for a) 2-four-features (Italian) /1-four-features (Brazilian), and b) 3-fourteen-features (Brazilian) / 1-fourteen-feature (Brazilian) with IV-perturbation, and IV-perturbation plus augmentation methods.
Baseline values were reported. The baseline data was generated using lower fuzziness with resampling in a single step, in contrast to perturbation and perturbation plus augmentation data, where higher fuzziness was applied using sequential resampling of the baseline data. The standard deviation values were shown in parentheses.
Fig 7.
Distributions of four hematological parameters across four different datasets (two training datasets–Dataset 1-four-features and Dataset 2-four-features and two test datasets–early and advance).
The hematological parameters are–a) platelet, b) leukocyte c) eosinophil and d) monocyte. These distributions indicate the proximity of the individual test datasets to the training datasets.
Table 5.
Blind prediction of XGBoost model trained on dataset 2-four-feature and tested on W.E.-early and W.E.-advanced datasets.
The early and advanced datasets contain only COVID-19-positive patient results; no negatives were available. Hence, only sensitivity values reported.