Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Use of machine learning to detect Escherichia coli in drinking water in Bangladesh

Abstract

Escherichia coli (E. coli) is a key indicator of fecal contamination in freshwater and can signal the presence of other harmful bacteria and viruses. The aim of the study is to evaluate the performance of machine learning (ML) tools to detect E. coli in drinking water in Bangladesh using surveillance data under two scenarios: an imbalanced dataset and a balanced dataset. We utilized data from the 2019 Bangladesh Multiple Indicator Cluster Survey, which included a total of 6,069 household drinking water samples. We used agglomerative hierarchical clustering with Ward’s linkage to identify district-level hotspots. Extreme Gradient Boosting with SHapley Additive exPlanations values were used for feature selection, and the Synthetic Minority Over-sampling Technique (SMOTE) was used to address class imbalance in the classification task. We applied nine classical ML models in this study: Adaptive Boosting (AdaBoost), Decision Trees (DT), Gradient Boosting Algorithm (GBA), k-Nearest Neighbors (KNN), Light Gradient-Boosting Machine (LightGBM), Logistic Regression (LR), Naïve Bayes (NB), Random Forest (RF), and Support Vector Machine (SVM), along with a Deep Learning Multi-Layer Perceptron (DL-MLP) model to predict the risk of E. coli contamination (REcC) in water. Model performance was evaluated using accuracy, precision, recall, F1 score, Cohen Kappa, area under the curve (AUC), and a violin plot. E. coli contamination in drinking water was detected in 39.2% (95% CI: 37.4–41.2) of households. Bandarban district had the highest REcC. After applying SMOTE and 10-fold cross-validation with hyperparameter tuning, model performance was more consistent across algorithms. In terms of model evaluation, AdaBoost slightly outperformed the others with an accuracy of 81.6%, Cohen kappa statistic of 19.4%, precision of 82.2%, recall of 99%, F1-score of 89.8%, and an AUC of 68.6%. Ensemble model for example AdaBoost and GBA models had ability to accurately classify drinking water samples with respect to the presence of E. coli using surveillance data than others selected model in this study.

Introduction

Enterobacteriaceae, particularly Escherichia coli (E. coli), are the dominant commensal bacteria found in the gastrointestinal tracts of humans and other warm-blooded animals, and are also considered important pathogens [1]. Contamination of household drinking water with E. coli is a significant public health concern, particularly amongst children [2]. Neonatal meningitis, urinary tract infections, enteritis, septicaemia, and other clinical illnesses are commonly caused by E. coli [1].

Over half of the global population uses groundwater as their primary drinking water source [3]. Conflicting demands have led to over-exploitation of 20% of aquifers, with contamination from chemicals, radionuclides, and microorganisms, as well as improper waste and wastewater management, posing significant risks to groundwater supplies [3]. Acute gastrointestinal illnesses (AGIs) brought on by drinking contaminated water spread irregularly throughout location and time as a result of microbiological groundwater contamination events. Contaminated groundwater is the leading cause of waterborne illness in the US, accounting for 64% of drinking water outbreaks between 1989 and 2002, with up to 150 million North American residents relying on groundwater systems [4].

The burden of waterborne illness is increased by the possibility of exposure to enteric pathogens [5]. Enteric infections are more common among sensitive subpopulations, such as the young, old, pregnant, and malnourished. Both low-and-middle income and high-income countries have high rates of infections caused by enteric pathogens. Groundwater contamination can occur in different ways, such as exposure to septic tanks, sewage, municipal wastewater treatment facilities, wildlife, and agricultural residues [5]. Nearly 60% of private well owners in Ireland do not use an appropriate domestic water treatment system [6].

In many countries, storing drinking water is a common practice due to intermittent or lack of direct access to drinking water within homes. Water storage has been identified as an important risk factor for diarrheal diseases [7]. Unhygienic handling during storage can lead to microbial contamination of drinking water [8]. The type of water source also affects contamination levels [9]. Additionally, water supplies may become contaminated in case of open defecation or subpar toilets are used. Social characteristics such as income will also be associated with contamination levels. Households categorised as middle-class or lower-class in Bangladesh were shown to be at a higher risk of drinking water contamination when they obtained their water from sources with elevated levels of E. coli [2]. The possibility of increased contamination levels was greatly decreased by treating drinking water, although pet ownership was strongly linked to recontamination.

From a statistical perspective, different approaches were used to identify the presence of E. coli contamination among household drinking water. A study in North America used pooled analysis to investigate groundwater contamination in the US and Canada [5]. Another researcher used association rule learning, regression analysis, and variable discretization techniques to identify the relationship between contamination and other predictors [3].

Machine learning (ML) has emerged as a significant solution for various applications such as bioinformatics [10,11], and activity recognition in recent decades. Deep learning (DL) models are powerful tools for computation on large datasets [12]. Since the early 2000s, DL has greatly improved predictive models by automatically extracting and analyzing valuable insights from raw data [13]. A hybrid model is created for general classification tasks, showing better performance than individual ML or DL models [14].

ML techniques were used to investigate the growth of E. coli O157:H7 in spiked raw ground beef and evaluated prediction performance using coefficient of determination (R²) and mean squared error [15]. A second study employed an impedimetric electrochemical aptasensor with gold interdigitated electrodes to measure E. coli O157:H7 in surface water for hydroponic lettuce irrigation. A ML framework was used to evaluate the effectiveness of existing methods for contamination detection and model performance was assessed using the root mean squared error [16]. Many ML algorithms were used to predict E. coli levels such as stochastic gradient boosting, random forest (RF), support vector machines (SVM), and k-nearest neighbor (KNN) [17]. Predictive models for E. coli were developed using regression techniques, artificial neural networks (ANN), and two adaptive neuro-fuzzy inference system (ANFIS) structures and their performance was compared using R² and the root mean squared log error [18]. RF, Categorical Boosting, and Naïve Bayes (NB) were used to classify surgery-ward E. coli isolates from susceptibility profiles, and performance was evaluated using accuracy, precision, recall, F1-score, and area under the curve (AUC) [19]. RF, logistic regression (LR), and SVM analyzed flow cytometry 2D histograms to detect E. coli associated microbial patterns, evaluated by accuracy, sensitivity, precision, and F1-score [20].

Although previous studies identified risk factors for E. coli contamination using traditional statistical methods [2,21], this study advances beyond risk factor identification by developing predictive models that classify contamination status in settings where routine laboratory surveillance is absent or infeasible. ML algorithms offer two critical advantages over conventional statistical approaches: (1) higher predictive accuracy through detection of complex non-linear relationships and high-order interactions among environmental, socioeconomic, and household-level predictors, and (2) the capacity to prioritize high-risk locations for targeted sampling when universal testing is not financially or logistically viable. To achieve this, we test the ability of classical ML [Adaptive Boosting (AdaBoost), Decision Trees (DT), Gradient Boosting Algorithm (GBA), KNN, Light Gradient-Boosting Machine (LightGBM), LR, NB, RF, and SVM] and the DL Multi-layer Perceptron (DL-MLP) algorithms using data from the 2019 Multiple Indicator Cluster Survey (MICS) to estimate the risk of E. coli contamination (REcC) in the drinking water under two scenarios: (1) an imbalanced dataset assessed using 10-fold stratified cross-validation (CV) without hyperparameter tuning, and (2) Synthetic Minority Over-sampling Technique (SMOTE)-balanced training data assessed using 10-fold stratified CV with hyperparameter tuning. We also identified the relevant predictors. The novelty of this work lies in transitioning from explanatory risk factor analysis to operational prediction tools that enable resource-constrained health systems to allocate limited diagnostic capacity based on model-predicted contamination probability rather than reactive testing. The Bangladeshi government is implementing advanced analytical approaches to predict E. coli contamination risk, aiming to improve water quality and protect public health. The 2019 MICS data can help identify hotspots and vulnerable populations, enhancing policymakers’ understanding of socio-demographic factors influencing contamination risk. This evidence-based approach can inform targeted interventions and resource allocation.

Methods

Data source and study design

The study analyzed secondary dataset from the 2019 MICS in Bangladesh to assess E. coli concentrations in drinking water. A two-stage, stratified cluster sampling method was implemented. Primary sampling units were selected from enumeration areas (EAs) defined by the 2011 Bangladesh Population and Housing Census in the first stage. Subsequently, 20 households (HHs) were selected from each EA via systematic random sample, with four households chosen for arsenic level assessment in their drinking water. Additionally, two HHs were randomly selected to measure E. coli concentrations in both household drinking water and its source. The study targeted a sample size of 6,440 households, with 6,069 (98.7%) successfully providing data for E. coli testing [22].

Outcome variable

In this paper, REcC in the HH drinking water is the response variable. The most recommended indicator for fecal contamination is the number of E. coli bacteria counts in a 100 mL sample. The outcome variable is E. coli contamination in household drinking water, which is categorized using WHO criteria into low (<1 colony forming units [CFU]/100mL), medium (1–10 CFU/100mL), and high risk (11–100 or more CFU/100 mL) categories [23,24]. To this study, medium and high-risk categories were grouped as increased risk category. Specifically, households with 1 or more CFU per 100 mL of drinking water were categorized as ‘Yes’ for the REcC, while those with less than 1 CFU per 100 mL were categorized as ‘No’ for the REcC. Mathematically, the outcome variable can be defined as:

Independent variables

In addition to the outcome variable, division, place of residence, age of household head, household head education, household size, wealth index, livestock ownership, types of toilet facility, sources of drinking water, location of the water source, treatment of drinking water, place of handwashing, mass media exposure, and women functional difficulties were selected as potential factors influencing REcC among drinking water in households (Table 1).

Statistical analysis

We first used a traditional approach based on Chi-squares tests to explore associations between REcC and explanatory variables. Agglomerative hierarchical clustering using Ward’s linkage was used to identify district-level hotspots for the REcC in drinking water. This method automatically groups data points into clusters based on their similarities, starting with each data point in its own cluster and merging the closest two until a single cluster remains.

XGBoost SHapley additive explanations (SHAP)

Prior to the ML analysis, a predictive model for imbalanced classification was developed using eXtreme Gradient Boosting (XGBoost), with the goal of reducing overfitting and improving accuracy through high speed [25,26]. We then used SHapley Additive exPlanations (SHAP) to interpret XGBoost model predictions, providing a comprehensive understanding of each feature’s contribution to the model’s predictions [27]. The SHAP summary plot was used to rank predictors by their mean absolute SHAP values (average absolute impact on the XGBoost output). We then dropped predictors with near-zero SHAP contribution and refit the model using the remaining variables to reduce noise and overfitting. Next, we reviewed SHAP dependence plots to identify clear non-linear patterns and strong interactions among key predictors. Based on these patterns, we updated feature coding where needed (for example, collapsing sparse categories and using more appropriate transformations) and refit the final model. Model performance was re-assessed using the same validation procedure to confirm that these SHAP-guided changes improved generalization and produced more stable predictions.

Additionally, Cramer’s V correlation was used to identify potential collinearity among features, ranging from 0 (no association) to 1 (perfect association) [28]. Values between 0.36 and 0.49 indicated a substantial correlation, while values of 0.50 or higher were considered strong correlation [29].

ML classifiers

This study employed a comprehensive approach, integrating nine (09) classical ML models and one (01) DL techniques, to predict the REcC in household drinking water in Bangladesh. In the training dataset, we applied nine classical ML models utilized in this study include AdaBoost, DT, GBA, KNN, LightGBM, LR, NB, RF, and SVM. Additionally, we incorporated DL-MLP model into our analysis. ML models and evaluation metrics are extensively utilized in the literature for classification tasks [25,3037].

Data preparation and model evaluation

For the ML approach, the MICS dataset was split into training (70%) and test (30%) sets using stratification to preserve the REcC class distribution. The dataset was checked for missing values, which were handled using KNN imputation [38] before training the selected ML and DL algorithms was implemented within a scikit-learn Pipeline so that the same steps were applied consistently during training and validation. Data preprocessing was performed using Python [35], Scikit-learn [39], pandas [40] and Keras [41]. Imbalanced data refers to unequal representation of classes in classification problems, including binary and multi-class tasks [42]. The training set was subjected to 10-fold CV to identify optimal hyperparameters. We used the test dataset to evaluate the performance of models measured by various evaluation parameters, including accuracy, precision, recall, F1 score, and AUC [33,34,4346]. Results were presented in both tabulated and graphical formats, with violin plots [33] being used to illustrate the distribution of model performances. The overall procedure of the ML approach is displayed in Fig 1.

thumbnail
Fig 1. Flow chart of the procedure of the nine ML classifiers and one DL classifier.

https://doi.org/10.1371/journal.pone.0343606.g001

Classification outputs

After training, each household sample was classified using two model outputs. Predicted class labels were obtained using predict(X) and used to generate the confusion matrix and compute accuracy, precision, recall, F1-score, and Cohen’s kappa. When available, predicted probabilities were obtained using predict_proba(X)[:,1]; for SVM, probability estimation was enabled (probability = True) to obtain probability outputs. Receiver operating characteristic (ROC)-AUC was computed using predicted probabilities, whereas the other metrics were computed using predicted class labels.

Hyperparameter

Hyperparameters control the learning process and help determine the best model parameters [47]. A ML model’s performance depends on optimal hyperparameters, which can be tuned using methods like grid search, random search, and Bayesian optimization [48]. Grid search divides the hyperparameter space into a grid and evaluates all combinations using CV to find the optimal values [49]. We used grid search with CV to optimize each model’s hyperparameters. We optimized model hyperparameters using GridSearchCV applied to the training set only. For each algorithm, we evaluated the parameter grids listed in S1 Table in S1 File. Model selection was performed using stratified 10-fold CV (StratifiedKFold, n = 10, shuffle = True, random_state = 42) and the mean F1-score was used as the tuning criterion. The best model (best_estimator_) was refitted on the full training set and evaluated once on the held-out test set.

SMOTE

Class imbalance arises when one class has far more samples than another, causing models to overfit the majority class and perform poorly on the minority class [5052]. SMOTE generates synthetic minority-class samples to balance the dataset, reducing majority-class bias and helping the classifier learn the minority class more effectively [5355]. Since the outcome variable (REcC in drinking water) was imbalanced (39.2% “Yes” and 60.8% “No”), we used the SMOTE to address class imbalance in this binary outcome [56]. After SMOTE, synthetic samples are generated for the minority class so that the class distribution becomes more balanced, with roughly equal numbers of instances in each class [57]. Several studies have shown that SMOTE usually boosts minority-class precision, recall, and F1-score, giving a more balanced and fair assessment of model performance [52,57]. For our analysis, we evaluated two training scenarios for each model: (i) the original imbalanced training data, assessed using 10-fold stratified CV without hyperparameter tuning, and (ii) SMOTE-balanced training data, assessed using 10-fold stratified CV with hyperparameter tuning. When SMOTE was used, it was applied only to the training split within each CV fold; the validation folds and the final held-out test set remained non-synthetic.

Results

Prevalence of E. coli in household drinking water

The overall prevalence of E. coli contamination in drinking water was 39.2% (95% CI: 37.4 to 41.2).

Association between sociodemographic factors and E. coli contamination

In univariate analysis, the REcC in drinking water in Bangladesh was significantly associated with several factors, including division, residence, household head education, household size, wealth index, livestock ownership, types of toilet facilities, sources of drinking water, treatment methods for drinking water, handwashing locations, mass media exposure, and women functional difficulties (p < 0.05) (Table 2).

thumbnail
Table 2. Characteristics of enrolled households according to REcC in drinking water.

https://doi.org/10.1371/journal.pone.0343606.t002

The highest percentage of positive households was observed in Dhaka division (51.4%), followed by Chattogram division (49.7%), Mymensingh division (43.4%), Khulna and Sylhet divisions (both at 37.1%), Rajshahi division (27.5%), Rangpur division (23.4%), and Barishal division, which had the lowest risk (15.8%). The proportion of positive households was higher in urban areas (46.9%) than rural areas (37%). There was a significant association between education level of household head and REcC, with 42.3% of households in Bangladesh having heads with no formal education. Approximately 42.2% of households with more than six members were positive for E. coli contamination in their drinking water. With respect to social-economic characteristics, most houses positive for E. coli contamination (42.6%) were of high-income families. Livestock ownership and toilet facilities were significant factors for E. coli contamination risk among households in Bangladesh. Households with flush toilet facilities had increased prevalence of E. coli contamination compared to those with other facilities. A larger proportion (58%) of households that treated their drinking water were at increased REcC. Additionally, a higher percentage of households where hands were washed inside the dwelling (45.1%) were at REcC, followed by those washing hands in the yard/plot (38%), on a mobile object (39.7%), and those with no designated place for hand washing (36.6%). Households with exposure to mass media had a higher REcC (40.7%) compared to those without exposure to mass media (37.3%) (Table 2).

District-wise REcC

Certain districts in Bangladesh had higher REcC compared to others. We detected geographical clustering in contamination results using agglomerative hierarchical clustering, as shown in the dendrogram in S1 Fig in S1 File. The clustering was based on data from all sixty-four districts affected by E. coli in drinking water. The analysis resulted in four distinct clusters, with sizes of 10, 15, 12, and 27 districts, respectively, which were sorted according to the level of contamination of drinking water. Cluster C1 contained districts with relatively high REcC (up to 60%), whereas districts in cluster C4 had low REcC (around 20%). Cluster 1 (highest prevalence of contamination) comprises Bandarban, Dhaka, Habigonj, Narsingdi, Feni, Tangail, Rangamati, Rajbari, Lakshmipur, and Netrokona. Cluster 2 includes Khagrachari, Kushtia, Khulna, Madaripur, Pirojpur, Noakhali, Chattogram, Satkhira, Lalmonirhat, Meherpur, Cumilla, Kurigram, Cox’s Bazar, Bagerhat, and Sirajganj. Cluster 3 covers Jessore, Joypurhat, Rangpur, Thakurgaon, Naogaon, Panchagarh, Bhola, Patuakhali, Barisal, Barguna, Bogura, and Gaibandha. Cluster 4 consists of all remaining districts not categorized in the first three clusters (Fig 2).

thumbnail
Fig 2. District-wise map of the proportion of E. coli contamination in household drinking water across Bangladesh, generated using agglomerative hierarchical clustering with Ward linkage.

Administrative district boundaries were obtained from the GADM database (Bangladesh ADM2) and the map was generated by the authors in R using sf and ggplot2 without any proprietary basemap tiles. Data: MICS 2019, Bangladesh. Map: GADM (ADM2).

https://doi.org/10.1371/journal.pone.0343606.g002

Fig 2 shows the district level variation of REcC in the drinking water using agglomerative hierarchical cluster approach. High REcC was observed in Bandarban, Dhaka, Habigonj, Narshingdi, Feni, Rangamati, Tangail, Rajbari, Lakshmipur, and Netrokona districts, which incorporated the first cluster with > 58% of households being positive for E. coli.

Feature selection

Fig 3 displays SHAP values and top fourteen features with highest impact, with yellow indicating high and dark blue indicating low, with horizontal axis representing each feature’s contribution to prediction. Positive SHAP values indicate higher contamination risk, while negative values indicate lower risk; for example, higher division differences were associated with increased contamination. Regarding SHAP, geographical division was found to be the most significant predictor, followed by the household size, wealth index, household head’s education level, place of handwashing, and livestock ownership.

thumbnail
Fig 3. Feature importance based on SHAP values.

The beeswarm summary plot displays the distribution of SHAP values for the top predictors of REcC in household drinking water, with features ordered by their overall contribution to the model.

https://doi.org/10.1371/journal.pone.0343606.g003

Cramér’s V correlation

Cramér’s V correlation analysis revealed strong associations between various factors, particularly the place of handwashing and the location of the water source. Consequently, we excluded the place of handwashing from further ML modeling to predict E. coli levels in water (S2 Fig in S1 File).

The distribution of drinking-water sources across divisions is presented in S3 Fig in S1 File. There was a significant association between division and drinking-water source (χ² = 497.52, p < 0.001). Tubewell water was the predominant source in every division, whereas the use of piped water varied by region. Piped water use was highest in Dhaka (20.5%) and substantially lower in the other divisions (2.5%−9.1%). This geographic variation in water-source profiles is consistent with the SHAP results, which identified both division and drinking-water source as important predictors of E. coli contamination risk.

REcC in the drinking water using ML models

For the ML models, we tested nine algorithms RF, LR, DT, GBA, AdaBoost, LightGBM, KNN, NB, and SVM and as well as a DL approach based on an DL-MLP. Model performance after applying SMOTE and CV with hyperparameter tuning is presented in Fig 4. The corresponding results for imbalanced dataset including 10-fold CV (before SMOTE and before hyperparameter tuning) are provided in the Supplementary Materials (S4 and S5 Figs in S1 File). The predictive performance of each algorithm was compared based on accuracy, Cohen’s Kappa (κ), precision, recall, and F1-score. In both scenarios, scenario 2 (SMOTE-balanced training with CV and hyperparameter tuning) performed better than scenario 1 (original imbalanced training data).

thumbnail
Fig 4. Classification performance measure of the algorithms and comparison.

Performance indicators (accuracy, Cohen’s kappa coefficient, precision, recall, and F1 score) for ML algorithms (SVM, RF, NB, LR, LightGBM, KNN, GBA, DT, AdaBoost and DL-MLP).

https://doi.org/10.1371/journal.pone.0343606.g004

For scenario 2 (after SMOTE), AdaBoost demonstrated the highest accuracy at 81.6%, correctly predicting the REcC in household drinking water (Fig 4). It also achieved the highest Cohen’s Kappa (κ = 0.194), recall (0.990), and F1-score (0.898). GBA and LightGBM performed similarly well, with accuracies of 0.799 and 0.791, respectively, and strong balance between sensitivity and precision (GBA: recall 0.965, F1 0.887; LightGBM: recall 0.943, F1 0.881). RF and DT showed moderate-to-strong performance (RF: accuracy 0.759, F1 0.858; DT: accuracy 0.739, F1 0.842), while KNN achieved reasonable accuracy (0.717) and F1 (0.830) but had the lowest κ (0.077), indicating weaker improvement over chance agreement.

In contrast, LR and NB had lower accuracies (0.556 and 0.578) and lower F1-scores (0.676 and 0.696), despite relatively high precision (both ≥ 0.841). SVM also showed low accuracy (0.587) and a comparatively low F1-score (0.707). The DL-MLP model showed intermediate performance (accuracy 0.625, recall 0.682, F1 0.749). Among the ten classifiers, the AdaBoost ML classifier slightly outperformed the others (Fig 4 and S6 Fig in S1 File).

To predict the REcC among households in Bangladesh, the AUC measure was estimated for the LR, RF, GBA, DT, DL-MLP, KNN, AdaBoost, LightGBM, SVM and NB model models (Fig 5). The estimated AUC values were 0.591 (LR), 0.634 (NB), 0.670 (DT), 0.678 (RF), 0.651 (SVM), 0.619 (KNN), 0.686 (AdaBoost), 0.685 (GBA), 0.676 (LightGBM), and 0.626 (DL-MLP). Overall, the boosting and tree-based models showed the best discrimination for predicting REcC in household drinking water. AdaBoost achieved the highest AUC (0.686), followed closely by GBA (0.685), RF (0.678), and LightGBM (0.676). In contrast, LR had the lowest AUC (0.591), indicating the weakest discriminative performance among the evaluated models.

thumbnail
Fig 5. ROC curves and AUC for ten classifiers (AdaBoost, DL-MLP, DT, GBA, KNN, LightGBM, LR, NB, RF, and SVM) using 10-fold CV to predict the REcC in household drinking water.

https://doi.org/10.1371/journal.pone.0343606.g005

AdaBoost model achieved the highest and most consistent median accuracy across folds, followed by GBA, LightGBM, and RF. Unlike a box plot, the violin plot in Fig 6 displays the full distribution of the 10-fold CV accuracy results.

thumbnail
Fig 6. Violin plots showing the distribution of accuracy across 10-fold CV after SMOTE and hyperparameter tuning for ten classifiers (AdaBoost, DL-MLP, DT, GBA, KNN, LightGBM, LR, NB, RF, and SVM) used to predict the REcC in household drinking water.

The central point indicates the median and the inner band shows the interquartile range.

https://doi.org/10.1371/journal.pone.0343606.g006

Feature importance

Fig 7 illustrates the feature importance based on the AdaBoost model, providing valuable insights into the factors influencing E. coli levels in drinking water in Bangladesh. The analysis identifies five key variables division, wealth index, types of toilet facility, household size, and household head education as significant predictors of E. coli contamination in household water.

Discussion

This cross-sectional study used 2019 MICS dataset and main purpose of this study is the assess the performance of ML algorithms to detect REcC in drinking water and identified the hotspot of highly contaminated district of risk of E. coli using the cluster dendrogram of the hierarchical clustering based on Wald’s criterion. To the best of our knowledge, there was no available study which have conducted in Bangladesh to focused on REcC based on the different ML classifiers and clustering approaches we have used in this study. A comparison table between our results and previous research is provided in S2 Table in S1 File.

This study also evaluated that prevalence of REcC in drinking water in Bangladesh was 39.2%. About 39% and 65% of pathogens were detected in point-of-drinking water samples and public domain water source samples, respectively, and classified as low-risk using PCR testing [58]. In a similar study researcher was done using API and conventional methods and concluded that the prevalence of E. coli from different water sources was 49.48% [59]. In our study REcC in drinking water in Bangladesh was significantly associated with several factors, including division, residence, household head education, household size, wealth index, livestock ownership, types of toilet facilities, sources of drinking water, treatment methods for drinking water [8], handwashing location, mass media exposure and maternal functional difficulties. Moreover, a univariate and multivariate analysis conducted in a study [4] which was found that demographic factors such as household heads, room occupancy rate and wealth status was significantly correlated with the occurrence of E. coli in stored drinking water [60].

We evaluated the performance of nine ML algorithms and one DL algorithm to predict REcC in drinking water across two scenarios: original (imbalanced) data and balanced data (SMOTE, with CV and hyperparameter tuning). The ML models included RF, LR, DT, GBA, AdaBoost, LightGBM, KNN, NB, and SVM; DL-MLP was used for DL. Across all ML and DL models, Scenario 2 (balanced data) outperformed Scenario 1 (imbalanced data). For Scenario 2, AdaBoost demonstrated the highest accuracy (81.6%) for correctly predicting REcC in drinking water, along with the highest recall (0.990) and F1-score (0.898). AdaBoost is often highly effective for classification because it combines weak learners to refine decision boundaries, capture feature interactions, and preserve good generalization, which frequently leads to superior accuracy and F1 scores compared with many other models [61,62]. From a practical perspective, our results suggest that boosting and other tree-based ensemble models (AdaBoost, GBA, and LightGBM) are strong first-choice methods for similar survey-based contamination studies, where predictors are largely categorical and relationships may be non-linear. In this dataset, LR showed weaker discrimination (lower AUC) than the ensemble approaches, which may reflect the importance of non-linear effects and interactions across household, environmental, and geographic factors.

Additionally, in terms of estimated AUC values, both the AdaBoost and GBA models outperformed LR, RF, GBA, DT, KNN, LightGBM, SVM, NB and DL-MLP models in predicting the REcC among households in Bangladesh. AdaBoost is versatile because it can use different types of weak learners (for example, decision stumps, SVMs, or neural networks) and be tuned to specific problems, which often improves its performance [63,64]. In the Milwaukee River system, physicochemical water-quality variables explained only a limited proportion of the variation in E. coli (R² = 0.29–0.42). The hybrid ANN–GBM model performed best, but prediction accuracy remained low (about 42%) [65]. RF achieved the highest R² value (0.933) when combining water quality and RGB data in the USA [66], as well as in predicting E. coli concentrations in the Marne River (Paris Area, France) [67]. KNN was the most accurate model for predicting E. coli concentrations in Midmar Dam, with XGBoost and SVM performing next best, while RF and ANN had the largest errors [68].

Moreover, RF has been more effective in predicting E. coli concentrations in agricultural pond waters, as it provided the smallest RMSE in almost all cases [17]. Furthermore, both RF and TPOT exhibited the best performance in predicting E. coli concentrations in the Göta älv river in Gothenburg, Sweden [69], and in Ethiopia [70]. On the other hand, SVM outperforms LR in 26 cities with an accuracy of 70%, except in Kathmandu, where it achieves 79%. In contrast, LR demonstrates a significantly lower accuracy of 61% across all cities [71].

For feature importance in the AdaBoost model, division was the strongest predictor of REcC in drinking water, followed by wealth index, types of toilet facility, household size, household head education, and source of drinking water. Recent study showed that climate, land use, and population density were key predictors of E. coli contamination in shallow tubewell water in Bangladesh [72]. A recent study using a Bayesian censored model on a similar dataset found that household E. coli levels were significantly associated with division, water source/location, treatment, handwashing, toilet type, education and wealth, and livestock ownership [73].

Despite all efforts to optimize ML algorithms even after class balancing for predicting REcC using the 2019 MICS dataset, models achieved a maximum accuracy of 81.6% but only a modest AUC of 0.686. This discrepancy reveals ML’s limited discriminatory ability to reliably predict contamination in household drinking water, even at moderate prevalence levels.

Furthermore, a study conducted in Ireland explored spatial mapping of E. coli occurrence in the community [74] to examine local geographic patterns of E. coli contamination. This study identified hotspot areas with high levels of contamination by using a cluster dendrogram from hierarchical clustering based on Wald’s criterion. In our study, we concluded that the highest REcC in household drinking water was observed in districts such as Bandarban, Dhaka, Habiganj, Narsingdi, Feni, Rangamati, Tangail, Rajbari, Lakshmipur, and Netrokona. These districts were marked as the first cluster, where more than 58% of cases were identified, with red spots indicating the highest levels of contamination. In contrast, districts represented in sky blue were grouped into the fourth cluster, which was characterized as low risk for E. coli contamination in household drinking water across Bangladesh.

This study has limitations. In terms of limitations, it is a cross-sectional study, so no cause-and-effect relationships can be established. Additionally, the analysis relied on socio-demographic variables available in the MICS dataset, excluding environmental factors such as distance from sewage treatment plants and population density because these variables were absent from the dataset and could impact contamination risks.

Conclusion

In this study, we applied nine ML classifiers and one DL-MLP classifier to predict the REcC in household drinking water in Bangladesh using the 2019 MICS dataset. We also used agglomerative hierarchical clustering with Ward’s linkage to identify district-level hotspots. More than one-third of households were classified as having increased REcC in their drinking water. We compared model performance using accuracy, precision, recall, F1-score, Cohen’s kappa, ROC-AUC, and violin plots to summarize the distribution of performance across models. Across all models, SMOTE combined with CV and hyperparameter tuning substantially improved performance compared to the original imbalanced data. Based on ROC-AUC, the models showed limited discrimination between contaminated and non-contaminated household drinking-water samples, even though the best-performing model achieved 81.6% accuracy. Feature-importance results indicated that division, wealth index, type of toilet facility, household size, education, and source of drinking water were the most influential predictors of REcC. Future work should focus on improving predictive performance by incorporating additional environmental and spatial exposure variables where possible and by validating models on independent data.

Supporting information

S1 File. Supporting information figures and tables (S1-S6 Figs; S1-S2 Tables).

https://doi.org/10.1371/journal.pone.0343606.s001

(DOCX)

References

  1. 1. Allocati N, Masulli M, Alexeyev MF, Di Ilio C. Escherichia coli in Europe: An overview. Int J Environ Res Public Health. 2013;10(12):6235–54. pmid:24287850
  2. 2. Khan JR, Bakar KS. Spatial risk distribution and determinants of E. coli contamination in household drinking water: A case study of Bangladesh. Int J Environ Health Res. 2020;30(3):268–83. pmid:30924350
  3. 3. White K, Dickson-Anderson S, Majury A, McDermott K, Hynds P, Brown RS, et al. Exploration of E. coli contamination drivers in private drinking water wells: An application of machine learning to a large, multivariable, geo-spatio-temporal dataset. Water Res. 2021;197:117089. pmid:33836295
  4. 4. Fong T-T, Mansfield LS, Wilson DL, Schwab DJ, Molloy SL, Rose JB. Massive microbiological groundwater contamination associated with a waterborne outbreak in Lake Erie, South Bass Island, Ohio. Environ Health Perspect. 2007;115(6):856–64. pmid:17589591
  5. 5. Hynds PD, Thomas MK, Pintar KDM. Contamination of groundwater systems in the US and Canada by enteric pathogens, 1990–2013: A review and pooled-analysis. PLoS One. 2014;9(5):e93301.
  6. 6. O’Dwyer J, Hynds PD, Byrne KA, Ryan MP, Adley CC. Development of a hierarchical model for predicting microbiological contamination of private groundwater supplies in a geologically heterogeneous region. Environ Pollut. 2018;237:329–38. pmid:29499576
  7. 7. Roberts L, Chartier Y, Chartier O, Malenga G, Toole M, Rodka H. Keeping clean water clean in a Malawi refugee camp: A randomized intervention trial. Bull World Health Organ. 2001;79(4):280–7. pmid:11357205
  8. 8. Dada N, Vannavong N, Seidu R, Lenhart A, Stenström TA, Chareonviriyaphap T, et al. Relationship between Aedes aegypti production and occurrence of Escherichia coli in domestic water storage containers in rural and sub-urban villages in Thailand and Laos. Acta Trop. 2013;126(3):177–85. pmid:23499713
  9. 9. Bain R, Cronk R, Wright J, Yang H, Slaymaker T, Bartram J. Fecal contamination of drinking-water in low- and middle-income countries: A systematic review and meta-analysis. PLoS Med. 2014;11(5):e1001644. pmid:24800926
  10. 10. Holder LB, Haque MM, Skinner MK. Machine learning for epigenetics and future medical applications. Epigenetics. 2017;12(7):505–14. pmid:28524769
  11. 11. Mavaie P, Holder L, Beck D, Skinner MK. Predicting environmentally responsive transgenerational differential DNA methylated regions (epimutations) in the genome using a hybrid deep-machine learning approach. BMC Bioinformatics. 2021;22(1):575. pmid:34847877
  12. 12. Deng J, Dong W, Socher R, Li L-J, Kai Li, Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009. 248–55. https://doi.org/10.1109/cvpr.2009.5206848
  13. 13. Mamoshina P, Vieira A, Lane E, Zhavoronkov A. Applications of deep learning in biomedicine. Mol Pharm. 2016;13(5):1445–54. pmid:27007977
  14. 14. Mavaie P, Holder L, Skinner MK. Hybrid deep learning approach to improve classification of low-volume high-dimensional data. BMC Bioinformatics. 2023;24(1):419. pmid:37936066
  15. 15. Al S, Uysal Ciloglu F, Akcay A, Koluman A. Machine learning models for prediction of Escherichia coli O157:H7 growth in raw ground beef at different storage temperatures. Meat Sci. 2024;210:109421. pmid:38237258
  16. 16. Qian H, McLamore E, Bliznyuk N. Machine learning for improved detection of pathogenic E. coli in hydroponic irrigation water using impedimetric aptasensors: A comparative study. ACS Omega. 2023;8(37):34171–9. pmid:37744804
  17. 17. Stocker MD, Pachepsky YA, Hill RL. Prediction of E. coli concentrations in agricultural pond waters: Application and comparison of machine learning algorithms. Front Artif Intell. 2022;4:768650. pmid:35088045
  18. 18. Abimbola OP, Mittelstet AR, Messer TL, Berry ED, Bartelt-Hunt SL, Hansen SP. Predicting Escherichia coli loads in cascading dams with machine learning: An integration of hydrometeorology, animal density and grazing pattern. Sci Total Environ. 2020;722:137894. pmid:32208262
  19. 19. Tolan HK, Aydın İ, Tanyildizi-Kokkulunk H, Karakuş M, Akkaya Y, Kaya O, et al. Machine learning model for predicting multidrug resistance in clinical Escherichia coli isolates: A retrospective general surgery study. Antibiotics (Basel). 2025;14(10):969. pmid:41148661
  20. 20. Erb IK, Gador N, Jinbäck M, Lindberg E, Paul CJ. A data-driven early warning system for Escherichia coli in water based on microbial community analysis using flow cytometry 2D histograms. Water Research X. 2025;29:100404.
  21. 21. Hasan MM, Hoque Z, Kabir E, Hossain S. Differences in levels of E. coli contamination of point of use drinking water in Bangladesh. PLoS One. 2022;17(5):e0267386. pmid:35544525
  22. 22. Bangladesh Bureau of Statistics BBS, UNICEF Bangladesh. Progotir Pathey Bangladesh: Multiple Indicator Cluster Survey 2019. Dhaka (Bangladesh): Bangladesh Bureau of Statistics (BBS). 2019.
  23. 23. World Health Organization. Guidelines for drinking-water quality. Geneva: World Health Organization. 2011.
  24. 24. Stauber C, Miller C, Cantrell B, Kroell K. Evaluation of the compartment bag test for the detection of Escherichia coli in water. J Microbiol Methods. 2014;99:66–70. pmid:24566129
  25. 25. Chen T, Guestrin C. XGBoost: a scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD’16), 2016. 785–94. https://doi.org/10.1145/2939672.2939785
  26. 26. Zhang C, Jia D, Wang L, Wang W, Liu F, Yang A. Comparative research on network intrusion detection methods based on machine learning. Comput Secur. 2022;121:102861.
  27. 27. Mangalathu S, Hwang S-H, Jeon J-S. Failure mode and effects analysis of RC members based on machine-learning-based SHapley Additive exPlanations (SHAP) approach. Engineering Structures. 2020;219:110927.
  28. 28. Zelko E, Švab I, Rotar Pavlič D. Quality of life and patient satisfaction with family practice care in a roma population with chronic conditions in Northeast Slovenia. Zdr Varst. 2014;54(1):18–26. pmid:27646618
  29. 29. Fan C, Oh DS, Wessels L, Weigelt B, Nuyten DSA, Nobel AB, et al. Concordance among gene-expression-based predictors for breast cancer. N Engl J Med. 2006;355(6):560–9. pmid:16899776
  30. 30. Breiman L. Random forests. Machine Learning. 2001;45(1):5–32.
  31. 31. Gupta V, Mishra VK, Singhal P, Kumar A. An Overview of Supervised Machine Learning Algorithm. In: 2022 11th International Conference on System Modeling & Advancement in Research Trends (SMART), 2022. 87–92. https://doi.org/10.1109/smart55829.2022.10047618
  32. 32. Jiang T, Gradus JL, Rosellini AJ. Supervised machine learning: A brief primer. Behav Ther. 2020;51(5):675–87. pmid:32800297
  33. 33. Nayan MdIH, Uddin MSG, Hossain MdI, Alam MdM, Zinnia MA, Haq I, et al. Comparison of the Performance of machine learning-based algorithms for predicting depression and anxiety among university students in Bangladesh. Asian Journal of Social Health and Behavior. 2022;5(2):75–84.
  34. 34. Rahman R, Talukder A, Das S, Saha J, Sarma H. Understanding and predicting pregnancy termination in Bangladesh: A comprehensive analysis using a hybrid machine learning approach. Medicine (Baltimore). 2024;103(26):e38709. pmid:38941421
  35. 35. Raschka S, Patterson J, Nolet C. Machine learning in python: main developments and technology trends in data science, machine learning, and artificial intelligence. Information. 2020;11(4):193.
  36. 36. Sarker IH. Machine learning: Algorithms, real-world applications and research directions. SN Comput Sci. 2021;2(3):160. pmid:33778771
  37. 37. Wang R, Kim J-H, Li M-H. Predicting stream water quality under different urban development pattern scenarios with an interpretable machine learning approach. Sci Total Environ. 2021;761:144057. pmid:33373848
  38. 38. Rizvi STH, Latif MY, Amin MS, Telmoudi AJ, Shah NA. Analysis of machine learning based imputation of missing data. Cybern Syst. 2023;54(8):1–15.
  39. 39. Abraham A, Pedregosa F, Eickenberg M, Gervais P, Mueller A, Kossaifi J, et al. Machine learning for neuroimaging with scikit-learn. Front Neuroinform. 2014;8:14. pmid:24600388
  40. 40. Gupta P, Bagchi A. Introduction to pandas. Essentials of Python for artificial intelligence and machine learning. Cham (Switzerland): Springer Nature Switzerland. 2024. 161–96.
  41. 41. Schneider P, Xhafa F. Anomaly detection and complex event processing over IoT data streams: with application to eHealth and patient data monitoring. Cambridge (MA): Academic Press. 2022.
  42. 42. Han Q, Gui C, Xu J, Lacidogna G. A generalized method to predict the compressive strength of high-performance concrete by improved random forest algorithm. Constr Build Mater. 2019;226:734–42.
  43. 43. Hicks SA, Strümke I, Thambawita V, Hammou M, Riegler MA, Halvorsen P, et al. On evaluation metrics for medical applications of artificial intelligence. Sci Rep. 2022;12(1):5979. pmid:35395867
  44. 44. Khan MU, Aziz S, Iqtidar K, Zaher GF, Alghamdi S, Gull M. A two-stage classification model integrating feature fusion for coronary artery disease detection and classification. Multimed Tools Appl. 2021;81(10):13661–90.
  45. 45. Roth GA, Mensah GA, Johnson CO, Addolorato G, Ammirati E, Baddour LM, et al. Global burden of cardiovascular diseases and risk factors, 1990–2019. J Am Coll Cardiol. 2020;76(25):2982–3021.
  46. 46. Rainio O, Teuho J, Klén R. Evaluation metrics and statistical tests for machine learning. Sci Rep. 2024;14(1):6086. pmid:38480847
  47. 47. Pravin PS, Tan JZM, Yap KS, Wu Z. Hyperparameter optimization strategies for machine learning-based stochastic energy efficient scheduling in cyber-physical production systems. Digit Chem Eng. 2022;4:100047.
  48. 48. Yasir M, Karim AM, Malik SK, Bajaffer AA, Azhar EI. Prediction of antimicrobial minimal inhibitory concentrations for Neisseria gonorrhoeae using machine learning models. Saudi J Biol Sci. 2022;29(5):3687–93. pmid:35844400
  49. 49. Wu Z, Tran A, Rincon D, Christofides PD. Machine learning‐based predictive control of nonlinear processes. Part I: Theory. AIChE Journal. 2019;65(11).
  50. 50. Niaz NU, Shahariar KMN, Patwary MJA. Class Imbalance Problems in Machine Learning: A Review of Methods And Future Challenges. In: Proceedings of the 2nd International Conference on Computing Advancements, 2022. 485–90. https://doi.org/10.1145/3542954.3543024
  51. 51. Brishti F, Zhang F, Mohammed S, Bai L, Wu F, Chen B. Imbalanced classification with label noise: A systematic review and comparative analysis. ICT Express. 2025;11(6):1127–45.
  52. 52. Wibowo P, Fatichah C. An in-depth performance analysis of the oversampling techniques for high-class imbalanced dataset. regist j ilm teknol sist inf. 2021;7(1):63.
  53. 53. Heroza RI, Gan JQ, Raza H. Sia-SMOTE: a SMOTE-based oversampling method with better interpolation on high-dimensional data by using a Siamese network. In: International Work-Conference on Artificial Neural Networks (IWANN). Cham (Switzerland): Springer Nature Switzerland; 2023. 448–460.
  54. 54. Hairani H, Widiyaningtyas T, Dwi Prasetya D. Addressing class imbalance of health data: A systematic literature review on modified synthetic minority oversampling technique (SMOTE) Strategies. JOIV : Int J Inform Visualization. 2024;8(3):1310.
  55. 55. Fernández A, Garcia S, Herrera F, Chawla NV. SMOTE for learning from imbalanced data: progress and challenges, marking the 15-year anniversary. J Artif Intell Res. 2018;61:863–905.
  56. 56. Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP. SMOTE: Synthetic Minority Over-sampling Technique. J Artif Intell Res. 2002;16:321–57.
  57. 57. Coşkuner A, Rençber ÖF. The imbalanced data problem: investigating factors affecting financial freedom using data mining techniques with SMOTE method. Machine learning in finance: trends, developments and business practices in the financial sector. Cham (Switzerland): Springer Nature Switzerland. 2025. 87–100.
  58. 58. Saima S, Ferdous J, Sultana R, Rashid RB, Almeida S, Begum A, et al. Detecting Enteric Pathogens in Low-Risk Drinking Water in Dhaka, Bangladesh: An Assessment of the WHO Water Safety Categories. Trop Med Infect Dis. 2023;8(6):321. pmid:37368739
  59. 59. Odonkor ST, Addo KK. Prevalence of multidrug-resistant escherichia coli isolated from drinking water sources. Int J Microbiol. 2018;2018:7204013. pmid:30210545
  60. 60. Vannavong N, Overgaard HJ, Chareonviriyaphap T, Dada N, Rangsin R, Sibounhom A, et al. Assessing factors of E. coli contamination of household drinking water in suburban and rural Laos and Thailand. Water Supply. 2017;18(3):886–900.
  61. 61. Martinez W, Gray JB. Noise peeling methods to improve boosting algorithms. Computational Statistics & Data Analysis. 2016;93:483–97.
  62. 62. Maji K, Gupta S, Dutta PK. Enhancing heart disease prediction accuracy: Comprehensive analysis of XGBoost and AdaBoost. IET Conf Proc. 2025;2024(37):130–6.
  63. 63. Nabi RM, Saeed SAB, Haron H. Enhanced AdaBoostM1 with multilayer perceptron for stock price prediction. KJAR. 2023;8(1):73–85.
  64. 64. Gai F, Li Z, Jiang X, Guo H. In: 2016. 27–37.
  65. 65. Nafsin N, Li J. Prediction of total organic carbon and E. coli in rivers within the Milwaukee River basin using machine learning methods. Environ Sci: Adv. 2023;2:278–93.
  66. 66. Hong SM, Morgan BJ, Stocker MD, Smith JE, Kim MS, Cho KH, et al. Using machine learning models to estimate Escherichia coli concentration in an irrigation pond from water quality and drone-based RGB imagery data. Water Res. 2024;260:121861. pmid:38875854
  67. 67. Naloufi M, Lucas FS, Souihi S, Servais P, Janne A, Wanderley Matos De Abreu T. Evaluating the performance of machine learning approaches to predict the microbial quality of surface waters and to optimize the sampling effort. Water. 2021;13(18):2457.
  68. 68. Ibrahim AAMS, Nkonyane M, Ngcobo M, Walingo T, Tapamo J-R. Data-Driven machine learning models for E. coli concentration prediction. Sustainability. 2025;18(1):179.
  69. 69. Sokolova E, Ivarsson O, Lillieström A, Speicher NK, Rydberg H, Bondelind M. Data-driven models for predicting microbial water quality in the drinking water source using E. coli monitoring and hydrometeorological data. Sci Total Environ. 2022;802:149798. pmid:34454142
  70. 70. Ambel AA, Bain R, Degefu TB, Donmez A, Johnston R, Slaymaker T. Addressing gaps in data on drinking water quality through data integration and machine learning: Evidence from Ethiopia. npj Clean Water. 2023;6(1).
  71. 71. Kuroki S, Ogata R, Sakamoto M. Predicting the presence of E. coli in tap water using machine learning in Nepal. Water Environ J. 2023;37(3):402–11.
  72. 72. Wu J, Cao Y, Islam MdS, Emch M. Application of machine learning to identify influential factors for fecal contamination of shallow groundwater. Water. 2025;17(2):160.
  73. 73. Haq I, Rahman A, Akter MR, Hossain D, Nobrega D. Bayesian modeling of Escherichia coli contamination in household drinking water in Bangladesh: evidence from the Multiple Indicator Cluster Survey 2019. Int Health. 2025;ihaf138.
  74. 74. Samarasundera E, Walsh T, Cheng T, Koenig A, Jattansingh K, Dawe A, et al. Methods and tools for geographical mapping and analysis in primary health care. Prim Health Care Res Dev. 2012;13(1):10–21. pmid:22024314