Fig 1.
Flowchart of the proposed methodology.
Fig 2.
The 4-hourly variation of sensor readings from January, 2019, to March, 2021, for the water quality indicators in the outlet of the Aitoutan (ATT) watershed.
The red dotted lines represent the boundary of environmental quality standards for surface water in China. The water quality levels gradually deteriorate from level Ⅰ to level Ⅴ and the value of indicators exceeding the level Ⅴ is defined as “worse than Ⅴ”. For DO, the higher value represents the better water quality level, and for TP and NH4+-N, the higher value represents the worse water quality level.
Table 1.
Descriptive statistics of input indicators and output nutrients from the monitoring site located in the outlet of Aitoutan (ATT) watershed.
Fig 3.
Correlation analysis for the input and output indicators.
The statistical significance of rank correlations is denoted by asterisks for p < 0.05 (*) and p < 0.01 (**) (lower left). The different sizes and colors of circles represent the strength of the correlation between the indicators (upper right).
Table 2.
Comparison of the average estimation accuracy of the three machine-learning models (4-hourly frequency, testing step, n = 842).
Fig 4.
Comparison of the models’ performances by Taylor diagrams.
RF = “random forest”; SVM = “support vector machine”; BPNN = “back-propagation neural network”; TP = “total phosphorous”; TN = “total nitrogen”; and NH4+-N = “ammonia-nitrogen”.
Table 3.
Comparison of the average estimation accuracy of the RF model with three sampling frequencies (testing step).
Fig 5.
Scatterplots of the observations and average estimations with three sampling frequency scenarios.
The x-axis represents the observations while the y-axis represents the estimations. The grey dashed line represents the 1:1 fitted line of observations and estimations under ideal conditions. The red line represents the fitted line of observations and estimations in actual situation.
Fig 6.
Estimated R2 and Nash-Sutcliffe efficiency (NSE) values for the random-forest (RF) model under different sampling frequency scenarios.
The width of the violin shape indicates the frequency at which R2 and NSE appear at this value.
Table 4.
Results of the analysis-of-variance (ANOVA) test.
Fig 7.
Relative importance analysis results of five input indicators in the random-forest (RF) model.