Fig 1.
A screenshot of the billing table from SQL server containing unstructured data from patients.
Fig 2.
A screenshot of the MPS II dataset containing all symptoms from patients with dichotomous observations.
Fig 3.
Normal Q-Q plot of MPS II index from patients 21 years old or younger.
Red line represents a distribution reference line with μo equal to the sample mean for a normal distribution.
Fig 4.
Normal Q-Q plot of MPS II index from patients older than 21.
Red line represents a distribution reference line with μo equal to the sample mean for a normal distribution.
Fig 5.
The importance of features for MPS II disease forecasting by the NBC algorithm estimated using a ROC curve analysis conducted for each attribute.
Table 1.
Symptom combinations for potential patients diagnosed with MPS II disease by NBC algorithm.
Only the combinations with 1.6% incidence or higher have been presented here.
Table 2.
Features and their associated symptoms in MPS II disease.
The remained features in the final NBC model are show in bold.
Table 3.
Accuracy and Kappa values of features in the NBC model derived from Recursive Backward Feature Elimination algorithm and their positive predictive value.
Table 4.
The 2 × 2 contingency tables displays the performance evaluation using the bootstrapped resampling (n = 1000) and the Validation Set Approach technique on test dataset. Accuracy was used to select the optimal model by the largest value.
Table 5.
Performance comparison of Bayesian network classifiers using validation dataset.
Fig 6.
Top left: NBC = Naïve Bayes classifier; top right: TAN = Tree augmented Naïve-Bayes network; bottom left: BAN = Bayesian network augmented Naïve-Bayes network; bottom right: MBN = Markov blanket Bayesian network. Red circles are target variable (MPS II disease) and dark blue circles are features.