Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Fig 1.

In the “Cross-validation and testing” approach, the data are divided into two separate sets (cross-validation set and test set) only once.

First, different models are trained and validated with cross-validation and the best set of parameters is chosen. Prediction accuracy and statistical significance of the parameters are evaluated on the test set, after training on the cross-validation set.

More »

Fig 1 Expand

Fig 2.

In the “Nested cross-validation” approach, first (outer) cross-validation is performed to estimate predictability of the data.

In each iteration, data are divided into training and test sets. Before training, another (inner) cross-validation loop is used to optimize parameters. As model weights (fitted models) and parameters are different at every partition, it is not possible to report accuracy or statistical significance about a particular set of parameters or model weights.

More »

Fig 2 Expand

Fig 3.

For cross-validation and cross-testing, data are divided into two separate sets only once: a cross-validation set and a test set.

Similar to typical cross-validation, a number of iterations are carried out to choose the best parameters for the final model on the test set. Once the best combination of parameters has been chosen, the prediction accuracy and statistical significance can be evaluated on the test set with a modified cross-validation such that for each fold the original cross-validation set is repeatedly added to the training data. Due to the similarity to cross-validation, we term this approach cross-testing. While making it impossible to pick one final model on additional unseen data, the parameters that have been chosen remain interpretable.

More »

Fig 3 Expand

Table 1.

Comparison of the approaches.

More »

Table 1 Expand

Fig 4.

Batches of simulated data of different sizes are analyzed 1000 times with three different approaches.

Results show the mean accuracy (upper plot) and proportion of significant results (bottom plot) out of the 1000 runs. More data leads to higher average accuracy and increases the proportion of significant results. “Nested cross-validation” outperforms other approaches and the “cross-validation and testing” gives the worst performance in terms of average accuracy and proportion of significant results.

More »

Fig 4 Expand

Fig 5.

Batches of 100 simulated data points were analyzed 1,000 times with both approaches that contain a separate test set.

Results show the mean accuracy (upper plot) and proportion of significant results (bottom plot) out of the 1000 runs. A larger test set leads to smaller average accuracy because there is less data for choosing parameters and fitting a model. “Cross-validation and cross-testing” outperforms “cross-validation and testing” in terms of average accuracy and proportion of significant results as expected.

More »

Fig 5 Expand

Fig 6.

Analysis of real data (left: EEG dataset; right: spiking dataset) with three different approaches as a function of data size with test set size fixed at 50%.

Results show the mean accuracy (upper graphs) and proportion of significant results (bottom graphs) out of the 1000 runs. More data lead to higher average accuracy and increases the proportion of significant results. “Nested cross-validation” outperforms other approaches while “Cross-validation and testing” gives the worst performance in terms of average accuracy and proportion of significant results. The effect is smaller with EEG data suggesting that more efficient usage of data in the model fitting is not that important and the choice of parameters is actually the main influencer.

More »

Fig 6 Expand

Fig 7.

Analysis of neuroscience data (left: EEG dataset; right: spiking dataset) with three different approaches as a function of the relative test set size.

Results show the mean accuracy (upper graphs) and proportion of significant results (bottom graphs) out of the 1000 runs. Data set size was fixed to 50 for EEG and to 100 for spikes train data set. Larger test set leads to smaller average accuracy because there is less data for choosing parameters and fitting a model. “Cross-validation and cross-testing” outperforms “cross-validation and testing” in terms of average accuracy and proportion of significant results.

More »

Fig 7 Expand

Fig 8.

Accuracy and proportion of significant results are kept at chance levels when applying the novel approach to random data.

More »

Fig 8 Expand