Fig 1.
Label distribution in sleep accel dataset.
Number of labels in the Sleep Accel dataset per time interval over 8 hours (considering three sleep stages). The dataset shows class imbalance, with 9 “NREM” labels for every 3 “REM” labels and 1 “Wake” label.
Table 1.
Distribution of samples per class for each classification task.
The percentages indicate the proportion of each class in relation to the total dataset.
Fig 2.
General pipeline of the proposed method.
The process begins with time series data, which are transformed into visual representations. The final output consists of predictions for two scenarios: two-stage (wake/sleep) and sleep stages classification (wake/NREM/REM).
Fig 3.
The process begins with transforming raw accelerometer (ACC) data (pink) and heart rate (HR) data (blue) into visual representations. These visual representations serve as the input for training and validation. An ACC and HR data ensemble is created and validated (gray). The images are divided into patches used as inputs for their respective training sessions. Validation is carried out based on the ensemble results obtained from all patches. Finally, the ACC + HR ensemble is performed and validated again after obtaining the ensemble results from the patches (gray).
Fig 4.
Images generated from the accelerometer data for each type of representation and class. RGB images combine the x, y, and z axes to utilize all motion information. It is possible to observe visual differences between the classes, indicating that the visual representations capture specific motion patterns associated with each sleep stage.
Fig 5.
Images generated from the heart rate data for each type of representation and class. Grayscale images are used since heart rate data consists of a single value (in bpm). It is possible to observe visual differences between the classes, indicating that the visual representations capture specific heart rate patterns associated with each sleep stage.
Fig 6.
“Wake” images generated with GAF and accelerometer data.
Example of an original image (600600) and examples of patches (224
224 each).
Table 2.
Percentage of Wake, NREM, and REM samples in the training and validation sets across the five splits of the cross-validation. The data indicate that the class distribution remains stable across splits, suggesting that the random split does not introduce substantial bias.
Table 3.
Balanced accuracies obtained with each representation for sleep/wake classification.
Accelerometer data consistently outperformed heart rate data in all scenarios, with the GAF achieving the highest balanced accuracy (82.36% 3.24%) when using patch ensembles. Patch-based ensembles significantly improved balanced accuracy compared to original images.
Fig 7.
RP confusion matrices for sleep/wake classification.
The highest balanced accuracy (80.39% 2.43%) was achieved with accelerometer data and the ensemble of patches, while the highest sensitivity (84%) was observed with the ensemble combining accelerometer and heart rate. Ensemble combining the original accelerometer and heart rate data achieves higher sensitivity than other approaches. Comparing confusion matrices from original data versus patches highlights an improvement in classifying the “Wake” stage. For accelerometer data and heart rate data, the use of patches increased the correct classification of “Wake”. Sleep/wake classification is more balanced when using accelerometer data, and this balance is further enhanced in the ensemble combining accelerometer and heart rate patches.
Fig 8.
GAF confusion matrices for sleep/wake classification.
The highest balanced accuracy (82.36% 3.24%) was achieved with accelerometer data and the ensemble of patches, while the highest sensitivity (85%) was observed with the ensemble combining accelerometer and heart rate patches. Sleep/wake classification using original heart rate data is more balanced than with original accelerometer data, where both classes are confused to a similar extent. Classification with original accelerometer data tends to overestimate “Sleep”. However, by improving predictions for “Wake” using patches, the accelerometer-based classification becomes more balanced.
Fig 9.
MTF confusion matrices for sleep/wake classification.
The highest balanced accuracy (80.32% 3.43%) was achieved with the ensemble combining accelerometer and heart rate patches, while the highest sensitivity (87%) was observed with the ensemble combining accelerometer and heart rate. Using original accelerometer and heart rate data, “Sleep” is classified more accurately than “Wake”. This pattern is reflected in the ensemble of original data. As observed in the RP and GAF representations, the use of patches for accelerometer data leads to a more balanced classification.
Fig 10.
Spectrograms confusion matrices for sleep/wake classification.
The highest balanced accuracy (79.11% 3.93%) was achieved with accelerometer data and the ensemble of patches, while the highest sensitivity (80%) was observed with both ensembles combining accelerometer and heart rate. Spectrogram representation shows balanced classifications for both original accelerometer and heart rate data. However, the ensemble of these original data better classifies “Sleep” stages. Using patches for accelerometer data improves the classification of “Wake”.
Fig 11.
Sleep/wake classification over a night of sleep for a subject using the GAF representation.
Original data shows more “Sleep” errors for “Wake” at the beginning and around 6 hours, and frequent “Wake” errors for “Sleep” early on. Ensemble of patches reduces “Sleep” errors for “Wake”, with most “Wake” errors for “Sleep” around 2 and 6 hours.
Table 4.
Balanced accuracies obtained with each representation for sleep stages classification.
Heart rate data often outperformed accelerometer data in balanced accuracies (except for the Spectrogram), with the GAF achieving the highest balanced accuracy (62.18% 0.95%) when using patch ensemble. Patch-based ensembles significantly improved balanced accuracy compared to original images.
Fig 12.
RP confusion matrices for sleep stages classification.
The highest balanced accuracy (61.87% 1.67%) was achieved with heart rate data and the ensemble of patches. Accelerometer data, including the ensemble results, achieved a higher number of correct classifications for “Wake”. In contrast, matrices generated with heart rate data alone showed more accurate classifications of “NREM” and “REM”. Additionally, with accelerometer data (both original and patches), the most frequent misclassification was labeling “NREM” as “REM”. For heart rate data, the most common error was classifying “Wake” as “REM”.
Fig 13.
GAF confusion matrices for sleep stages classification.
The highest balanced accuracy (62.18% 0.95%) was achieved with heart rate data and the ensemble of patches. The confusion matrix with heart rate patches demonstrated an increase in correct classifications of “NREM” and “REM”. The most common misclassifications with accelerometer data were labeling “NREM” as “REM” and “REM” as “NREM”. Meanwhile, with heart rate data, the most frequent error was classifying “Wake” as “REM”.
Fig 14.
MTF confusion matrices for sleep stages classification.
The highest balanced accuracy (60.97% 1.85%) was achieved with the ensemble combining accelerometer and heart rate patches. Accelerometer data and both types of ensembles most frequently classified “Wake” correctly, similar to other representations. For heart rate data, the incorrect classification of “Wake” as “REM” observed with original data decreased with the use of patches, resulting in a more balanced classification.
Fig 15.
Spectrograms confusion matrices for sleep stages classification.
The highest balanced accuracy (57.36% 2.68%) was achieved with heart rate data and the ensemble of patches. Heart rate data, both in its original and patched forms, resulted in fewer classifications of “NREM” and “REM,” overestimating “Wake”. However, the patched configuration increased the number of correct “REM” classifications while reducing incorrect “Wake” predictions. Both ensemble configurations improved correct classifications of “Wake” but showed increased confusion for “NREM” when the true class was “REM”.
Fig 16.
Sleep stages classification over a night of sleep for a subject using the GAF representation.
Original data shows frequent “Wake” for “NREM” and “REM” for “NREM” errors. Ensemble of patches reduces these errors, but “REM” to “Wake” errors persist early on, and “NREM” to “Wake” errors appear around the 8-hour mark.
Table 5.
Balanced accuracies obtained with raw data for sleep/wake classification.
Table 6.
Balanced accuracies obtained with raw data for sleep stages classification.
Table 7.
Balanced accuracies obtained with feature extraction for sleep/wake and sleep stages classifications.
Table 8.
Comparison of the best-balanced accuracies obtained with different data representations for sleep/wake classification.
Table 9.
Comparison of the best-balanced accuracies obtained with different data representations for sleep stages classification.