Fig 1.
Human-shaped image indicating pupil dilation, electrodermal activity, blood volume pulse, and skin temperature of participants.
To explain the study [8], this image was created by only mimicking the shape of the original image, and it differs from the image actually used for training. The original image can be accessed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International license.
Fig 2.
A Mel-scaled spectrogram generated from a song.
On the x-axis, we have the time dimension, representing the duration of the audio segment. The y-axis denotes the frequency. The color intensity in the spectrogram indicates the amplitude (or energy) of different frequencies at each point in time, with warmer colors representing higher amplitudes and cooler colors indicating lower amplitudes.
Fig 3.
The process of converting a song into Mel-scaled spectrograms.
Initially, the song is segmented into discrete units, each spanning 10 seconds. Subsequently, each of these 10-second segments is individually transformed into a Mel-scaled spectrogram.
Table 1.
Dataset information includes the number of songs and the number of Mel-scaled spectrograms converted from songs.
SM and Non-SM stand for stress relief music and non-stress relief music, respectively.
Table 2.
Structures of a) ResNet-18, 50, 101 and b) DenseNet-161, 169, 201.
The architectural structures of two types of convolutional neural network models: a) ResNet and b) DenseNet. Specifically, it details the layer configurations, kernel sizes, and channel dimensions for three variants of ResNet (ResNet-18, ResNet-50, ResNet-101) and three variants of DenseNet (DenseNet-161, DenseNet-169, DenseNet-201).
Fig 4.
The design of the clinical study employing a 2 × 2 crossover methodology.
Participants were randomized into two sequence groups, A and B. Group A first experienced Individual Music (IM) followed by Researcher-selected Music (RM) after a washout period. Conversely, Group B started with RM and then transitioned to IM, also separated by a washout period.
Table 3.
A summary table of the participants’ basic biological information, categorized by age and sex.
It displays the mean and median ages, the age range (minimum and maximum values), and the distribution of participants by sex for each sequence group of the clinical study.
Table 4.
A table of baseline demographics, detailing the initial levels of stress, happiness, and satisfaction among participants before the clinical study commenced.
It includes mean and median values, as well as the range (minimum and maximum scores) for each emotional state across the two sequence groups.
Fig 5.
The comparative testing accuracy curves for ResNet-18, ResNet-50, ResNet-101, DenseNet-161, DenseNet-169, and DenseNet-201 models, using both custom and DEAM datasets.
The curves illustrate how the accuracy rates of each model vary over the testing period.
Table 5.
A comprehensive summary of testing accuracy, F1-score, Recall, and Precision metrics for the custom and DEAM datasets, as evaluated across a range of models including ResNet-18, ResNet-50, ResNet-101, DenseNet-161, DenseNet-169, and DenseNet-201.
Fig 6.
The distribution of Visual Analog Scale (VAS) scores for stress, happiness, and satisfaction, measured before and after the clinical test.
It provides a visual comparison of the emotional state changes experienced by participants as a result of the intervention.
Table 6.
The VAS scores for stress, happiness, and satisfaction before and after the clinical test.
The data is summarized to show the mean and standard deviation of participants’ scores, highlighting the changes in emotional states prompted by the clinical intervention.
Table 7.
The results of the non-inferiority test, comparing the effectiveness of Researcher Music (RM) to Individual Music (IM) based on stress, happiness, and satisfaction scores.
The data includes estimated means, differences between means, confidence intervals, p-values, and the assessment of non-inferiority.