Fig 1.
Overview of experiments and data types.
In the present study data related to the processing of a set of music and non-music sounds is obtained in three steps. In a behavioral experiment 14 participants gave continuous tension ratings while listening. This resulted in one time course of tension ratings for each stimulus and subject. In an EEG experiment the brain signals of nine subjects were recorded while they listened to the stimuli. Stimuli were repeated three times for each stimulus, resulting in 27 EEG recordings for each stimulus. From the waveform of each stimulus a set of nine acoustic/musical features was extracted from each stimulus.
Fig 2.
(1) In the first step of the analysis the 61-channel EEG signal (after generic preprocessing, see Methods) is temporally embedded and the power slope of the audio signal is extracted. In the training step (2) the embedded EEG features are regressed onto the audio power slope (Ridge Regression). After that (3) the resulting spatio-temporal filter (regression weight matrix) reducing the multichannel EEG to a one-dimensional projection is applied to a new presentation of the same stimulus. The regression filter can be transformed (4a) into a spatio-temporal pattern that indicates the distribution of information which is relevant for the reconstruction of the audio power slope. This spatio-temporal pattern, in turn, can be (4b) decomposed into components (derived with the MUSIC-algorithm) which have a scalp topography and a temporal signature. The EEG projections obtained in (3) subsequently are examined with respect to Cortico-Acoustic correlation (CACor).
Fig 3.
Example of stimulus waveform and tension ratings.
Top: Audio signal (blue) and sound intensity (red) for the Rachmaninov Prelude (entire stimulus). Middle: Tension ratings of single subjects (grey, N = 14), Grand Average (black, thick line) and standard deviation (black, thin line). Bottom: Activity index indicating the percentage of rising (blue) or falling (green) tension ratings at a given time point.
Fig 4.
CACor coefficients for each single subject and each stimulus presentation; GA: Grand Average (N = 27). Pink shading indicates significant positive correlation between EEG projection and audio power slope. Significance was determined by applying permutation tests and subsequently performing Bonferroni-correction for N = 27 presentations per stimulus.
Fig 5.
Three examples that show 3s-long segments of an extracted EEG projection (red) for a single stimulus presentation and a single subject and the respective audio power slope (blue) of Chords, Chopin Etude and Jungle noises. Note that in the optimization procedure a time lag between stimulus and brain response is integrated in the spatio-temporal filter, and that, consequently, the resulting EEG projections shown here are not delayed with respect to the audio power slope. The correlation coefficients indicate the magnitude of correlation for the shown segment of 3 s.
Fig 6.
CACor score profile and Coordination score profile.
The CACor score profile for the set of nine stimuli summarizes in how many of the 27 presentations significant Cortico-Acoustic Correlation was detected. Significance of correlation was determined (1) in a permutation-testing approach (darkblue bars) and (2) with Pyper et al.’s method (light blue bars) to estimate the effective degrees of freedom (see Methods). b) Behavioral Coordination score profile. All profiles are sorted according to the descending CACor score.
Table 1.
Relation between CACor scores and music features.
Spearman’s correlation coefficient (a) between CACor score profile and music feature profiles for nine acoustic/higher-level music features, (b) between Coordination score profile and music feature profiles. Bottom line: Correlation between CACor scores and Coordination scores. Asterisks indicate significant correlation.
Fig 7.
(a): Scalp topography (top) and time course (bottom) of ERP (single subject) derived by averaging channel Fz for all tone onsets of the stimulus 'Chord sequence'. The scalp topography corresponds to the shaded interval of 160–180 ms in the time course. (b): Spatial (top) and temporal (bottom) dimension of MUSIC component derived from spatio-temporal patterns for the same subject. c) Top: Reference pattern for selection of consistently occurring MUSIC components. This pattern is the Grand Average scalp topography obtained with classical ERP analysis. Individual scalp patterns that are contained in the Grand Average were determined by taking the mean scalp pattern across a 20 ms window that corresponds to the maximum amplitude of each subject’s onset ERP at channel Fz. The time windows of maximum amplitude were determined by visual inspection and ranged between 130 and 250 ms after tone onset. The bottom part of c) shows the individual time courses (grey) as well as the average time course (black).
Fig 8.
MUSIC components (Grand Average) for all stimuli.
Scalp topographies and time courses (thick line: average of all subjects, thin lines: standard deviation) of the MUSIC component that was extracted most consistently from the decomposed spatio-temporal patterns. Single-subject scalp topographies that are included in these averages are contained in the Supplementary Material.