Figures
Abstract
Interictal epileptiform discharges (IEDs) are diagnostically important EEG abnormalities, but after an epoch has already been confirmed as containing IED, a separate methodological question remains whether its scalp distribution can be assigned among predefined categories, and do routinely co-recorded auxiliary channels alter that classification. This conditional five-class tasks were evaluated using 2,514 expert-confirmed four-second IED epochs labelled as generalized, frontal, temporal, occipital, or centro-parietal. Identical stratified epoch-level partitions, training-only SMOTE, 26 handcrafted features per included channel, and multiple machine-learning classifiers were used across a staged channel ablation comparing 19-channel scalp EEG, 21-channel EEG with two ECG channels, and the complete 29-channel input containing scalp EEG, referential, ECG, and EMG channels. Linear discriminant analysis achieved the best EEG-only test accuracy (88.89%), whereas CatBoost achieved 93.25% on EEG with ECG signals and 94.44% with the whole channel. All eight directly comparable classifiers showed numerically higher test accuracy after ECG was added, although these differences were descriptive and were not subjected to formal paired significance testing. SHAP analysis identified the key channel-feature variables influencing the fitted models, complementing performance ablation. In the EEG with ECG on CatBoost model, ECG channels RA and LA contributed 15.79% and 15.12% of normalized global attribution, respectively, while beta-band power was the most important feature family, accounting for 18.76%. The study therefore provides a controlled analysis of downstream spatial categorization and auxiliary-channel dependence within expert-confirmed IED epochs. The results do not establish physiological localization, unseen-patient generalization, or a clinically validated ECG biomarker.
Citation: Plabon AM, Mukit A, Neyamul M, Jehady OF, Zuba FT, Mina MF, et al. (2026) Conditional spatial classification of expert-confirmed interictal epileptiform discharge epochs: An EEG-ECG ablation and SHAP analysis. PLoS One 21(9): e0358782. https://doi.org/10.1371/journal.pone.0358782
Editor: Teppei Matsubara, Athinoula A Martinos Center for Biomedical Imaging, UNITED STATES OF AMERICA
Received: November 27, 2025; Accepted: September 4, 2026; Published: September 22, 2026
Copyright: © 2026 Plabon et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: No new data were collected. Publicly available data. Data link: https://doi.org/10.6084/m9.figshare.28069568.v2.
Funding: The author(s) received no specific funding for this work.
Competing interests: The authors have declared that no competing interests exist.
Introduction
Epilepsy is a neurological disorder characterized by recurrent unprovoked seizures arising from abnormal neuronal activity [1]. Electroencephalography (EEG) is central to clinical assessment because it records cerebral electrical activity and can reveal interictal epileptiform discharges (IEDs) between seizures [2,3]. Beyond the presence of an IED, its morphology and scalp distribution contribute to electroclinical description and annotation [4]. The scalp-distribution label, however, should not be equated with source localization, the seizure-onset zone, or the epileptogenic zone.
In the source dataset, expert-confirmed IEDs were organized into five predefined scalp-distribution categories: generalized, frontal, temporal, occipital, and centro-parietal [5]. Once an epoch has already been identified as containing an IED, assigning such a distribution label becomes a distinct downstream classification problem. Studying this conditional task is methodologically useful because it isolates discrimination among expert-defined spatial categories from the separate problem of deciding whether an IED is present at all. It can therefore provide a controlled benchmark for downstream annotation and for examining what recorded channel groups contribute to the separation of these labels. Throughout this article, “spatial classification” refers only to these predefined scalp-distribution categories and not to source reconstruction or localization of the seizure-onset zone.
IED detection and conditional spatial classification address different questions [6,7]. Detection evaluates whether unselected EEG contains an IED in the presence of non-IED background, artifacts, and IED-like transients [8]. The present analysis begins only after expert confirmation of IED presence and asks how the confirmed epoch is categorized among five scalp-distribution labels. This restricted design intentionally holds the detection decision constant so that channel-set and model effects can be examined within the IED-positive subset.
Machine learning provides a practical analytical framework for this task because each epoch is represented by hundreds of channel-wise time-domain, frequency-domain, and nonlinear descriptors, and the mapping from these variables to five labels may be multivariate and model-dependent. Evaluating several classifier families under identical data partitions allows the study to test whether the conditional categories are separable from the handcrafted representation and whether observed channel-set effects are consistent across different modelling assumptions. The purpose is not to replace expert IED detection, but to evaluate a high-dimensional downstream categorization problem after expert annotation.
The source recordings include scalp EEG together with referential, electrocardiogram (ECG), and electromyogram (EMG) channels [5]. ECG was evaluated because it is a simultaneously recorded auxiliary modality that can be added without changing the IED labels or scalp EEG channels. The EEG-versus-EEG with ECG comparison therefore asks a focused ablation question whether does the fitted classifier obtain additional discriminatory information when the two ECG channels are available. This test does not assume that any gain is autonomic or physiologically localizing; rather, it helps determine whether the model depends on information outside scalp EEG and highlights the need to consider about patient, state, or recording-specific structure when interpreting multimodal models. The subsequent full-input comparison jointly adds four referential and four EMG channels and consequently cannot separate their individual effects.
SHapley Additive exPlanations (SHAP) were incorporated to answer a complementary question to the channel ablation [9]. Ablation quantifies whether predictive performance changes when a channel group is added, whereas SHAP identifies which channel-feature variables a fitted model uses to produce its outputs. This combination makes SHAP central to the interpretability aim of the study: it can show whether an apparent performance gain is accompanied by substantial model dependence on the added modality and which features account for that dependence. Because SHAP importance is conditional on the dataset, feature representation, model, and target [10], the resulting attributions are interpreted only as model-specific predictive contributions, not as physiological biomarkers, causal mechanisms, or evidence of clinical localization [11]. This caution is especially important for ECG and EMG variables, which may reflect patient identity, sleep-wake state, recording conditions, movement, contamination, or other dataset-specific structure [12–14].
Accordingly, this study aimed to: (1) evaluate whether expert-confirmed IED epochs can be separated among five predefined scalp-distribution categories using a common handcrafted-feature representation; (2) compare multiple machine-learning classifiers under identical epoch-level partitions; (3) quantify the descriptive performance change when two ECG channels are added to 19-channel scalp EEG; (4) compare EEG with ECG using the complete 29-channel input; and (5) utilize SHAP to characterize model dependence on specific channels and features, while keeping physiological and clinical claims explicitly outside the scope of the analysis.
Methodology
The study used a conditional spatial-classification pipeline restricted to expert-confirmed IED epochs in five predefined classes: generalized, frontal, temporal, occipital, and centro-parietal. The same preprocessing, channel-wise feature extraction, data partitions, imbalance handling, classifier evaluation, and SHAP framework were applied across channel configurations so that the ablation compared channel availability rather than different modelling pipelines.
Proposed methodology
For the proposed work, the same expert-confirmed IED epochs followed the workflow shown in Fig 1. Twenty-six handcrafted features were extracted independently from every included channel. Three primary feature matrices were created: (1) 19 scalp EEG channels (19 x 26 = 494 variables); (2) 19 scalp EEG channels plus the two ECG channels, RA and LA (21 x 26 = 546 variables); and (3) the full 29-channel input comprising 19 scalp EEG, four referential channels, two ECG channels, and four EMG channels (29 x 26 = 754 variables). Each classifier received a flat vector of channel-feature variables. Electrode coordinates, adjacency matrices, inter-electrode distances, topographic maps, and source-reconstruction outputs were not used; spatial identity was represented only through channel-specific feature names.
Dataset
There are 29 channels on the recordings: 19 scalp EEG channels positioned using the international 10–20 system illustrated in Fig 2, four scalp referential channels (T1, T2, A1, and A2), 2 channels for ECG (RA and LA), and 4 channels for deltoid EMG (bilateral) depicted in Fig 3. Table 1 lists the channel groups and the location where they were recorded. The ECG RA and ECG LA are referred to as lead24 and lead25 in the EEG with ECG SHAP outputs, while this was the original feature naming convention.
The “Fp” refers to frontopolar, “F” to frontal, “C” to central, “P” to parietal, “O” to occipital, and “T” to temporal locations.
The right and left arm ECGs are called RA and LA. Bilateral Deltoid EMG channels D1-D4 are indicated.
The open dataset was obtained from the Epilepsy Center of Peking Union Medical College Hospital [5] and contains recordings from 84 participants. Fifty-two recordings contained at least one expert-confirmed IED, whereas 32 recordings were reported as normal EEG. Each participant contributed 20 minutes of recording, giving 28 hours in total. The source dataset contains 22,933 four-second epochs not labelled as IEDs and 2,516 four-second epochs labelled as IEDs. During quality control, two IED epochs were discarded, leaving 2,514 expert-confirmed IED epochs for the present analysis. Non-IED epochs were not included as a target class because the study was designed specifically to evaluate conditional categorization after IED presence had already been established.
The recordings and epoch labels were provided in MATLAB and NumPy formats. Table 2 summarizes the consciousness state of the original 2,516 epochs labelled as IEDs.
EEG signal preprocessing
The four-second epochs were linearly detrended, notch filtered at 50 Hz and 100 Hz, and band-pass filtered between 0.5 and 40 Hz using a Butterworth filter [15]. Each channel configuration underwent the same applicable preprocessing before feature-matrix construction.
The composite criterion of robust log-variance outliers and low interchannel correlation was used to identify noisy or defective EEG channels. Detected bad EEG channels were interpolated from neighboring channels [13]. Independent component analysis was only used on the EEG channels to remove ocular artifacts while retaining the ECG and EMG signals [14,16]. All scalp EEG channels were then re-referenced to a common average.
Two epochs of wake state were rejected due to the peak-to-peak amplitude being above a strong median absolute deviation (MAD) threshold. The final dataset therefore contained 2,514 IED epochs: 762 wake epochs (30.31%) and 1,752 sleep epochs (69.69%).
Feature extraction
For each retained four-second epoch, the same database of 26 handcrafted descriptors was computed independently for every included channel, yielding 494 variables for EEG only, 546 for EEG with ECG, and 754 for the full 29-channel input (scalp EEG, referential, ECG, and EMG). The features were grouped into time-domain, frequency-domain, and nonlinear categories. The same descriptor set was intentionally used across EEG and ECG channels to keep the representation fixed during channel-set ablation: if modality-specific feature engineering were introduced only for ECG, any performance change would reflect both the added channels and a changed feature representation. This design choice does not imply that EEG and ECG have equivalent physiology or that every descriptor has the same physiological interpretation across modalities; it provides a controlled, modality-agnostic statistical representation for the comparative analysis.
Time-domain features.
The variables used in time domain are mean, median, standard deviation, skewness, kurtosis, peak-to-peak amplitude, waveform length, slope sign changes, Hjorth mobility and Hjorth complexity.
Frequency-domain features.
Welch power spectral density estimates were used to compute frequency domain variables. Absolute band power was calculated for each band: delta (0.5–4 Hz), theta (4–8 Hz), alpha (8–13 Hz) and beta (13–30 Hz) as well as the ratio of each band power, peak frequency and median frequency.
Nonlinear features.
Nonlinear variables were approximate entropy, sample entropy, permutation entropy, spectral entropy, correlation dimension, detrended fluctuation analysis and the Hurst exponent [17].
Data splitting
The 2,514 IED-labelled epochs were partitioned using a stratified 80:10:10 epoch-level split which contains 2,011 training epochs, 251 validation epochs, and 252 held-out test epochs. Class proportions were maintained, and identical epoch assignments were used across all channel configurations and classifiers. The partition was not patient-disjoint, so epochs from the same participant could occur in more than one subset. Consequently, the evaluation measures internal held-out-epoch performance and cannot establish generalization to unseen patients. A patient-disjoint or leave-one-patient-out design would be required for that stronger claim. Table 3 summarizes the distribution of the five IED categories in the retained dataset.
Addressing class imbalance with SMOTE analysis
Synthetic Minority Over-sampling Technique (SMOTE) was applied only to the training subset to reduce imbalance among the five IED categories [18]. Validation and test epochs were not resampled. Applying SMOTE after the split prevented synthetic information from entering the validation or test sets.
For a minority-class feature vector , a synthetic vector was generated by interpolation between
and a randomly selected nearest neighbour x_nn from the same class:
where:
is the original minority-class sample;
is a randomly selected nearest neighbour from the same class;
- Lambda
is sampled from a uniform distribution on [0,1]; and
is the synthetic sample on the line segment joining
and
.
Machine learning classifier evaluation
Machine learning was used here as a comparative framework for mapping the high-dimensional handcrafted feature vectors to five conditional spatial labels. The evaluated classifiers were XGBoost, LightGBM, CatBoost, multinomial logistic regression, linear discriminant analysis (LDA), decision tree, extra trees, random forest, and support vector machine (SVM). Including linear, tree-based, ensemble, and boosting approaches allowed the analysis to assess whether channel-set patterns were specific to one model family. Random-forest results were available for the EEG-only and EEG with ECG analyses, while eight classifiers had results for all three primary channel configurations. The same procedure and partitions were used for each ablation. Hyperparameters were selected using the validation set, and final performance was reported on the held-out test set.
The performance was summarized in terms of overall accuracy, class-specific precision, recall and F1 score along with macro-averaged F1 score [16]. Macro-F1 was highlighted due to the fact that it gives equal importance to each spatial category irrespective of the class frequencies.
Channel-set ablation design
The channel-set ablation compared feature matrices derived from identical IED epochs and an unchanged feature-extraction pipeline. The EEG-only matrix contained 26 features from each of 19 scalp EEG channels. The EEG with ECG matrix added both ECG RA and ECG LA channels. The full matrix contained all 29 recorded channels: 19 scalp EEG, four referential, two ECG, and four EMG channels. Labels, split assignments, training-only SMOTE, classifier families, and evaluation procedures were held constant. Thus, the EEG-versus-EEG with ECG comparison describes the performance change associated with making ECG-derived variables available to the model. The EEG with ECG-versus-full comparison reflects the joint addition of four referential and four EMG channels and cannot isolate either group independently.
Performance differences between channel configurations are reported descriptively. Formal paired hypothesis tests and confidence intervals were not calculated because the epoch-level prediction outputs needed for paired error analysis were not retained for every configuration. The reported percentage-point differences therefore must not be interpreted as statistically significant effects, and the study does not claim that ECG provides a statistically validated improvement. Inferential testing using retained paired predictions or repeated patient-disjoint resampling is required to establish reproducible performance differences [19].
An additional exploratory EEG with EMG configuration, comprising the 19 scalp EEG channels and four bilateral deltoid EMG channels, was evaluated as a supplementary analysis. The same preprocessing, feature-extraction, data-partitioning, SMOTE based training, classifier-evaluation, and SHAP-analysis procedures were applied. Detailed results are provided in S1 File.
SHapley Additive exPlanations (SHAP) explainability analysis
SHAP values were used to describe how strongly each channel-feature variable contributed to the outputs of a fitted model [9]. Global importance was calculated as the mean absolute SHAP value across test epochs and output classes and was then normalized to percentages within each feature set. This analysis complements, rather than replaces, the channel ablation: ablation addresses whether performance changes when channels are added, while SHAP identifies where the fitted model places predictive dependence after those channels are available. The normalized percentages are neither statistical effect sizes nor physiological effect sizes. They should not be compared between configurations as though they reflected the same causal quantity since they are relative model-attribution measures inside a certain dataset, feature space, classifier, and target.
Shapley value for a feature i.
- For feature
, the Shapley value is the weighted average change in model output produced by adding that feature to all possible subsets of the remaining features.
- The calculation considers feature subsets that do not contain
and compares the prediction before and after
is added.
- The difference between these predictions is the marginal contribution of feature
.
- Weighted averaging across subsets yields the Shapley value.
Multi-Class SHAP (per sample and class).
For multi-class classification, SHAP values are obtained for each sample, feature, and output class. A positive or negative SHAP value indicates whether the feature shifts the model output toward or away from a class. Class-specific SHAP values were converted to absolute magnitudes for global importance summaries. This formulation extends Shapley attribution to the five-class output.
SHAP importance for feature
.
The overall significance of feature has been taken as interest here rather than a single sample or class. Thus, the absolute SHAP values are averaged across all samples (data points) all
classes. Taking the absolute value makes sure both positive influence and negative influence are counted as importance. This gives a global importance score for each feature.
Model performance metrics
Model performance was derived from the multi-class confusion matrix. For each class, true positives, false positives, and false negatives were defined using a one-versus-rest formulation. Overall accuracy measured the proportion of correctly classified epochs.
- True positives (TP): epochs of a class correctly assigned to that class.
- True negatives (TN): epochs outside a class correctly not assigned to that class.
- False positives (FP): epochs from other classes incorrectly assigned to the class.
- False negatives (FN): epochs of the class incorrectly assigned elsewhere.
Precision, recall, and F1-score were calculated separately for each class.
- Accuracy was calculated as the number of correct predictions divided by the total number of test epochs.
- Precision was calculated as TP/(TP+FP) and measures the proportion of predictions for a class that were correct.
- Recall was calculated as TP/(TP+FN) and measures the proportion of epochs from a class that were correctly identified.
- F1-score was calculated as the harmonic mean of precision and recall. Macro-F1 was the unweighted mean of the five class-specific F1-scores.
Together, these metrics summarize overall and class-balanced conditional classification performance.
Results
Class-wise metrics were calculated for training, validation, and held-out test subsets. Model selection used the validation set, while principal conclusions are based on the untouched test set. Table 4 summarizes the original full 29-channel results. Tables 5 and 6 compare test accuracy and macro-F1 across channel configurations, and Table 7 reports CatBoost class-specific F1-scores. All between-configuration differences are descriptive.
For the full 29-channel configuration, CatBoost produced the highest test accuracy (94.44%) and macro-F1 (94.44%). Its class-specific F1-scores were 0.965 for generalized, 0.903 for frontal, 0.943 for temporal, 0.951 for occipital, and 0.960 for centro-parietal epochs. These values quantify conditional categorization among expert-confirmed IED epochs and are not measures of IED detection.
Channel-set ablation results
Adding ECG produced numerically higher held-out test accuracy for all eight classifiers evaluated across the three primary configurations. CatBoost increased from 86.90% with EEG only to 93.25% with EEG with ECG, a descriptive difference of 6.35 percentage points. The full 29-channel CatBoost result was 94.44%, a further descriptive increase of 1.19 percentage points. The full configuration exceeded EEG with ECG for six classifiers, whereas LDA and decision tree performed better with EEG with ECG. Thus, the fixed-split results show a consistent numerical association between ECG inclusion and higher internal epoch-level accuracy, while the effect of the remaining channel groups was model-dependent. Because paired inferential testing was not performed, these differences are not presented as statistically significant improvements.
Detailed validation- and test-set results for the EEG with EMG, EEG-only, and EEG with ECG configurations are provided in S1 File. These supplementary results include class-specific precision, recall, and F1-score, together with overall classification accuracy for the evaluated classifiers. Table 5 depicts the test accuracy across channel configurations and descriptive differences for EEG signals only, EEG with ECG signals and full 29 channel signals while Table 6 summarizes macro-F1 across channel configurations of test set. And finally, Table 7 summarizes CatBoost class-specific test F1-scores across different channel configurations
Relative to EEG-only CatBoost, EEG with ECG yielded higher F1-scores for generalized, frontal, temporal, occipital, and centro-parietal categories. The full configuration yielded the highest generalized, temporal, and occipital F1-scores; EEG with ECG yielded the highest centro-parietal F1-score; and frontal F1 was 0.903 in both ECG-containing configurations. These observations describe the fixed test split and were not subjected to formal inferential testing.
EEG with ECG SHAP analysis
Within the EEG with ECG CatBoost model as stated in Table 8, ECG RA and ECG LA received 15.79% and 15.12% of normalized global SHAP attribution, respectively. Together they accounted for 30.90% of normalized attribution in that feature set. Beta-band power was the leading feature family (18.76%), followed by peak-to-peak amplitude (9.72%), theta power (8.08%), alpha-beta ratio (7.98%), and alpha power (7.28%). Table 8 lists the largest individual channel-feature attributions.
Within the EEG with ECG using CatBoost model, the two ECG channels received substantial normalized attribution, indicating that the fitted classifier used ECG-derived variables when separating the five expert-assigned labels. This attribution pattern does not demonstrate a clinically meaningful interictal biomarker or an autonomic localization signal. Patient identity, cardiac morphology, sleep-wake state, recording conditions, movement, contamination, and other dataset-specific structure remain plausible alternative explanations, which is precisely why the SHAP analysis is interpreted as a model-dependence analysis rather than physiological validation.
SHAP feature-importance analysis
Table 9 summarizes aggregated SHAP attribution for the full 29-channel models. The values identify variables used by the trained classifiers in the conditional five-class task; they do not show that a feature independently improved accuracy or represents a validated biomarker.
Table 10 reports channel-wise SHAP attribution for the full 29-channel configuration.
CatBoost achieved the highest full-channel test accuracy. Table 11 lists its ten largest combined channel-feature attributions.
Across the full-channel models in Tables 9–11, spectral-power and waveform-morphology variables frequently received high attribution. Beta power and peak-to-peak amplitude were prominent in several tree-based classifiers. These results indicate that the models used these variables to separate the five expert-assigned labels; they do not establish that the variables detect IEDs or identify their physiological source.
In the full configuration, ECG RA and ECG LA received high attribution in several classifiers, while scalp EEG channels including Fz, F8, and Fp1 also contributed. When combined, the SHAP and ablation analyzes offer two distinct descriptive perspectives: the ablation shows that ECG availability coincided with higher fixed-split accuracy across the eight comparable classifiers, and SHAP shows that the fitted models assigned appreciable predictive attribution to ECG-derived variables. Neither result establishes a causal autonomic mechanism or statistically validated modality benefit, and normalized SHAP percentages can change when the feature set changes.
In EEG-only CatBoost, beta power contributed 17.93% to normalized feature-family attribution and the most dominant channels were Fz (12.18%), F3 (10.74%), Fp1 (9.06%) and F8 (8.34%). In EEG with ECG CatBoost beta power was ranked 18.76%, while ECG RA and ECG LA were the two top ranked channels. These percentages are descriptive features of the fitted models.
Additional SHAP analyses are presented in S1 File, including aggregated feature-family importance, channel-wise importance, and the ten highest-ranked channel-feature attributions for each analysed classifier. These normalized SHAP values represent model-specific dependence within the corresponding feature set and should not be interpreted as physiological effect sizes, validated biomarkers, or evidence of causal mechanisms.
Discussion
This study addressed a focused downstream question whether after expert confirmation that an epoch contains an IED, can the epoch be classified among five predefined scalp-distribution categories, and does access to routinely co-recorded auxiliary channels alter that classification. Under a common handcrafted-feature and machine-learning pipeline, the five labels were separable on the internal held-out epoch split, and ECG inclusion was associated with numerically higher accuracy for every directly comparable classifier. SHAP then identified the variables on which the fitted models depended, providing an attribution-level complement to the channel-set ablation. These results define the study as a methodological analysis of conditional categorization and multimodal model dependence, rather than an IED detector or clinical localization system.
Scope and methodological contribution
The principal methodological contribution is the controlled separation of three questions that are often mixed up which includes IED detection, downstream spatial categorization, and interpretation of auxiliary-channel dependence. By fixing IED presence through expert annotation, the analysis isolates discrimination among five scalp-distribution labels. By holding epochs, feature extraction, partitions, training strategy, and classifier procedures constant across channel sets, the ablation asks whether making additional recorded modalities available changes internal classification. This restricted setting is useful for benchmarking downstream annotation models and for revealing whether a classifier may exploit information outside the expected scalp EEG channels. It does not reproduce routine clinical EEG review, in which IED presence is unknown, and it does not constitute source localization.
Performance of the classifiers in the context of prior work
CatBoost test accuracy was 86.90% with EEG only, 93.25% with EEG with ECG, and 94.44% with the full 29-channel input. The numerical increase associated with adding ECG was larger than the subsequent increase associated with adding the remaining channels for CatBoost, but the pattern was classifier-dependent. LDA and decision tree performed better on EEG with ECG channels than with the full channel set. Because no paired confidence intervals or hypothesis tests were available, these differences are descriptive. The results therefore show a dataset-specific fixed-split association between ECG availability and classification performance and do not establish a statistically significant, universal, causal, or clinically validated benefit of ECG.
Table 12 places the present work alongside representative literature to clarify task and validation differences, not to claim benchmark superiority. Most prior studies address binary IED detection or use different spatial targets, datasets, architectures, and validation strategies; Chung et al. [9], for example, used a patient-disjoint approach for focal IED detection and lobe classification. Because those methods were not reimplemented on the present dataset under the same partitions, their published performance values are contextual only and are not direct baselines for the current conditional five-class task.
Previous studies indicate that performance may decline as epileptiform classification moves from binary decisions to finer spatial categories [6,20]. The present results show that handcrafted-feature classifiers can separate five predefined labels within a curated collection of IED-positive epochs. This observation does not imply that the task is more clinically difficult than binary detection, because non-IED background, artifacts, and IED-like transients were excluded.
Classical and boosting-based methods remained competitive for the structured handcrafted features. CatBoost performed best with the ECG-containing configurations, whereas LDA performed best with EEG only. This configuration-dependent ranking argues against treating any model or channel set as universally superior.
Earlier feature-based studies showed that handcrafted EEG variables can support epileptiform analysis, often in binary detection settings [24]. The contribution of the present analysis is different and narrower: it applies a common handcrafted representation and fixed modelling workflow to quantify how conditional five-class performance changes across scalp EEG, EEG with ECG, and the complete recorded channel set, and then uses SHAP to identify which variables the fitted models depend on. The study does not demonstrate superiority over prior methods because those methods were not benchmarked on this dataset, and it does not test whether ECG improves IED-versus-non-IED detection.
Cross-study comparisons are limited by differences in annotation, class definitions, preprocessing, sampling, validation, channel composition, and target definition. More importantly, the current epoch-level split may place epochs from one participant in multiple subsets, whereas patient-disjoint schemes such as leave-one-patient-out validation are more appropriate for estimating performance on unseen individuals. The reported values should therefore be interpreted only as internal held-out-epoch performance. Direct reimplementation of prior methods on the same dataset and patient-disjoint evaluation would be required for a rigorous comparative or clinical-generalization claim.
Top features and leads (SHAP analysis)
The channel ablation provides a descriptive performance pattern that complements the SHAP rankings. All eight comparable classifiers achieved numerically higher fixed-split test accuracy with ECG than with scalp EEG alone. This consistency motivates examination of how the fitted models used the added modality, but it is not equivalent to statistical evidence of an ECG effect. ECG was not evaluated for separating IED from non-IED activity, and the full-input comparison cannot identify an independent EMG effect because referential and EMG channels were added together.
SHAP is useful in this design because performance alone cannot show which variables account for a model’s decisions after a channel group is added. By decomposing model outputs into channel-feature attributions, SHAP indicates whether the fitted classifier actually depends on the newly available modality and identifies the features through which that dependence appears. For example, an EEG frontal electrode may rank highly because it helps separate frontal-labelled epochs from other IED-labelled categories, while ECG ratios may contribute to class separation for dataset-specific reasons. These attributions explain the fitted model; they do not validate physiology, detect IEDs, or localize an epileptogenic source.
The prominence of beta power, waveform morphology, and selected EEG and ECG channels is consistent with the models exploiting spectral and amplitude differences among labelled epochs. This compatibility does not prove a physiological explanation. Normalized ECG attribution may also be influenced by feature-space composition. Patient-disjoint validation and controlled confound analyses are required before clinical or biological meaning can be assigned.
Interpretation of ECG and EMG contributions
High ECG attribution may reflect patient-specific cardiac morphology, sleep-wake state, medication, recording conditions, movement, cross-channel contamination, or other dataset structure; a genuine interictal neurocardiac association is only one possible explanation. The current analysis cannot distinguish among these alternatives. Seizure-related autonomic literature is therefore not treated as direct evidence for ECG-based spatial classification of interictal discharges.
The performance ablation shows that ECG inclusion was associated with improved conditional classification across the tested classifiers, but the mechanism remains unresolved. The full configuration does not isolate EMG because referential and EMG channels were introduced together. ECG SHAP prominence should therefore be regarded as hypothesis-generating rather than proof of an autonomic mechanism, a clinical biomarker, or spatial source information.
Imbalanced dataset handling
SMOTE was only used on the training subset. Validation and test distributions were left unchanged. This reduced training imbalance but did not eliminate unequal class representation or guarantee equivalent performance across categories.
Strengths of the study
Strengths include expert-confirmed epoch labels, identical epoch partitions across channel configurations, training-only oversampling, multiple classifier families, class-specific metrics, a staged channel ablation, and model-level interpretation with SHAP. In particular, the design keeps the feature representation and modelling workflow fixed when ECG is added, so the observed descriptive performance differences can be attributed to channel availability within this dataset rather than to a simultaneous change in feature engineering. The paired use of ablation and SHAP also distinguishes the question of whether performance changes from the question of which variables the fitted model uses.
Limitations
Several limitations define the interpretation of this work. First, only expert-confirmed IED epochs were analyzed; consequently, the study does not evaluate IED detection, specificity, false-positive reduction, artifact rejection, or clinical screening. Second, the epoch-level split was not patient-disjoint. Participant overlap may allow models, particularly ECG-containing models, to exploit patient- or recording-specific information and may yield optimistic estimates relative to unseen-patient evaluation. Third, configuration differences are descriptive because paired prediction outputs were not retained for formal hypothesis tests or confidence intervals; therefore, no statistically significant ECG improvement is claimed. Fourth, prior published methods were not reimplemented on this dataset, so Table 12 provides context rather than a direct performance benchmark and the study does not claim superiority over existing approaches. Fifth, although EEG with ECG isolates the addition of the two ECG channels, the full configuration jointly adds referential and EMG channels, so independent EMG and referential effects remain unknown. Sixth, the same modality independent descriptor library was used for EEG and ECG to preserve a controlled ablation; this supports comparability but does not exploit modality-specific feature engineering. Seventh, spatial information is represented implicitly through channel-specific feature names rather than explicit electrode geometry or topographic maps. Finally, SHAP attributions are model-specific associations and do not establish physiological biomarkers or causal mechanisms.
Future directions
Future work should prioritize patient-disjoint evaluation, including leave-one-patient-out or comparable subject-level validation, followed by external validation. Retaining subject identifiers and per-epoch prediction outputs would also permit paired confidence intervals, significance testing, and analysis of patient-specific confounding. A rigorous benchmark should reimplement representative prior methods, including approaches mentioned in Table 12, on the same dataset and compare them under identical patient-disjoint partitions, with and without ECG where technically applicable. Additional extensions should include representative non-IED background, artifacts, and IED-like transients; separate ablations of referential, ECG, and EMG channels; modality-specific feature representations tested against the current common-feature baseline; and explicit topographic or graph-based spatial representations. These studies are needed before the descriptive ECG association observed here can be interpreted as reproducible or clinically useful.
Conclusion
This study demonstrates that conditional five-class categorization of expert-confirmed IED epochs is a separable downstream machine-learning problem within the evaluated dataset and fixed epoch-level split. A controlled channel ablation showed numerically higher test accuracy after ECG was added to scalp EEG for all eight directly comparable classifiers; CatBoost increased from 86.90% with EEG only to 93.25% with EEG with ECG and 94.44% with the full 29-channel input. SHAP analysis further showed that the fitted ECG-containing models placed substantial attribution on ECG-derived variables, helping explain how the models used the added modality. These findings establish a methodological basis for studying auxiliary-channel dependence in conditional IED categorization, but the performance differences are descriptive rather than statistically significant, and no direct benchmark against prior methods was performed. Patient-disjoint validation, paired inferential testing, direct benchmarking, and evaluation that includes non-IED epochs are required before unseen-patient or clinical utility can be assessed.
Supporting information
S1 File. The file contains validation- and test-set analyses for the EEG with EMG, EEG-only, and EEG with ECG configurations, including class-specific performance metrics, aggregated feature-family importance, channel-wise importance, and classifier-specific channel-feature rankings.
https://doi.org/10.1371/journal.pone.0358782.s001
(DOCX)
References
- 1. Fisher RS, et al. ILAE official report: A practical clinical definition of epilepsy. Epilepsia. 2014;55:475–82.
- 2. Davak A. A comprehensive review on epilepsy and its diagnosis/ treatment using eeg signals. EIJMHS. 2025.
- 3. Islam T, Basak M, Islam R, Roy AD. Investigating population-specific epilepsy detection from noisy EEG signals using deep-learning models. Heliyon. 2023;9(12):e22208.
- 4. Kural MA, Duez L, Sejer Hansen V, Larsson PG, Rampp S, Schulz R, et al. Criteria for defining interictal epileptiform discharges in EEG: A clinical validation study. Neurology. 2020;94(20):e2139–47. pmid:32321764
- 5. Lin N, Zheng M, Li L, Hu P, Gao W, Sun H, et al. An EEG dataset for interictal epileptiform discharge with spatial distribution information. Sci Data. 2025;12(1):229. pmid:39920162
- 6. Zhang L, Wang X, Jiang J, Xiao N, Guo J, Zhuang K, et al. Automatic interictal epileptiform discharge (IED) detection based on convolutional neural network (CNN). Front Mol Biosci. 2023;10:1146606. pmid:37091867
- 7. da Silva Lourenço C, Tjepkema-Cloostermans MC, van Putten MJAM. Machine learning for detection of interictal epileptiform discharges. Elsevier Ireland Ltd; 2021.
- 8. Lin N, Gao W, Li L, Chen J, Liang Z, Yuan G, et al. vEpiNet: A multimodal interictal epileptiform discharge detection method based on video and electroencephalogram data. Neural Netw. 2024;175:106319. pmid:38640698
- 9. Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv. Neural Inf. Process. Syst. 2017;2017:4766–75. https://arxiv.org/pdf/1705.07874
- 10. Janzing D, Minorics L, Bloebaum P. Feature relevance quantification in explainable AI: A causal problem. PMLR. 2020. Accessed: Jul 26 2026. https://proceedings.mlr.press/v108/janzing20a.html
- 11. Jin W, Li X, Fatehi M, Hamarneh G. Guidelines and evaluation of clinical explainable AI in medical image analysis. Med Image Anal. 2023;84:102684. pmid:36516555
- 12. Brookshire G, et al. Data leakage in deep learning studies of translational EEG. Front Neurosci. 2024;18.
- 13. Bigdely-Shamlo N, Mullen T, Kothe C, Su K-M, Robbins KA. The PREP pipeline: Standardized preprocessing for large-scale EEG analysis. Front Neuroinform. 2015;9:16. pmid:26150785
- 14. Delorme A, Makeig S. EEGLAB: An open source toolbox for analysis of single-trial EEG dynamics including independent component analysis. J Neurosci Methods. 2004;134(1):9–21. pmid:15102499
- 15. Stancin I, Cifrek M, Jovic A. A review of eeg signal features and their application in driver drowsiness detection systems. Sensors; 2021;21(11).
- 16. Sokolova M, Lapalme G. A systematic analysis of performance measures for classification tasks. Information Processing & Management. 2009;45(4):427–37.
- 17. Richman JS, Moorman JR. Physiological time-series analysis using approximate and sample entropy. Am J Physiol Heart Circ Physiol. 2000;278.
- 18. Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP. SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research. 2002;16:321–57.
- 19. Nadeau C, Bengio Y. Inference for the generalization error. Machine Learning. 2003;52(3):239–81.
- 20. Chung YG, Lee W-J, Na SM, Kim H, Hwang H, Yun C-H, et al. Deep learning-based automated detection and multiclass classification of focal interictal epileptiform discharges in scalp electroencephalograms. Sci Rep. 2023;13(1):6755. pmid:37185941
- 21. Bagheri E, Jin J, Dauwels J, Cash S, Westover MB. A fast machine learning approach to facilitate the detection of interictal epileptiform discharges in the scalp electroencephalogram. J Neurosci Methods. 2019;326:108362. pmid:31310822
- 22. Thomas J, et al. Automated Detection of Interictal Epileptiform Discharges from Scalp Electroencephalograms by Convolutional Neural Networks. Int J Neural Syst. 2020;30(11).
- 23. Tjepkema-Cloostermans MC, Carvalho RCV, van Putten MJAM. Deep learning for detection of focal epileptiform discharges from scalp EEG recordings. Clinical Neurophysiology. 2018;129(10):2191–6.
- 24. Wang X, et al. A two-stage automatic system for detection of interictal epileptiform discharges from scalp electroencephalograms. eNeuro. 2023;10(11).
- 25. Fürbass F, Kural MA, Gritsch G, Hartmann M, Kluge T, Beniczky S. An artificial intelligence-based EEG algorithm for detection of epileptiform EEG discharges: Validation against the diagnostic gold standard. Clin Neurophysiol. 2020;131(6):1174–9. pmid:32299000