Figures
Abstract
Objectives
To determine whether bacterial- and viral-associated whole-blood host-response modules discovered in GSE211567 retain their prespecified directions when applied unchanged to two external cohorts, including a second cross-platform cohort, and to quantify their robustness to analytical choices and gene deletion.
Methods
Five biologically guided modules were defined and fixed in GSE211567 before external analysis. The same modules were applied unchanged to GSE73461 and GSE72810 using an unweighted mean gene-wise z-score rule. Primary contrasts included 52 bacterial and 94 viral samples in GSE73461 and 23 definite bacterial and 28 definite viral samples in GSE72810. Wilcoxon tests were adjusted across the five modules using the Benjamini-Hochberg method. Hodges-Lehmann shifts, rank-biserial effects and 10,000-replicate bootstrap confidence intervals were calculated. Sensitivity analyses assessed z-score reference populations, probable-case inclusion, probe summarisation, GSVA scoring and exhaustive leave-one/two-gene deletion.
Results
All five modules retained their expected directions in both projection cohorts. BACT_M2, VIR_M1a, VIR_M1b and VIR_M2 had confidence intervals excluding zero and passed false-discovery-rate correction in both cohorts, whereas BACT_M1 remained directionally concordant but borderline in GSE73461 and GSE72810 (adjusted P = 0.0799 in each). All 30 z-reference, case-definition and probe-collapse sensitivity estimates retained the expected direction. Under GSVA, BACT_M1 gained statistical support, whereas VIR_M2 retained its viral-higher direction but lost confidence-interval and adjusted-P-value support. Across 29,826 leave-one/two-gene variants, every variant retained the expected direction and the minimum Pearson correlation with its complete-module score was 0.9940.
Conclusions
The predefined host-response modules retained their expected directions across two external cohorts measured on different Illumina platforms, with the strongest reproducible support for BACT_M2 and the three viral-associated modules. BACT_M1 and the GSVA behaviour of VIR_M2 demonstrate that direction preservation does not imply uniform statistical support across cohorts or scoring algorithms.
Citation: Maghembe RS (2026) External transportability of bacterial- and viral-associated host-response modules across public transcriptomic cohorts. PLoS One 21(9): e0353559. https://doi.org/10.1371/journal.pone.0353559
Editor: Benjamin M. Liu, Children's National Hospital, George Washington University, UNITED STATES OF AMERICA
Received: June 23, 2026; Accepted: August 25, 2026; Published: September 10, 2026
Copyright: © 2026 Reuben S. Maghembe. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All transcriptomic datasets reanalysed in this study are publicly available from the NCBI Gene Expression Omnibus. The discovery analysis used GSE211567, and the formal external projection analysis used GSE73461. Candidate or technical-rehearsal datasets considered during workflow development included GSE161731, GSE261482 and GSE68310. Derived data underlying the reported findings are provided as S1 Dataset, including the cohort audit, locked module definitions, identifier coverage, sample-level module scores and projection statistics.
Funding: The author(s) received no specific funding for this work.
Competing interests: The author has declared that no competing interests exist.
Introduction
Distinguishing bacterial from viral infection remains important for infectious-disease medicine, antimicrobial stewardship and host-response biology [1–4]. Although pathogen detection is central to diagnosis, microbiological results may be delayed, insensitive or difficult to interpret when colonisation, co-infection or previous antimicrobial exposure complicate pathogen attribution [2,5,6]. Host transcriptomic studies have therefore been widely explored as complementary approaches, but much work has emphasised diagnostic classifiers and prediction performance rather than whether the underlying host-response biology is transportable across heterogeneous cohorts [7–10].
A complementary question is whether biologically interpretable response programmes discovered in one setting behave similarly when tested in a different cohort. In this study, transportability means preservation of the prespecified biological direction when the same module definition and scoring rule are applied to another dataset. Fixed-module projection means applying predefined gene sets to an external dataset without changing their genes, expected directions or weights and without retraining a model. This is distinct from diagnostic validation. A classifier may perform well because of cohort-specific, technical or sampling structure, whereas the present analysis asks whether the same biological modules retain their expected bacterial-higher or viral-higher behaviour outside the discovery dataset [8,10,11].
Public infection transcriptomic cohorts are valuable for this purpose, but samples may differ by geography, clinical syndrome, pathogen spectrum, platform, preprocessing and case definition [7,9,11]. Examining the discovery contrast separately within the available geographic strata can reduce the risk that a pooled bacterial-versus-viral signal is driven mainly by one setting. The study therefore kept discovery and external testing separate: module definitions, expected directions and scoring rules were fixed before either external cohort was analysed.
Here, GSE211567 was used to discover bacterial- and viral-associated host-response programmes while requiring cross-site directional support. Five conservative modules were defined before external testing and then applied unchanged to GSE73461 and GSE72810. The latter provided a second accession-level, sample-level and cross-platform cohort. The aim was not to train or validate a diagnostic classifier, but to test whether predefined modules retained their expected directions, quantify cross-cohort effect sizes and determine sensitivity to score reference, probe handling, case definition, alternative gene-set scoring and gene deletion.
Methods
Study design and datasets
This study used a staged design that kept discovery separate from external testing. GSE211567 was used for discovery, biological interpretation and definition of the five modules. GSE73461 was analysed only after the module genes, expected directions and scoring rules had been fixed. GSE72810 was subsequently analysed as a second accession-level and sample-level cohort providing cross-platform validation. Neither external dataset was used to reselect genes, rename modules, alter module composition, tune weights or train a diagnostic classifier.
Public host-transcriptomic datasets were considered if they contained infection-relevant human transcriptomic data, recoverable sample metadata, interpretable pathogen-class labels, usable feature identifiers and sufficient sample structure for the intended analysis. GSE161731 was retained only as a technical rehearsal dataset, while GSE261482 and GSE68310 were audited but not selected for formal projection. GSE211567, GSE73461 and GSE72810 were obtained from the Gene Expression Omnibus [12–14].
GSE211567 discovery and module definition
GSE211567 provided the discovery dataset and contained whole-blood transcriptomic samples from Sri Lanka and the United States [12]. After metadata and expression-quality auditing, the prespecified primary discovery set contained 224 samples: 101 bacterial and 123 viral. The primary limma model evaluated bacterial versus viral infection while adjusting for site and sequencing batch, and separate bacterial-versus-viral models were also fitted within Sri Lanka and the United States. For each modelled feature, directional concordance was defined as agreement in the sign of the bacterial-versus-viral log2 fold change between the contrasts being compared. Percentage directional concordance was the number of same-sign features divided by the number of compared features, multiplied by 100. Spearman correlations of log2 fold changes were used as a complementary measure of ranked effect-size agreement across the pooled and site-specific analyses.
Differential-expression results were treated as a discovery-ranking layer rather than a diagnostic signature [15]. Positive log2 fold-change values were interpreted as bacterial-higher and negative values as viral-higher. Transcript-level features were mapped to gene identifiers, summarised at gene level and carried into Gene Ontology biological-process enrichment [16,17]. Cross-site directionally supported features were prioritised, overlapping or redundant enriched terms were grouped, and candidate biological programmes were reviewed using the documented curation hierarchy. Module selection was therefore a biologically guided and reproducible curation step rather than mathematical optimisation for separation or prediction accuracy. The external-cohort results were not used to choose or modify module genes. Five modules were fixed before external testing: BACT_M1, BACT_M2, VIR_M1a, VIR_M1b and VIR_M2.
GSE73461 formal external projection
GSE73461 expression, annotation, metadata and group labels were audited independently of module discovery [13]. The cohort was measured on the Illumina GPL10558 platform. The 52 DefiniteBacterial, 94 DefiniteViral and 55 Control samples (n = 201) constituted the main z-score reference population. The primary inferential contrast remained restricted to the 52 DefiniteBacterial versus 94 DefiniteViral samples; Control samples contributed to the reference population but not to the bacterial-versus-viral test. Inflammatory, Kawasaki and Unknown groups were excluded from both the main z-score reference population and the primary contrast.
Illumina probes were mapped to gene symbols, and locked genes were checked against predefined module-coverage thresholds. The numbers of scored genes were 24/25 for BACT_M1, 21/21 for BACT_M2, 128/128 for VIR_M1a, 33/33 for VIR_M1b and 105/106 for VIR_M2.
GSE72810 cross-platform validation and probe selection
GSE72810 contained 146 paediatric whole-blood samples measured using the Illumina HumanHT-12 v3 platform [14]. All 146 samples were retained in the main z-score reference population, whereas the prespecified primary inferential contrast was restricted to 23 definite bacterial versus 28 definite viral samples. Seventeen probable bacterial and seven probable viral samples were reserved for expanded-case sensitivity analysis. Sixteen controls and 55 uncertain samples contributed to the main z-score reference population but were excluded from the primary bacterial-versus-viral test.
Predefined module genes were reconciled to the GSE72810 platform through Entrez identifiers. Of 313 module-gene instances, 303 were mapped and 10 were unmapped. Module coverage was BACT_M1 24/25, BACT_M2 20/21, VIR_M1a 125/128, VIR_M1b 33/33 and VIR_M2 101/106. When multiple authorised probes represented the same Entrez gene, a prespecified rule selected the probe with the highest median expression across all 146 samples, with lexicographic ordering used only to resolve exact ties. Probe selection was completed before testing group differences.
Module scoring and statistical analysis
For the primary mean-z analyses, expression for each mapped gene g in sample i was standardised within the prespecified reference population as z_gi = (x_gi – mean_g)/ SD_g. For a module containing K mapped and available genes, the sample-level module score was the unweighted arithmetic mean, score_i = (1/K) sum_g z_gi. Thus, every available gene contributed equally to the primary module score. Missing genes were omitted only after module coverage had been documented. No module was retrained, reweighted or direction-flipped.
Wilcoxon rank-sum tests compared bacterial and viral module scores, with Benjamini-Hochberg correction across the five modules [18]. Positive effects denote bacterial-higher scores and negative effects denote viral-higher scores. In addition to median bacterial-minus-viral differences, the revision analyses calculated Hodges-Lehmann location shifts and rank-biserial effects. Confidence intervals were obtained using 10,000 bootstrap replicates.
Sensitivity and robustness analyses
Z-reference sensitivity was evaluated by repeating scoring using only the primary bacterial and viral samples as the reference population. GSE72810 sensitivity analyses additionally included probable bacterial and viral cases and compared the locked representative-probe scores with mean scores across all authorised probes. Score concordance was assessed using Pearson and Spearman correlations.
A GSVA sensitivity analysis was performed in GSE73461 using the unchanged module gene sets and both z-reference populations. Because mean-z and GSVA scores have different numerical scales, cross-method interpretation focused on effect direction, rank-biserial effect, confidence intervals and adjusted P values rather than direct comparison of raw score magnitudes.
Gene-deletion robustness was assessed exhaustively in GSE73461 by recalculating every module after removal of each individual gene and every pair of genes. Variant scores were compared with the corresponding complete-module scores, and expected-direction retention, Wilcoxon results and Pearson and Spearman correlations were recorded.
Cross-cohort interpretation and reproducibility boundaries
The GSE72810 and GSE73461 GEO sample accession sets were disjoint and the cohorts were measured on different Illumina array platforms. However, direct participant overlap could not be assessed because participant identifiers were not deposited, and the studies arose from the same broad investigator network. GSE72810 is therefore described as a second accession-level and sample-level cross-platform cohort rather than as a fully investigator-independent replication cohort.
Analysis scripts, decision logs, quality gates, source-data tables, manuscript-facing figures and supplementary outputs were organised in the project repository. The analyses evaluate transportability and robustness of the predefined biological modules. They do not constitute diagnostic classifier discovery, clinical-performance validation, clinical implementation evidence or causal validation.
Ethics approval
This study reanalysed publicly available de-identified transcriptomic datasets and did not involve new recruitment of human participants, new collection of human biospecimens or access to identifiable private information. No new ethics approval was required for this secondary analysis of public data.
Declaration of generative AI and AI-assisted technologies in the manuscript preparation process
During preparation and revision of this manuscript, the author used ChatGPT, an AI-assisted tool provided by OpenAI, to assist with editorial organisation, language refinement, workflow planning, code drafting and checking, formatting checks, and preparation of manuscript and submission-support materials. AI-generated suggestions were not treated as scientific evidence. Analysis scripts were executed against the stated public datasets, and numerical and graphical outputs were checked using the documented reproducibility, source-lock and quality-control procedures. The author reviewed and verified the analysis code, results, references, biological interpretation, figures, tables, manuscript text and submission materials and takes full responsibility for the final work.
Results
GSE211567 discovery identifies bacterial- and viral-associated programmes across sites
The GSE211567 discovery analysis used a predefined bacterial-versus-viral contrast while keeping module definition separate from external testing. The primary limma analysis ranked host-transcriptomic features before cross-site concordance checks across the pooled and site-stratified analyses (Fig 1A-B). This procedure reduced the likelihood of carrying forward features driven mainly by one geographic or technical stratum.
(A) Primary bacterial-versus-viral discovery analysis in GSE211567. The volcano-style plot summarises the limma-ranked host-transcriptomic contrast used as the discovery starting point before cross-site concordance assessment and biological module definition. Positive log2 fold-change values indicate bacterial-higher features, whereas negative values indicate viral-higher features. (B) Cross-site concordance of the GSE211567 bacterial-versus-viral discovery contrast. Spearman correlations summarise ranked agreement of log2 fold changes between the pooled, Sri Lanka and United States analyses. Directional concordance is the percentage of compared features whose bacterial-versus-viral log2 fold changes have the same sign in the two indicated analyses. This analysis was used to prioritise signals whose direction was not restricted to a single site before pathway interpretation and module definition. (C) The five predefined GSE211567 discovery modules used in external testing. Two modules were bacterial-higher, representing cytoplasmic translation/ribosomal protein activity and mitochondrial respiration/oxidative phosphorylation. Three modules were viral-higher, representing broad antiviral/interferon-stimulated defence, viral restriction/type I interferon signalling and cytokine/innate immune regulation. Module genes and expected directions were fixed before external analysis, with no later gene reselection or module redefinition.
Conservative biological curation defines five fixed modules
Directionally eligible features were mapped to genes and assessed using Gene Ontology biological-process enrichment. Redundancy-reduced biological groups were reviewed using the documented curation hierarchy before final module locking. Two modules were bacterial-higher: BACT_M1, representing cytoplasmic translation and ribosomal activity, and BACT_M2, representing mitochondrial respiration and oxidative phosphorylation. Three modules were viral-higher: VIR_M1a, VIR_M1b and VIR_M2, representing broad antiviral/interferon defence, viral restriction/type I interferon activity and cytokine/innate immune regulation, respectively (Fig 1C).
GSE73461 formal external projection retains all expected directions
The primary GSE73461 contrast contained 52 DefiniteBacterial and 94 DefiniteViral samples. All modules passed the predefined coverage threshold and were scored without gene reselection, module redefinition, reweighting or model training.
All five modules retained their discovery directions (Fig 2; Table 1). BACT_M2 was bacterial-higher, with a median bacterial-minus-viral score difference of +0.3328 and BH-adjusted Wilcoxon P = 0.0202. BACT_M1 was also bacterial-higher but remained borderline after correction, with a median difference of +0.2067 and adjusted P = 0.0799. VIR_M1a, VIR_M1b and VIR_M2 were viral-higher, with median differences of −0.4629, −0.6739 and −0.2596 and adjusted P values of 4.77 x 10^-6, 1.41 x 10^-6 and 0.00848, respectively. Primary-only z-reference scoring retained all expected directions and the same four-module pattern of adjusted statistical support.
(A) Distribution of locked module scores in the independent GSE73461 external projection cohort. Module scores were calculated using the pre-specified unweighted mean z-score rule in DefiniteBacterial and DefiniteViral samples. Genes were scored without gene reselection, reweighting, module renaming or diagnostic model training. (B) Median bacterial-minus-viral module-score differences in the main GSE73461 projection and in the primary-only z-score sensitivity analysis. Positive values indicate higher scores in DefiniteBacterial samples, whereas negative values indicate higher scores in DefiniteViral samples. All five modules retained the expected discovery direction in both analyses. (C) BH-adjusted Wilcoxon P values for the main projection and the primary-only z-score sensitivity analysis are shown as independent points for each categorical module. Circles denote the main projection and triangles denote the primary-only z-score sensitivity analysis. The dashed horizontal line indicates BH-adjusted P = 0.05. The module categories are not connected because they are distinct predefined gene sets rather than points on a continuous sequence. BACT_M1 remained directionally concordant but borderline, whereas BACT_M2 and all three viral-associated modules passed the adjusted-P-value threshold in both mean-z analyses.
Table 1 reports fixed-module external projection. It should not be interpreted as diagnostic classifier discovery, model training, gene rediscovery, module redefinition or causal validation.
GSE72810 provides accession-level and cross-platform validation
The GSE72810 primary contrast contained 23 definite bacterial and 28 definite viral samples. Coverage of the predefined module genes remained high after Entrez reconciliation and prespecified representative-probe selection. All five modules retained their expected directions.
BACT_M2 had a Hodges-Lehmann bacterial-minus-viral shift of +0.425 (95% CI + 0.161 to +0.676; adjusted P = 0.0020). VIR_M1a, VIR_M1b and VIR_M2 had shifts of −0.726 (95% CI −0.907 to −0.577), −0.941 (95% CI −1.201 to −0.752) and −0.576 (95% CI −0.682 to −0.475), with adjusted P values of 7.16 x 10^-8, 8.77 x 10^-8 and 8.77 x 10^-8, respectively. BACT_M1 was bacterial-higher but borderline, with a shift of +0.266 (95% CI −0.024 to +0.722; adjusted P = 0.0799).
Cross-cohort effects support four modules in both cohorts
For harmonised cross-cohort comparison, Fig 3 and Table 2 report Hodges-Lehmann bacterial-minus-viral shifts with bootstrap 95% confidence intervals; these estimates are distinct from the median score differences reported for the original GSE73461 projection in Fig 2 and Table 1. All ten cohort-module effects retained the expected direction (Fig 3; Table 2). BACT_M2, VIR_M1a, VIR_M1b and VIR_M2 had confidence intervals excluding zero and passed false-discovery-rate correction in both GSE73461 and GSE72810. BACT_M1 was directionally concordant in both cohorts, but both confidence intervals included zero and both adjusted P values were approximately 0.08.
Hodges-Lehmann bacterial-minus-viral module-score shifts are shown with stratified nonparametric bootstrap 95% confidence intervals for the five modules predefined in the GSE211567 discovery analysis.
Positive estimates indicate higher module scores in bacterial samples, whereas negative estimates indicate higher scores in viral samples. GSE73461 contained 52 DefiniteBacterial and 94 DefiniteViral samples measured on GPL10558. GSE72810 contained 23 definite bacterial and 28 definite viral samples measured on GPL6947.
BACT_M2, VIR_M1a, VIR_M1b and VIR_M2 retained their expected directions with confidence intervals excluding zero and BH-adjusted Wilcoxon P values below 0.05 in both cohorts. BACT_M1 retained the expected bacterial-higher direction in both cohorts but remained statistically borderline.
The modules were scored without gene reselection, module redefinition, gene reweighting, direction flipping or diagnostic-model training.
GSE72810 was analysed as a second accession-level and sample-level cohort providing cross-platform validation. Its GSM accession set was disjoint from GSE73461 and it used a different Illumina platform. Direct participant overlap could not be assessed because participant identifiers were not deposited, and the studies arose from the same broad investigator network.
Effects are Hodges-Lehmann bacterial-minus-viral location shifts with bootstrap 95% confidence intervals. Positive values indicate bacterial-higher scores and negative values indicate viral-higher scores. Benjamini-Hochberg adjustment was applied across the five modules within each cohort.
The cross-cohort pattern therefore supported four modules under both effect-size and adjusted-significance criteria. The strongest and most consistent signals were the viral-associated modules, followed by the bacterial mitochondrial respiration/OXPHOS module.
Sensitivity analyses retain direction but identify method-dependent support
Across the two GSE73461 z-reference analyses and four GSE72810 z-reference, case-definition and probe-collapse analyses, all 30 module estimates retained the expected direction (S1 Fig, panel A). Twenty-four estimates passed false-discovery-rate correction and 25 rank-biserial confidence intervals excluded zero. GSE72810 score representations were highly concordant; the minimum Pearson correlation was 0.9874 and the minimum Spearman correlation was 0.9814.
The GSVA analysis retained all expected directions but showed non-uniform inferential support (S1 Fig, panel B). BACT_M2, VIR_M1a and VIR_M1b remained supported under both mean-z and GSVA scoring. BACT_M1 was borderline under mean-z scoring but gained confidence-interval and adjusted-P-value support under GSVA. In contrast, VIR_M2 was supported under mean-z scoring but was near zero and statistically unsupported under GSVA in both reference populations. VIR_M2 is therefore scoring-method-sensitive rather than uniformly robust across algorithms.
Exhaustive deletion analysis supports distributed module signal
The leave-one/two-gene analysis evaluated 29,826 module variants across the two GSE73461 scoring populations (S1 Fig, panel C). Every variant retained the expected module direction. The minimum Pearson correlation with the corresponding complete-module score was 0.9940. These results indicate that the observed module behaviour was not driven by removal-sensitive dependence on one gene or one gene pair, while not establishing causal sufficiency of individual genes.
Discussion
This study evaluated whether biologically curated whole-blood host-response modules discovered in GSE211567 retained their prespecified bacterial-higher or viral-higher behaviour when the same definitions and scoring rules were applied to two external cohorts. All five modules retained their expected directions in both GSE73461 and GSE72810. Four modules – BACT_M2, VIR_M1a, VIR_M1b and VIR_M2 - had confidence intervals excluding zero and passed adjusted-significance thresholds in both cohorts. BACT_M1 was directionally concordant but borderline in each cohort, indicating a reproducible direction with less stable inferential support.
The viral-associated modules showed the strongest transportability. VIR_M1a and VIR_M1b captured broad antiviral, interferon-stimulated and viral-restriction programmes, while VIR_M2 represented cytokine and innate immune regulation. Their cross-cohort behaviour is consistent with the central role of interferon-linked responses in viral infection and with prior studies in which interferon-inducible biomarkers contributed to viral-versus-bacterial discrimination [3,4,19]. BACT_M2 also transported consistently, supporting a bacterial-associated mitochondrial respiration and oxidative-phosphorylation programme in these datasets and aligning with evidence that inflammatory activation can remodel leukocyte immunometabolism [20,21].
The sensitivity analyses refine rather than uniformly strengthen this interpretation. Direction was stable across z-reference populations, GSE72810 case definitions and probe handling, and GSE72810 score correlations remained high. Exhaustive deletion analysis also supported a distributed module signal. However, GSVA changed inferential support for two modules: BACT_M1 became supported, whereas VIR_M2 lost statistical support despite retaining the expected sign. Directional robustness, effect-size robustness and statistical robustness should therefore be reported as related but distinct properties.
The principal contribution is not a diagnostic classifier, but a transparent framework for asking whether predefined biological modules retain their behaviour in new datasets. Fixing the module genes, expected directions and scoring rules before external testing, and prohibiting gene reselection, sign reversal, reweighting and model retraining, reduces rediscovery bias. Hodges-Lehmann shifts, rank-biserial effects, confidence intervals, cross-platform testing and deletion analyses provide complementary evidence about the magnitude and stability of the observed programmes.
Several limitations remain. Public metadata and infection adjudication were imperfect, and the cohorts differed in age, syndrome, pathogen composition, platform and preprocessing. Whole-blood scores may reflect both cell-composition changes and cell-intrinsic activation. The GSE72810 and GSE73461 GEO sample accession sets were disjoint and were measured on different Illumina platforms, but direct participant overlap could not be assessed because participant identifiers were not deposited; the studies also arose from the same broad investigator network. GSE72810 should therefore not be treated as a fully investigator-independent replication cohort. The retrospective analyses do not establish clinical performance, clinical readiness or causal mechanisms. Prospective, investigator-independent and multi-cohort studies are required to determine how these modules behave across age groups, syndromes, pathogens, sampling times and treatment contexts.
Transparency declaration
Code availability.
Analysis scripts, decision logs, quality gates, source-data tables, manuscript-facing figures and supplementary outputs are available in the public project repository at https://github.com/rmaghembe1/host-response-module-transportability. The revision-round code and outputs used for this resubmission are available on the public plosone_revision_round1_2026 branch at https://github.com/rmaghembe1/host-response-module-transportability/tree/plosone_revision_round1_2026.
Supporting information
S1 Dataset. Supplementary Tables S1–S5.
External-cohort search register, locked GSE211567 module definitions, GSE73461 identifier coverage and probe choices, GSE73461 sample-level scores, and complete GSE73461 primary and z-reference sensitivity statistics.
https://doi.org/10.1371/journal.pone.0353559.s001
(XLSX)
S2 Dataset. Supplementary Table S6.
GSE72810 cohort audit, sample-classification framework, locked primary and expanded contrasts, and cross-cohort independence assessment.
https://doi.org/10.1371/journal.pone.0353559.s002
(XLSX)
S3 Dataset. Supplementary Table S7.
GSE72810 Entrez reconciliation, module coverage, complete module-gene mapping, and prespecified representative-probe choices.
https://doi.org/10.1371/journal.pone.0353559.s003
(XLSX)
S4 Dataset. Supplementary Table S8.
Privacy-safe GSE72810 sample-level module scores using anonymized analysis identifiers, primary and sensitivity tests, bootstrap effect-size confidence intervals, and score-concordance analyses.
https://doi.org/10.1371/journal.pone.0353559.s004
(XLSX)
S5 Dataset. Supplementary Table S9.
Harmonised GSE73461–GSE72810 cross-cohort effect-size source data and five-module summary.
https://doi.org/10.1371/journal.pone.0353559.s005
(XLSX)
S6 Dataset. Supplementary Table S10.
GSE73461 mean-z/GSVA comparison, score concordance, module-level deletion summary, worst-case deletion variants, and complete exhaustive leave-one/two-gene results.
https://doi.org/10.1371/journal.pone.0353559.s006
(XLSX)
S1 Fig. Sensitivity and robustness of locked host-response modules.
A, Z-reference, case-definition and probe-collapse sensitivities. Rank-biserial bacterial-versus-viral effects and bootstrap 95% confidence intervals are shown for the two GSE73461 mean-z analyses and four GSE72810 analyses. All 30 cohort-module estimates retained the expected direction. GSE72810 score concordance across the tested z-reference and probe-collapse representations remained high, with minimum Pearson r = 0.9874 and minimum Spearman rho = 0.9814. B, Mean-z versus GSVA scoring-method sensitivity in GSE73461. Rank-biserial effects are displayed because they are comparable across scoring methods with different raw score scales. BACT_M2, VIR_M1a and VIR_M1b retained confidence-interval and BH-adjusted statistical support under both methods and z-reference populations. BACT_M1 was borderline under mean-z scoring but supported under GSVA. VIR_M2 retained a small viral-higher direction under GSVA but its confidence intervals included zero and its BH-adjusted P values were not significant. VIR_M2 is therefore described as scoring-method-sensitive. C, Exhaustive leave-one/two-gene robustness in GSE73461. Points show the minimum Pearson correlation between each deletion variant and the corresponding complete-module score. Across 29,826 variants, every leave-one and leave-two analysis retained the expected module direction. The minimum Pearson correlation was 0.9940 and the minimum Spearman correlation was 0.9897. Positive rank-biserial effects indicate bacterial-higher scores and negative effects indicate viral-higher scores. These analyses evaluate robustness of the predefined modules and do not constitute gene reselection, module redefinition, diagnostic-model training or causal validation.
https://doi.org/10.1371/journal.pone.0353559.s007
(PNG)
Acknowledgments
The author thanks the investigators and participants of the public transcriptomic studies reanalysed in this work, including GSE211567, GSE73461 and GSE72810. The author also acknowledges the public repositories and database maintainers that made these datasets available for secondary analysis. This acknowledgement does not imply endorsement of the present analysis by the original dataset generators.
References
- 1. Metlay JP, Waterer GW, Long AC, Anzueto A, Brozek J, Crothers K. Diagnosis and treatment of adults with community-acquired pneumonia. Am J Respir Crit Care Med. 2019;200:e45-67.
- 2. Halabi S, Shiber S, Paz M, Gottlieb TM, Barash E, Navon R, et al. Host test based on tumor necrosis factor-related apoptosis-inducing ligand, interferon gamma-induced protein-10 and C-reactive protein for differentiating bacterial and viral respiratory tract infections in adults: diagnostic accuracy study. Clin Microbiol Infect. 2023;29(9):1159–65. pmid:37270059
- 3. Papan C, Argentiero A, Porwoll M, Hakim U, Farinelli E, Testa I, et al. A host signature based on TRAIL, IP-10, and CRP for reducing antibiotic overuse in children by differentiating bacterial from viral infections: a prospective, multicentre cohort study. Clin Microbiol Infect. 2022;28(5):723–30. pmid:34768022
- 4. Rhedin S, Eklundh A, Ryd-Rinder M, Peltola V, Waris M, Gantelius J, et al. Myxovirus resistance protein A for discriminating between viral and bacterial lower respiratory tract infections in children - The TREND study. Clin Microbiol Infect. 2022;28(9):1251–7. pmid:35597507
- 5. Jain S, Self WH, Wunderink RG, Fakhran S, Balk R, Bramley AM, et al. Community-Acquired Pneumonia Requiring Hospitalization among U.S. Adults. N Engl J Med. 2015;373(5):415–27. pmid:26172429
- 6. Rutjes AWS, Reitsma JB, Coomarasamy A, Khan KS, Bossuyt PMM. Evaluation of diagnostic tests when there is no gold standard. A review of methods. Health Technol Assess. 2007;11(50):iii, ix–51. pmid:18021577
- 7. Herberg JA, Kaforou M, Wright VJ, Shailes H, Eleftherohorinou H, Hoggart CJ, et al. Diagnostic Test Accuracy of a 2-Transcript Host RNA Signature for Discriminating Bacterial vs Viral Infection in Febrile Children. JAMA. 2016;316(8):835–45. pmid:27552617
- 8. Andres-Terre M, McGuire HM, Pouliot Y, Bongen E, Sweeney TE, Tato CM, et al. Integrated, Multi-cohort Analysis Identifies Conserved Transcriptional Signatures across Multiple Respiratory Viruses. Immunity. 2015;43(6):1199–211. pmid:26682989
- 9. Sweeney TE, Wong HR, Khatri P. Robust classification of bacterial and viral infections via integrated host gene expression diagnostics. Sci Transl Med. 2016;8(346):346ra91. pmid:27384347
- 10. Oved K, Cohen A, Boico O, Navon R, Friedman T, Etshtein L, et al. A novel host-proteome signature for distinguishing between acute bacterial and viral infections. PLoS One. 2015;10(3):e0120012. pmid:25785720
- 11. Leticia Fernandez-Carballo B, Escadafal C, MacLean E, Kapasi AJ, Dittrich S. Distinguishing bacterial versus non-bacterial causes of febrile illness - A systematic review of host biomarkers. J Infect. 2021;82(4):1–10. pmid:33610683
- 12.
National Center for Biotechnology Information. Gene Expression Omnibus Accession GSE211567. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE211567
- 13.
National Center for Biotechnology Information. Gene Expression Omnibus Accession GSE73461. Gene Expression Omnibus. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE73461
- 14.
National Center for Biotechnology Information. Gene Expression Omnibus Accession GSE72810. Gene Expression Omnibus. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE72810
- 15. Ritchie ME, Phipson B, Wu D, Hu Y, Law CW, Shi W, et al. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 2015;43(7):e47. pmid:25605792
- 16. Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, et al. Gene Ontology: tool for the unification of biology. Nat Genet. 2000;25:25–9.
- 17. Gene Ontology Consortium. The Gene Ontology resource: enriching a GOld mine. Nucleic Acids Res. 2021;49:D325–34.
- 18. Benjamini Y, Hochberg Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society Series B: Statistical Methodology. 1995;57(1):289–300.
- 19. Schneider WM, Chevillotte MD, Rice CM. Interferon-stimulated genes: a complex web of host defenses. Annu Rev Immunol. 2014;32:513–45. pmid:24555472
- 20. O’Neill LAJ, Kishton RJ, Rathmell J. A guide to immunometabolism for immunologists. Nat Rev Immunol. 2016;16(9):553–65. pmid:27396447
- 21. Russell DG, Huang L, VanderVen BC. Immunometabolism at the interface between macrophages and pathogens. Nat Rev Immunol. 2019;19(5):291–304. pmid:30679807