Figures
Abstract
The COVID-19 pandemic has shown the value of large-scale community PCR tests for epidemic surveillance, but the viral load measurements they provide have seldom been exploited to reconstruct within-host trajectories. Because these data are collected for diagnostic rather than research purposes, they are characterized by sparse longitudinal follow-up, heterogeneous sampling, and missing metadata, raising uncertainty about their usefulness to reconstruct with-host viral dynamics. We first conducted a simulation study to assess the feasibility of estimating the viral dynamics patterns from such datasets. Across multiple scenarios replicating realistic sampling patterns, we found that peak viral load and clearance time could be estimated with good accuracy, although uncertainty tended to be underestimated. Parameters driving early viral kinetics, e.g., incubation and proliferation time, were mostly estimated with poor precision, as most community tests are conducted after symptom onset, later in infection. We then applied this framework to a large dataset of 322,218 PCR tests associated with symptomatic SARS-CoV-2 infections in France between July 2021 and March 2022, encompassing both Delta-variant circulation and the emergence of first Omicron variants. We quantified the associations of age, vaccination status, and variant type with viral load trajectories. Age ≥ 65 years was consistently associated with a longer clearance time (2–6 days), while vaccination was associated with a shorter clearance time (2–4 days). Infections with Omicron variants were associated with lower peak viral load (2–3 Ct) and shorter clearance times (1–2 days) compared with pre-Omicron (Delta) infections. Thus, community PCR tests can be leveraged to identify key parameters of viral dynamics. As multiplex PCR testing becomes increasingly widespread, establishing robust frameworks for data collection, sharing, and privacy protection will be essential to support the use of these data for modelling purposes.
Author summary
In this study, we explored whether viral load data collected routinely in community laboratories can help us understand how factors like age, vaccination, and viral variant influence the course of infection. These data, collected for diagnostic purposes rather than research, are large but often incomplete, with most individuals tested only after the time of symptoms onset. We first used simulations to assess whether, despite these limitations, key aspects of viral dynamics can be reliably estimated. We then applied our approach to millions of test results collected in France during periods dominated by Delta and Omicron variants. We found that older age was consistently associated with longer infection duration, while vaccination was associated with a shorter duration of detectable virus without differences in peak viral levels. Omicron infections were generally associated with lower peak viral load and somewhat faster clearance than pre-Omicron infections. Overall, our work demonstrates that community testing data can provide valuable insights into viral dynamics and help monitor the effects of new variants or interventions. As multiplex tests become more common, facilitating and improving data collection, sharing, and privacy protection will be essential to make the most of these resources in future outbreaks.
Citation: Beaulieu M, Hozé N, Vieillefond V, Goetschy T, Cosentino G, Blanquart F, et al. (2026) Quantitative analysis of massive SARS-CoV-2 testing in the community in France in 2021–2022 reveals the associations of variant, vaccination, and age with viral dynamics in symptomatic individuals. PLoS Comput Biol 22(7): e1013811. https://doi.org/10.1371/journal.pcbi.1013811
Editor: Rustom Antia, Emory University, UNITED STATES OF AMERICA
Received: December 1, 2025; Accepted: July 8, 2026; Published: July 27, 2026
Copyright: © 2026 Beaulieu et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: The data underlying this study consist of individual-level human data collected in the context of routine clinical care and are not publicly available due to ethical, legal, and regulatory restrictions, including compliance with data protection regulations (e.g., GDPR). The data are owned and controlled by Biogroup, which serves as the independent data controller. While some employees of the company are co-authors of this study, data access requests are handled independently through the company’s established data access procedures. Data are available upon reasonable request to qualified researchers for legitimate scientific purposes, subject to review and approval in accordance with applicable regulations and the company’s data access policies. Requests for access should be directed to: Biogroup Email: recherche.nouveauprojet@biogroup.fr.
Funding: This work was supported by Université Paris Cité through a PhD scholarship (M.B.). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: I have read the journal’s policy and the authors of this manuscript have the following competing interests: V.V., T.G., and G.C. are employees of Biogroup; this affiliation did not influence the design, analysis, or interpretation of the study. The other authors have declared that no competing interests exist.
Introduction
Virological tests done in the general community have become an essential component of the public health surveillance system during an outbreak. During the Covid-19 pandemic, the design, volume and level of detail of these data have also opened new avenues for mathematical modellers. As a striking example, theoretical modelling performed in 2021 showed that the analysis of random large scale cross sectional distribution of SARS-CoV-2 viral loads could inform on the dynamics of the epidemics [1].
The emergence of variants of concern (VOCs) and the implementation of large-scale vaccination campaigns have largely influenced both the severity and the transmission of SARS-CoV-2 [2–5]. Yet, and despite dozens of studies, the impact of Omicron emergence and vaccination on viral load remains debated. In general, there is a consensus towards an effect of vaccination on viral load with pre-Omicron variants [6–8]. The effect of vaccination on Omicron variants is less clear, with some studies suggesting that vaccination may reduce viral load [7] while others suggest no effects [8–10].
Here we hypothesize that the millions of PCR tests done in the general population could be analysed in a systematic and quantitative way to leverage the information and reduce the potential bias present in small or retrospective cohort studies. Identifying signals associated with a change in viral dynamics is nonetheless highly challenging with such data. First, these data have been collected for diagnostic, and not for research purposes. Therefore, they are typically characterized by a large proportion of missing metadata, high variability in sample collection and laboratory procedures, making it challenging to compare and analyse viral load values. In addition, sampling times are skewed, with most individuals being sampled close to symptom onset, and only very little longitudinal follow-up, questioning whether such data can be used to reliably infer on viral dynamics.
In this study, we first evaluated through simulation the feasibility of inferring viral dynamics from sparse and heterogeneous data. We then analysed a large dataset of community-based PCR tests collected in France in symptomatic individuals during the circulation of both Delta and Omicron variants to assess associations between variant, vaccination status, and age and viral load dynamics.
Materials and methods
Data collection
The dataset includes all PCR tests performed in Biogroup laboratories, one of the largest groups of private community laboratories in France, between July 1, 2021, and March 13, 2022 (S1 Fig). This dataset extends a previously described dataset by integrating additional data collected during summer and autumn 2021 [11–13]. Only individuals for which the following information were available were analysed: cycle threshold (Ct) value, sex, age, vaccination status, symptomatic status, time since symptom onset (see below). All these pieces of information (except Ct value) were self-reported.
As these data were not collected for research purposes, we made a number of assumptions to make them suitable for analysis. We considered an individual as infected if at least one PCR test with quantified Ct value was available. All tests with the same variant of infection within 30 days after the first detection were considered to be the same infection. When available, viral sequencing was used to determine the variant of infection; otherwise, we relied on mutation information obtained from the PCR tests (S1-S2 Tables). When the virus genotype was not available, the variant of infection was assumed to be the dominant variant in the population, i.e., representing more than 95% of circulating variants at the time of symptom onset [14,15]. As Beta and Gamma, as well as Omicron BA.1 and BA.2, cannot be distinguished by mutation profiles, we classified variants only as pre-Omicron or Omicron. Notably, the vast majority of pre-Omicron infections corresponded to the Delta variant. The time of symptom onset was reported by intervals (0–1, 2–4, 5–7, 8–14, 15–28, and >28 days) and we used in the model the middle of the interval as the time of onset. Individuals were considered vaccinated upon self-report of receiving at least one dose of the vaccine.
Modelling viral kinetics
We characterized the viral kinetics, from time of infection to viral clearance using a piecewise linear mixed-effects model, which we fitted to cycle thresholds (Fig 1). The time of infection is reconstructed from the incubation period (, estimated from the time of symptom onset. Upon infection, the viral load increases for a duration TP, called the proliferation period, and the viral load reaches its peak noted VP. Then the viral load declines to clearance for a duration noted TC. Let
denote the vector of observations for the individual
, with the
Ct value
observed at time
since symptom onset, described by the following equations:
Schematic representation of viral kinetics from infection to clearance, showing in bold the estimated parameters: incubation period (TI), proliferation period (TP), peak viral load (VP), and clearance period (TC). LOD denotes the limit of detection, and Vinf the viral load at the time of infection.
where represents the structural model, a function of
, and
the vector of individual parameters. Each individual parameter
(with
the parameter index) is composed of
, a fixed effect (or population parameter) shared by all the individuals, and
, the random effect of the parameter
specific to the individual
, which is assumed to follow a lognormal distribution with variance-covariance matrix Ω. We considered a residual error model with σ the standard deviation of the residual error, and
. We set the Ct value at time of infection to 50, which is 10 Ct above the limit of detection.
Simulation study
Scenarios.
We conducted a simulation study to evaluate the performance of parameter estimation in terms of bias, uncertainty, and computational efficiency. A total of 50 datasets were simulated with various population sizes, but keeping constant the ratio of 50% of individuals being infected and 50% being non-infected. The viral kinetics of infected individuals were simulated using the same population parameters: an incubation period () of 5 days, a proliferation period (
) of 6 days, a peak viral load (
) of 25 Ct, and a clearance period (
) of 15 days. The standard deviation of individual random effects (
) was set to 0.15 for all parameters (S2 Fig). The measurements were assumed to be noisy, with a variance of 4 Ct, of the same order of magnitude as that estimated from the real dataset (S3 Fig). The sensitivity of PCR tests is imperfect and decreases as viral load decreases. Thus, modelling all individuals regardless of whether they have a positive PCR test enables the screening of infected individuals who only test negative because they were sampled beyond the window of detectable viral load.
Three scenarios for the sampling times were considered. In the first scenario, we assumed a rich sampling design in a population of 120 individuals (i.e., 60 infected individuals), with daily observations from infection to clearance for each individual. The second scenario assumed a population of 4,000 individuals (i.e., 2,000 infected individuals) and that the sampling times were uniformly distributed during the infection period, i.e., from one week before symptom onset to three weeks after symptom onset. The distribution of the number of tests per individual was similar to that in the real dataset (1 test: 84%, 2 tests: 14%, 3 tests: 2%, 4 and more: < 1%), reflecting limited longitudinal measurements. Restricting the analysis to individuals with at least one positive test captures only 80–85% of all infections. The third scenario assumed a population of 4,000 individuals (i.e., 2,000 infected individuals), and best reflected the data collected in community laboratories, with the distribution of the number of tests and sampling times both similar to those in the real dataset. Thus, in this scenario tests were mainly performed just after the onset of symptoms (see Fig 2). Restricting the analysis to individuals with at least one positive test captures more than 99% of the infections.
Number of positive and negative PCR tests available according to the time since symptom onset, and percentage of PCR tests that were negative. Blue: positive PCR tests; orange: negative PCR tests.
Inference.
In all three scenarios, we evaluated two different modelling approaches. The first approach consisted in modelling only individuals with at least one positive PCR test, assuming that they were all genuinely infected. Thus, we excluded from the modelling all individuals with exclusively negative PCR tests. With this approach, the likelihood of observing can be written as follows:
where is the predicted Ct value under the viral dynamics model,
the density of the normal distribution evaluated at the
th observation of individual
,
, with mean
and observation error
,
the normal cumulative distribution function evaluated at the limit of detection (LOD) [16]. Finally,
and
are the numbers of positive and negative PCR tests for each individual.
The second approach consisted in modelling both individuals with at least one positive PCR test, and individuals with only negative tests, thereby including individuals who may not be infected. In this situation, both the infection status of each individual and the viral kinetic parameters of infected individuals are inferred. The infection status of each individual can be inferred by integrating PCR test results with their timing relative to symptom onset. In this case the likelihood of observing is expressed as the sum of contributions from the cases where the individual is infected vs. not infected:
where is defined in (Eq 2),
denotes a population parameter corresponding to the estimated proportion of infected individuals, and
is defined as follows:
with denotes the probability of true negative, the probability of testing negative given that the individual is not infected, and was fixed to 0.9998 in line with the high specificity of PCR testing [17].
We estimated the model parameters under two different inference frameworks: a Bayesian framework and a frequentist framework. The Bayesian inference relied on the HMC-NUTS algorithm [18] implemented in Stan (via the rstan package, version 2.32.6 [19]), and was applied to both modelling approaches. The frequentist framework used the stochastic approximation expectation–maximization (SAEM) algorithm [20] implemented in Monolix 2023R1 [21], but could only be applied to the first approach (i.e., restricted to infected individuals), because the likelihood accounting for infection status cannot be specified in Monolix. The simulation codes and models are available in the repository [22].
In Stan, we ran four chains in parallel, with random initial value sampled in the priors, with 400 iterations each including 200 warm-up iterations, leading to a posterior sample of 800 replicates. We considered the mean of the posterior distribution as the estimate of the parameter. We used moderately informative priors, using Gaussian distributions centred on the simulated values (S3 Fig). We restricted the analysis to fits with convergent chains, defined by the criterion < 1.05 [23], which compares the between- and within-chain estimates.
Evaluation.
The performance of both frequentist and Bayesian inference, for each modelling approach in each scenario, was evaluated by the relative error of estimation (REE), the coverage rate of the 95% confidence (or credibility) intervals of the parameters, and computation time [24].
The relative error of estimation assesses the accuracy of estimation, and is defined as:
With denoting the estimate of parameter
(posterior mean for Bayesian inference, or maximum likelihood estimate for frequentist inference) in simulation
, and
the true simulated value of
.
Coverage rates were used to evaluate the validity of the 95% confidence (or credibility) intervals, reflecting both the potential bias and the uncertainty of parameter estimates. The coverage rate was calculated from 50 simulations in which the population parameters were identical across all scenarios and datasets. The 95% confidence interval of the coverage rate was computed assuming a binomial distribution with number of trials 50. The nominal target coverage is 0.95, and intervals including this value suggest appropriate uncertainty quantification.
Inference in PCR data in the general community
Only symptomatic individuals with at least one positive PCR test were included in the model, with all data related to their infection (see above). All observations below the LOD were considered as censored. The analysis focused on assessing the influence of variant of infection (Pre-Omicron vs. Omicron), vaccination status (vaccinated vs. unvaccinated), and age (<65 vs. ≥ 65 years) on viral load dynamics. Given the large number of individuals (Table 1), and the potential interaction between these covariates, a composite covariate with eight subgroups, defined by the combinations of these three factors, was constructed. Thus, a single parameter inference was performed, and the population parameters (the fixed effects , the variance of random effects
, and the residual error
) were shared by subgroups. Covariates associations on viral dynamics were illustrated using a forest plot of parameter estimates with their 95% confidence intervals for each subgroup, together with the predicted mean trajectory and its 95% confidence interval. Due to the prohibitive computational time for a dataset of this size (S8 Fig), parameter estimation was performed using the stochastic approximation expectation-maximization (SAEM) algorithm implemented in Monolix 2023R1.
Results
Simulation study
The first simulation scenario considered a rich sampling design (daily Ct values) in a cohort of N = 120 individuals (50% being infected), and, as expected in such a rich setting, all parameters were estimated with good precision and accuracy and coverage of 95% in both Bayesian and frequentist inference frameworks (Figs 3A and 4A).
Boxplots represent the distribution of the relative error of estimation of the population parameters for the incubation period (, in days), the proliferation period (
, in days), the peak viral load (
in Ct), and the clearance period (
, in days). The central line in each box indicates the median, the box represents the interquartile range (IQR; 25%–75%), and the whiskers extend to 1.5 × IQR or to the most extreme values. Values beyond this range are shown as outliers (black dots), while the diamond represents the mean of relative error of estimation (relative bias). Results are given for three scenarios of data generation (see methods). Scenario 1: daily observations from infection to clearance in 120 individuals (50% infected); scenario 2: sparse data uniformly distributed from infection to clearance in 4,000 individuals (50% infected); scenario 3: sparse data mainly after symptom onset in 4,000 individuals (50% infected). Light green: both positive and negative tests for all individuals are included in the Bayesian framework using Stan; dark green: only positive individuals are included in the Bayesian framework using Stan; orange: only positive individuals are included in the frequentist framework using Monolix.
The dots (bars) represent the coverage rate (and 95% confidence interval) of the population parameters for the incubation period (, in days), the proliferation period (
, in days), the peak viral load (
in Ct), and the clearance period (
, in days). Scenario 1: daily observations from infection to clearance in 120 individuals (50% infected); scenario 2: sparse data uniformly distributed from infection to clearance in 4,000 individuals (50% infected); scenario 3: sparse data mainly after symptom onset in 4,000 individuals (50% infected). Light green: both positive and negative tests for all individuals are included in the Bayesian framework using Stan; dark green: only positive individuals are included in the Bayesian framework using Stan; orange: only positive individuals are included in the frequentist framework using Monolix.
We then tested a more realistic scenario with N = 4,000 individuals (50% being infected), with most individuals having only one observation, randomly sampled during the infection period (see methods). The mean relative error was larger than in the first scenario but remained below 15% for all parameters, regardless of the modelling approach and inference framework, and equal to 2 and 5% for peak viral load () and time to clearance (
), respectively, corresponding to an absolute error of 0.5 Ct and 0.8 day, respectively (S5B Fig). Interestingly, the coverage rates were lower than 95% for most parameters, regardless of the modelling approach and inference framework, with the lowest values equal to 25 and 10% for peak viral load and time to clearance respectively, (Fig 4B). Together these findings indicate that with a large population of individuals with random sampling times, all methods tended to underestimate uncertainty, but the parameters of viral kinetics could be accurately estimated, with low absolute errors (S5B Fig).
Finally, the third scenario mimicked the real dataset with N = 4,000 individuals (50% being infected), with most individuals having only one observation, sampled mainly after symptom onset. In spite of the skewed distribution in sampling times, the mean relative error was close to that observed in the second scenario, with values below 20% for all parameters. Errors were larger for the approach restricted to individuals with at least one positive PCR test. The mean relative error did not exceed 2 and 5% for peak viral load () and time to clearance (
), respectively, corresponding to an absolute error of 0.4 Ct and 0.8 day, respectively (S5C Fig). The coverage rates were improved for peak viral load and time to clearance, equal to 80 and 52%, respectively (Fig 4C), thanks to denser sampling post-symptoms. Conversely, as almost no information was available before symptom onset, the coverage rate was largely deteriorated for parameters driving the early viral kinetics, with values equal to 22 and 10% for the incubation period (
) and the proliferation period (
), respectively, using a frequentist framework. The coverage rate was much better when using a Bayesian framework, and a tendency of Monolix to underestimate the standard errors of parameter estimates (S6C Fig). Detailed results on the relative estimation errors for all parameters across the scenarios are presented in the Supplementary Material (S7 Fig).
Finally, while the Bayesian inference framework with Stan provided a better estimation of the uncertainty, it came at a substantial computational cost due to the large sample size and the use of the HMC-NUTS algorithm. In Stan, when modelling the 2,000 individuals with confirmed infection, runtimes were approximately 40 times longer than those in Monolix (3.5 h vs 5.4 mins; S8 Fig). The analysis of the full population of 4,000 individuals was extremely cumbersome, with an average runtime of 14 hours in the third scenario, limiting the use of Bayesian inference for our data. Consequently, we performed parameter estimation on the large community-based PCR dataset using the SAEM algorithm implemented in Monolix 2023R1, which provided reliable estimates of key parameters (peak viral load and clearance period) with limited bias and acceptable computation times.
Inference of the within-host kinetics of infections by pre-Omicron and Omicron variants in the general community
Description of the population.
We then applied our framework to a large real-life dataset of all PCR tests results done in Biogroup community laboratories in France between July 1, 2021, and March 13, 2022. The initial dataset comprised 6,668,557 PCR tests from 4,242,747 individuals—a remarkably large dataset thanks to the very frequent testing in that period in France. After data processing and cleaning (see methods), 383,874 positive tests were analysed, corresponding to 322,218 symptomatic infections, for which age, sex, variant of infection, vaccination status, and time since symptom onset were available (S9 Fig). The Pre-Omicron infection subgroups consisted predominantly of Delta infections, which was by far the dominant circulating variant in France during the summer and autumn of 2021 [14,15].
The largest group involved vaccinated Omicron cases aged <65 years (N = 166,867), followed by unvaccinated Omicron cases aged <65 years (N = 53,349), while the smallest groups were unvaccinated individuals aged ≥65 years with pre-Omicron (N = 2,022) or Omicron (N = 1,613 for Omicron) infections. Most positive tests were within 4 days after symptom onset (Fig 2), and only a small fraction of individuals, 16%, had repeated measurements (Table 1). In the population of individuals with at least one positive test, the virus cleared progressively, with 25 and 62% of individuals turning negative during the second week and the third week after symptom onset respectively (Fig 2).
Estimation of the viral load dynamics parameters.
To analyse the large community dataset, we estimated model parameters using the maximum-likelihood estimator obtained with the SAEM algorithm implemented in Monolix Software 2023R1. Stan was not used because computation times were prohibitively long for a dataset of this size (S6 Fig), and Monolix yielded estimates of the key parameters (peak viral load and clearance period) with minimal bias (Fig 3C).
The model fitted distinct Ct trajectories across the eight subgroups (Table 2). Estimation uncertainty was slightly higher in older age classes, particularly for incubation and proliferation phases, reflecting the smaller sample sizes in these groups.
The clearance period was consistently longer among individuals aged ≥65 years across variants and vaccination groups, by approximately 2–5 days (Fig 5), corresponding to a longer duration of detectable viral load (Fig 6). In pre-Omicron infections, the clearance period among unvaccinated individuals aged ≥65 years was 22.5 days (95% confidence interval: 21.2–23.8), compared with 17.5 days (17.3–17.7) in individuals aged <65 years. A similar pattern was observed for Omicron infections, with clearance of 21.1 days (19.7–22.6) in unvaccinated individuals aged ≥65 years versus 15.9 days (15.7–16.2) in individuals aged <65 years. Differences in peak viral load by age were more limited, with lower Ct values (reflecting higher viral load) in individuals aged ≥65 years, e.g., 16.0 (15.6–16.5) versus 17.6 (17.5–17.7) for Omicron infections in unvaccinated individuals aged ≥65 years versus <65 years, respectively.
Estimated parameters of the viral dynamic model in the different population subgroups. Dots show the mean value, and vertical lines indicate the 95% confidence interval. Red: Pre-Omicron variants; blue: Omicron variants. Solid line: unvaccinated; dashed lines: vaccinated.
Model-based kinetics of the viral load value (in Ct) in the different population subgroups. Lines show the mean predicted Ct trajectory, with shaded areas indicating the 95% confidence interval. Red: Pre-Omicron variants; blue: Omicron variants. Solid line: unvaccinated; dashed lines: vaccinated.
Shorter clearance periods were observed among vaccinated individuals across variants and age groups, by approximately 2–4 days (Fig 5), corresponding to a shorter duration of detectable viral load (Fig 6). In pre-Omicron infections, clearance time was 22.5 days (21.2–23.8) in unvaccinated versus 18.9 days (18.1–19.7) in individuals aged ≥65 years, and 17.5 days (17.3–17.7) versus 14.7 days (14.5–14.9) in vaccinated individuals aged <65 years. In Omicron infections, differences by vaccination status were smaller among individuals aged <65 years, with clearance times of 15.9 days (15.7-16.2) in unvaccinated, and 15.4 days (15.3–15.5) in vaccinated individuals. Larger differences were observed among individuals aged ≥65 years, with clearance times of 21.1 days (19.7-22.6) in unvaccinated and 16.9 days (16.6-17.2) in vaccinated individuals, corresponding to a difference of 4.2 days. Peak viral load was similar across vaccination groups, with differences consistently <0.5 cycle threshold.
Omicron infections were consistently associated with lower peak viral load across vaccination and age groups, by approximately 2–3 Ct (Fig 5) compared with pre-Omicron variants. Pre-Omicron infections showed Ct values around 14–15, e.g., 14.8 (14.7–14.9) in unvaccinated individuals aged <65 years, compared with values around 17–18 in Omicron infections, e.g., 17.6 (17.5–17.7) in the same age group. Clearance period also tended to be shorter in Omicron compared with pre-Omicron infections, with the largest difference observed among vaccinated individuals aged ≥65 years, with a duration of 18.9 days (18.1-19.7) in pre-Omicron infections and 16.9 days (16.6-17.2) in Omicron infections.
Discussion
We showed by simulation that key patterns of viral kinetics can be estimated from community-based testing data with good accuracy and precision. The modelling framework was subsequently applied to a large dataset of PCR results collected in French community laboratories between July 2021 and March 2022, and showed that age > 65 years was associated with differences in viral kinetics parameters, including a longer duration of infection of 2–6 days. Vaccination was associated with shorter clearance periods across variants and age groups, by 2–4 days, but was not associated with differences in peak viral load. Infections with Omicron variants were associated with lower peak viral load across vaccination and age groups, by approximately 2–3 Ct compared with Pre-Omicron variants. Clearance times also tended to be shorter in Omicron compared with pre-Omicron infections, with differences of up to 2 days.
The differences identified here are consistent with findings from previous studies, including age-related variation in peak viral load and clearance duration [25,26], with a lower peak and faster clearance observed in Omicron compared with Delta infections [9,27,28], as well as faster viral clearance among vaccinated individuals infected with pre-Omicron variants [29–31]. Although several studies have used community-based data to model viral dynamics and identify factors shaping these trajectories [11,12,32], our study however accounts for time since symptom onset to date infections and reconstruct viral kinetics. Moreover, it robustly characterizes, using community-based data, the independent associations of Omicron and vaccination on within-host viral dynamics. Our work demonstrates that, beyond their role in epidemic surveillance, routinely collected PCR testing data can be leveraged to uncover factors associated with viral dynamics.
Our simulation study also highlights the challenges of modelling large-scale datasets from community laboratories, where most data are either negative, or positive but measured after symptom onset with no or very few longitudinal data. While modelling both positive and negative tests in a Bayesian inference framework provided unbiased estimates, this approach was computationally cumbersome making it not practical for large datasets (S8 Fig). In addition, Bayesian inference can be sensitive to prior specifications; however, in our case, posterior distributions appeared only weakly influenced by the priors (S10 Fig). We here decided to use a pragmatic strategy, focusing on modelling confirmed infections using a frequentist inference framework. The model simulations show that this choice comes at the cost of a reduced accuracy and precision for parameter estimates, in particular for the pre-symptomatic period, however the magnitude of these effects was limited for the parameters driving the post-symptomatic period, i.e., peak viral load and time to viral clearance (S5, S6C Figs). Further our model assumes a Ct value of 50 at the time of infection. To assess the robustness of this assumption, we conducted a sensitivity analysis using an alternative model that did not rely on it. The results were consistent with those of the main analysis (S3 Table). The model can also be used in other settings with richer dataset or to assess differences across treatment groups in clinical trials. For instance, we performed an external validation using data from the PLATCOV trial [33], comparing viral load in untreated vs remdesivir-treated individuals. Consistent with previously published results [33], the model identified a significantly shorter clearance phase in the remdesivir arm, with estimated viral clearance half-lives consistent with those reported in the original study (S4 Table, S11–S12 Figs).
Beyond methodological issues, the main limitation of these data is that they were not collected for research purposes and may be associated with various limiting or confounding factors. For instance, the (self-reported) occurrence or the period of symptom onset, which is used to reconstruct viral trajectories, is prone to recall bias (S9 Fig) [34]. This uncertainty was not accounted for in our model and the data of symptom onset was imputed at the median of declared interval of symptom onset. Although most individuals declared an interval for symptom onset of one day, a Bayesian inference framework could naturally incorporate this uncertainty in future work. Although our selected population did not appear to differ from the overall dataset based on observed characteristics, restricting the analysis to individuals with complete information, in particular symptom and vaccination status, may have induced selection bias (S5 Table). More generally, data were collected during a period where the vast majority of PCR tests were fully reimbursed by the national health insurance [35], limiting financial barriers to testing. Nevertheless, the motivation for undergoing testing have evolved over time, particularly in the context of the French “passe sanitaire” implemented during the study period, which may have resulted in differential testing behaviour between vaccinated and unvaccinated individuals. This policy may also partly explain the high prevalence of unvaccinated individuals in our dataset (S13 Fig). Moreover, several sources of residual confounding may persist. First, disease severity could act as an unmeasured confounder, as information on symptom type and intensity was not available in our dataset. Similarly, no clinical data were available, while presence of risk factors (e.g., chronic disease) may impact both the use of tests and viral load. In addition, information on the time since last vaccine dose was rarely recorded, preventing us from accounting for potential waning immunity effects. Furthermore, during the study period, reimbursed rapid antigen tests were widely available in pharmacies, raising the possibility that individuals seeking PCR testing may not be representative of the general population. Another factor limiting the generalization of our findings to the overall population is that the analysis focused exclusively on symptomatic infections, which exclude the 55% of infections in our dataset with no declared symptoms. Of these individuals, 33% corresponded to individuals who consistently reported being asymptomatic at all tests, and 22% to those with unknown or missing symptom status at all tests. The proportion of asymptomatic cases is consistent with prevalence estimates of around 40% reported in other studies [36]. Note that the Ct distribution of symptomatic individuals was shifted toward lower values (S14A Fig) when compared to asymptomatic individuals (S14B Fig), indicating potential differences in their underlying viral kinetics or the timing of their test. Finally, we also investigated potential differences according to temporal trends, irrespective of circulating variant or immune background [37]. To explore this, we examined the estimated random effects of peak viral load and viral clearance duration according to the date of testing (S15 Fig), but did not observe any temporal pattern.
To conclude, our study contributes to the broader effort of expanding methodological tools for analysing viral dynamics during epidemics. Such community-based laboratory data are increasingly available for several acute respiratory infections, particularly with the widespread adoption of multiplex PCR assays, which allow the systematic detection of multiple viruses from a single sample [38]. Because these data are collected routinely and continuously, they represent a rich source of information at low cost. Leveraging them with appropriate statistical tools could help identify vulnerable subpopulations, monitor the impact of a new variant, or evaluate vaccine effectiveness in a population broadly representative of the general population. To support the wider use of community laboratory datasets, we recommend establishing clear frameworks for data collection, data sharing with research teams, and privacy protection.
Supporting information
S1 Fig. Distribution of dates of PCR tests in the community dataset.
https://doi.org/10.1371/journal.pcbi.1013811.s001
(TIFF)
S1 Table. Mutation and variants correspondence.
https://doi.org/10.1371/journal.pcbi.1013811.s002
(DOCX)
S2 Table. Rule for attributing mutation detection to the infection variant.
https://doi.org/10.1371/journal.pcbi.1013811.s003
(DOCX)
S2 Fig. Simulated individual parameters from the simulation study.
https://doi.org/10.1371/journal.pcbi.1013811.s004
(TIFF)
S3 Fig. Examples of simulated datasets in each scenario.
https://doi.org/10.1371/journal.pcbi.1013811.s005
(TIFF)
S4 Fig. Prior for the Bayesian inference framework.
https://doi.org/10.1371/journal.pcbi.1013811.s006
(TIFF)
S5 Fig. Absolute error of estimation of parameters.
https://doi.org/10.1371/journal.pcbi.1013811.s007
(TIFF)
S6 Fig. Estimates and associated uncertainty of population parameters.
https://doi.org/10.1371/journal.pcbi.1013811.s008
(TIFF)
S7 Fig. Relative error of estimation of parameters.
https://doi.org/10.1371/journal.pcbi.1013811.s009
(TIFF)
S8 Fig. Computation times in the simulation study.
https://doi.org/10.1371/journal.pcbi.1013811.s010
(TIFF)
S3 Table. Estimated parameters in the model without hypothesis on viral load at the time of infection.
https://doi.org/10.1371/journal.pcbi.1013811.s013
(DOCX)
S4 Table. Estimated parameters for patients in the “no study drug” and “remdesivir” arms of the PLATCOV trial.
https://doi.org/10.1371/journal.pcbi.1013811.s014
(DOCX)
S11 Fig. Distribution of cycle threshold values according to symptomatic status declared at the time of testing.
https://doi.org/10.1371/journal.pcbi.1013811.s015
(TIFF)
S12 Fig. Individual estimate of the viral clearance half-lives.
https://doi.org/10.1371/journal.pcbi.1013811.s016
(TIFF)
S5 Table. Age and sex characteristics by presence or absence of vaccination status information.
https://doi.org/10.1371/journal.pcbi.1013811.s017
(DOCX)
S13 Fig. Proportion of vaccinated individuals by week of testing by age category.
https://doi.org/10.1371/journal.pcbi.1013811.s018
(TIFF)
S14 Fig. Distribution of cycle threshold values according to symptomatic status declared at the time of testing.
https://doi.org/10.1371/journal.pcbi.1013811.s019
(TIFF)
S15 Fig. Distribution of the estimated individual random effects of peak viral load and clearance period according to the date of testing.
https://doi.org/10.1371/journal.pcbi.1013811.s020
(TIFF)
Acknowledgments
We thank Bob Carpenter and Charles Margossian (Flatarion Institute) for valuable discussions on modelling large datasets in Stan and on accounting for infection status in a Bayesian framework. We also thank the reviewers for their constructive comments, which contributed to improving the manuscript.
References
- 1. Hay JA, Kennedy-Shaffer L, Kanjilal S, Lennon NJ, Gabriel SB, Lipsitch M, et al. Estimating epidemiologic dynamics from cross-sectional viral load distributions. Science. 2021;373(6552).
- 2. Kirwan PD, Charlett A, Birrell P, Elgohari S, Hope R, Mandal S, et al. Trends in COVID-19 hospital outcomes in England before and after vaccine introduction, a cohort study. Nat Commun. 2022;13(1):4834. pmid:35977938
- 3. Watson OJ, Barnsley G, Toor J, Hogan AB, Winskill P, Ghani AC. Global impact of the first year of COVID-19 vaccination: a mathematical modelling study. Lancet Infect Dis. 2022;22(9):1293–302. pmid:35753318
- 4. Allen H, Vusirikala A, Flannagan J, Twohig KA, Zaidi A, Chudasama D, et al. Household transmission of COVID-19 cases associated with SARS-CoV-2 delta variant (B.1.617.2): national case-control study. The Lancet Regional Health – Europe. 2022;12. pmid:34729548
- 5. Lyngse FP, Mortensen LH, Denwood MJ, Christiansen LE, Møller CH, Skov RL, et al. Household transmission of the SARS-CoV-2 Omicron variant in Denmark. Nat Commun. 2022;13(1):5573. pmid:36151099
- 6. Thompson MG, Burgess JL, Naleway AL, Tyner H, Yoon SK, Meece J. Prevention and attenuation of Covid-19 with the BNT162b2 and mRNA-1273 vaccines. New England Journal of Medicine. 2021;385(4):320–9.
- 7. Woodbridge Y, Amit S, Huppert A, Kopelman NM. Viral load dynamics of SARS-CoV-2 Delta and Omicron variants following multiple vaccine doses and previous infection. Nat Commun. 2022;13(1):6706. pmid:36344489
- 8. Puhach O, Meyer B, Eckerle I. SARS-CoV-2 viral load and shedding kinetics. Nat Rev Microbiol. 2023;21(3):147–61. pmid:36460930
- 9. Yang Y, Guo L, Yuan J, Xu Z, Gu Y, Zhang J. Viral and antibody dynamics of acute infection with SARS-CoV-2 omicron variant (B.1.1.529): a prospective cohort study from Shenzhen, China. The Lancet Microbe. 2023;4(8):e632–41. pmid:37459867
- 10. Lunt R, Quinot C, Kirsebom F, Andrews N, Skarnes C, Letley L, et al. The impact of vaccination and SARS-CoV-2 variants on the virological response to SARS-CoV-2 infections during the Alpha, Delta, and Omicron waves in England. J Infect. 2024;88(1):21–9. pmid:37926118
- 11. Blanquart F, Abad C, Ambroise J, Bernard M, Cosentino G, Giannoli JM, et al. Characterisation of vaccine breakthrough infections of SARS-CoV-2 Delta and Alpha variants and within-host viral load dynamics in the community, France, June to July 2021. Eurosurveillance. 2021;26(37).
- 12. Cosentino G, Bernard M, Ambroise J, Giannoli J-M, Guedj J, Débarre F, et al. SARS-CoV-2 viral dynamics in infections with Alpha and Beta variants of concern in the French community. J Infect. 2022;84(1):94–118. pmid:34329672
- 13. Blanquart F, Abad C, Ambroise J, Bernard M, Débarre F, Giannoli J-M, et al. Temporal, age, and geographical variation in vaccine efficacy against infection by the Delta and Omicron variants in the community in France, December 2021 to March 2022. Int J Infect Dis. 2023;133:89–96. pmid:37182550
- 14.
Hodcroft EB. CoVariants: SARS-CoV-2 Mutations and Variants of Interest. https://covariants.org/ 2021. Accessed 2023 October 1.
- 15.
Données de laboratoires pour le dépistage (A COMPTER DU 18/05/2022) - SI-DEP. https://www.data.gouv.fr/datasets/donnees-de-laboratoires-pour-le-depistage-a-compter-du-18-05-2022-si-dep Accessed 2023 October 1.
- 16. Jacqmin-Gadda H, Thiébaut R, Chêne G, Commenges D. Analysis of left-censored longitudinal data with application to viral load in HIV infection. Biostatistics. 2000;1(4):355–68. pmid:12933561
- 17. Snedden CE, Lloyd-Smith JO. Predicting the presence of infectious virus from PCR data: A meta-analysis of SARS-CoV-2 in non-human primates. PLoS Pathog. 2024;20(4):e1012171. pmid:38683864
- 18. Hoffman MD, Gelman A. The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo. arXiv. 2014. http://arxiv.org/abs/1111.4246
- 19.
RStan: the R interface to Stan. https://mc-stan.org/ 2025.
- 20. Delyon B, Lavielle M, Moulines E. Convergence of a stochastic approximation version of the EM algorithm. Ann Statist. 1999;27(1).
- 21.
Antony. Monolix 2023R1, Simulations Plus. https://doi.org/10.5281/zenodo.11401936
- 22.
Beaulieu M. Code and resources for simulation. https://github.com/Germante1/Code_SARS_CoV_viral_dynamics_in_community 2025.
- 23.
Gelman A, Carlin JB, Stern HS, Dunson DB, Vehtari A, Rubin DB. Bayesian data analysis. 3rd ed. London: Taylor and Francis. 2014.
- 24. Morris TP, White IR, Crowther MJ. Using simulation studies to evaluate statistical methods. Statistics in Medicine. 2019;38(11):2074–102.
- 25. Jones TC, Biele G, Mühlemann B, Veith T, Schneider J, Beheim-Schwarzbach J. Estimating infectiousness throughout SARS-CoV-2 infection course. Science. 2021;373(6551):eabi5273.
- 26. Néant N, Lingas G, Le Hingrat Q, Ghosn J, Engelmann I, Lepiller Q. Modeling SARS-CoV-2 viral kinetics and association with mortality in hospitalized patients from the French COVID cohort. Proceedings of the National Academy of Sciences. 2021;118(8):e2017962118.
- 27. Hay JA, Kissler SM, Fauver JR, Mack C, Tai CG, Samant RM, et al. Quantifying the impact of immune history and variant on SARS-CoV-2 viral kinetics and infection rebound: A retrospective cohort study. eLife. 2022.
- 28. Fryer HR, Golubchik T, Hall M, Fraser C, Hinch R, Ferretti L, et al. Viral burden is associated with age, vaccination, and viral variant in a population-representative study of SARS-CoV-2 that accounts for time-since-infection-related sampling bias. PLoS Pathog. 2023;19(8):e1011461. pmid:37578971
- 29. Kissler SM, Fauver JR, Mack C, Tai CG, Breban MI, Watkins AE, et al. Viral Dynamics of SARS-CoV-2 Variants in Vaccinated and Unvaccinated Persons. N Engl J Med. 2021;385(26):2489–91. pmid:34941024
- 30. Garcia-Knight M, Anglin K, Tassetto M, Lu S, Zhang A, Goldberg SA, et al. Infectious viral shedding of SARS-CoV-2 Delta following vaccination: A longitudinal cohort study. PLoS Pathog. 2022;18(9):e1010802. pmid:36095030
- 31. Singanayagam A, Hakki S, Dunning J, Madon KJ, Crone MA, Koycheva A, et al. Community transmission and viral load kinetics of the SARS-CoV-2 delta (B.1.617.2) variant in vaccinated and unvaccinated individuals in the UK: a prospective, longitudinal, cohort study. The Lancet Infectious Diseases. 2022;22(2):183–95. pmid:34756186
- 32. Elie B, Roquebert B, Sofonea MT, Trombert-Paolantoni S, Foulongne V, Guedj J, et al. Variant-specific SARS-CoV-2 within-host kinetics. J Med Virol. 2022;94(8):3625–33. pmid:35373851
- 33. Jittamala P, Schilling WHK, Watson JA, Luvira V, Siripoon T, Ngamprasertchai T, et al. Clinical antiviral efficacy of remdesivir in coronavirus disease 2019: an open-label, randomized controlled adaptive platform trial (PLATCOV). The Journal of Infectious Diseases. 2023;228(10):1318–25.
- 34. Coughlin SS. Recall bias in epidemiologic studies. J Clin Epidemiol. 1990;43(1):87–91. pmid:2319285
- 35.
info.gouv.fr. Fin de la gratuité systématique des tests Covid-19. https://www.info.gouv.fr/actualite/fin-de-la-gratuite-systematique-des-tests-covid-19 2021.
- 36. Buitrago-Garcia D, Egli-Gany D, Counotte MJ, Hossmann S, Imeri H, Ipekci AM, et al. Occurrence and transmission potential of asymptomatic and presymptomatic SARS-CoV-2 infections: A living systematic review and meta-analysis. PLoS Med. 2020;17(9):e1003346. pmid:32960881
- 37. Wongnak P, Schilling WHK, Jittamala P, Boyd S, Luvira V, Siripoon T, et al. Temporal changes in SARS-CoV-2 clearance kinetics and the optimal design of antiviral pharmacodynamic studies: an individual patient data meta-analysis of a randomised, controlled, adaptive platform study (PLATCOV). Lancet Infect Dis. 2024;24(9):953–63. pmid:38677300
- 38. Elnifro EM, Ashshi AM, Cooper RJ, Klapper PE. Multiplex PCR: optimization and application in diagnostic virology. Clin Microbiol Rev. 2000;13(4):559–70. pmid:11023957