This is an uncorrected proof.
Figures
Abstract
The Mediterranean population of tomato leaf curl New Delhi virus (ToLCNDV-ES) is characterized by a high genetic uniformity, distinguishing it from its Asian counterparts. ToLCNDV-ES is thought to have a monophyletic origin, likely resulting from a single recombination event, prior to its spread throughout the Mediterranean region. Following its first detection in southeastern France in 2020, ToLCNDV-ES re-emerged in France in 2022. Our analysis based on advanced long-read sequencing, circular DNA profiling, and phylogeny indicates both local persistence of French ToLCNDV-ES and multiple independent introduction events. Signatures of positive selection were identified in French ToLCNDV-ES populations, whereas no clear evidence of recombination was found. Bayesian time-structured phylogenetic analyses suggest that introductions in France occurred between 2018 and 2021 from the major ToLCNDV-ES clade, while several Italian ToLCNDV-ES isolates diverged prior to the virus introduction in the Mediterranean basin. Overall, this study demonstrates the value of an optimized long-read sequencing approach for resolving circular DNA virus diversity, and sheds light on the complex evolutionary history of ToLCNDV-ES in the Mediterranean Basin, particularly in southeastern France.
Author summary
Understanding how emerging plant viruses invade new regions and evolve after introduction is a central challenge in pathogen biology. Here, we investigate the emergence and spread of the Mediterranean strain of tomato leaf curl New Delhi virus (ToLCNDV-ES) in southeastern France following its first detection in 2020 and subsequent re-emergence in 2022. While ToLCNDV-ES populations across the Mediterranean basin have previously been reported to exhibit remarkable genetic uniformity, the processes underlying their persistence and dissemination remain poorly understood. Using an advanced long-read sequencing workflow combined with circular DNA profiling and phylogenetic analyses, we characterize the evolutionary dynamics of ToLCNDV-ES populations at unprecedented resolution. Our analyses reveal evidence for both local persistence of viral populations and multiple independent introduction events into France, highlighting the complex invasion dynamics of this emerging begomovirus. We detect signatures of positive selection in French ToLCNDV-ES populations. Our analyses suggest that introductions into France likely occurred between 2018 and 2021 from the major Mediterranean ToLCNDV-ES lineage. Moreover, our study demonstrates the value of an advanced long-read sequencing strategy for studying circular DNA viruses, building upon a framework previously described. By enabling accurate reconstruction of full viral genomes and population diversity, this approach provides a powerful tool to investigate the emergence, spread, and evolution of plant viruses. We believe that this methodological advance, combined with our analysis of ToLCNDV-ES introduction dynamics, offers insights that are broadly relevant for understanding the evolutionary trajectories of emerging viral pathogens.
Citation: Patthamapornsirikul A, Golyaev V, Verdin E, Goillon C, Boolell H, Dierickx S, et al. (2026) Novel insights into tomato leaf curl New Delhi virus introduction and evolution in Southeastern France using an advanced long-read sequencing workflow. PLoS Pathog 22(8): e1014463. https://doi.org/10.1371/journal.ppat.1014463
Editor: Adi Stern, Tel Aviv University, ISRAEL
Received: March 6, 2026; Accepted: July 12, 2026; Published: August 7, 2026
Copyright: © 2026 Patthamapornsirikul et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: The viral genome sequences used in this study have been deposited in the NCBI GenBank database and the accession number can be found in supplementary materials. The sequences will be released in July 2026 (according to the GenBank staff, July 12) Oxford Nanopore sequencing data generated in this study have now been deposited in the European Nucleotide Archive (ENA) under study accession PRJEB114816 (https://www.ebi.ac.uk/ena/browser/view/PRJEB114816).
Funding: This work was supported by the Virtigation project (virtigation.eu) funded by European Union’s Horizon 2020 research and innovation programme under grant agreement No 101000570 all authors and AP doctoral scholarship was funded by Franco-Thai program co-operated by French embassy in Thailand and Campus France. The funders had no role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: We have read the journal’s policy, and the authors of this manuscript declare the following competing interests. Wim Dumon and Koen Deforche are the founders of Emweb (https://www.emweb.be/). Sam Dierickx and Sara Cleemput are the employees of Emweb. Emweb is the author of Genome Detective (https://www.genomedetective.com/). The remaining authors declare that no competing interests exist.
1. Introduction
Southern France represents an important agricultural hub for vegetable and horticultural crop production, particularly cucurbit crops such as courgette, melon, and cucumber, which are cultivated in either open fields or greenhouses. This economically important sector is regularly threatened by a wide range of plant viruses, predominantly aphid-transmitted species [6]. These endemic viral agents are not only widespread but also exhibit complex molecular population structures characterized by high mutation rates, frequent genetic recombination, and repeated introductions leading to the emergence of novel variants [Desbiez et al. 2009, 8,31,6,19]. In addition to the endemic aphid-transmitted viruses, tomato leaf curl New Delhi virus (ToLCNDV) was first reported in southeastern France in 2020 [7] and has since become more frequent in the region.
ToLCNDV (Begomovirus solanumdelhiense), a bipartite begomovirus transmitted by the whitefly Bemisia tabaci [58,43], has emerged over the last decades as one of the most significant plant viruses in Asia and the Mediterranean Basin. First identified in the Indian subcontinent, Asian ToLCNDV isolates display a broad host range, infecting major crops notably within the Solanaceae (tomato, pepper, potato), Malvaceae (cotton and okra), and Cucurbitaceae (zucchini, melon, cucumber, and squash) [58,57,43]. In the Mediterranean Basin, ToLCNDV was first detected in Spain in 2012 and subsequently reported to spread throughout Europe and the wider Mediterranean region [26,58,14,7]. The ToLCNDV strain found in the Mediterranean—commonly referred to as ToLCNDV-ES—has a more restricted host range than Asian isolates, primarily infecting cucurbit species under natural conditions [16,49,43]. ToLCNDV-ES exhibits a relatively homogenous molecular population, distinct from the Asian isolates, and is thought to have arisen through interspecific recombination [16,39]. Its low genetic diversity suggests that the virus was introduced into the Mediterranean region through a single event, followed by rapid spread across the Basin [49,25,55,43]. Despite the overall genetic uniformity, the virus shows evidence of ongoing evolution through the accumulation of mutations and occasional recombination events [25,37].
As in other plant viruses, selection acting on the viral diversity generated by genetic reassortment, mutations, and/or recombination may enhance the adaptability of the virus to diverse environmental conditions and host species, supporting its persistence and spread [8], and potentially enabling it to overcome resistance and evade control strategies [32,54]. For an emerging pathogen such as ToLCNDV, investigating molecular diversity using robust analytical tools is essential for understanding viral ecology, outbreak dynamics and mechanisms of emergence, as well as for developing sustainable management strategies. Advances in molecular techniques, particularly high-throughput sequencing (HTS), have enabled detailed characterization of viral genomes by uncovering fine-scale genetic variations and population structure. HTS-based analyses of viral genomes allow in-depth profiling of inter- and intra-species mixed infections, detection of recombination events, and assessment of within-host diversity, far beyond the capabilities of first-generation sequencing techniques [48,29]. However, accurate genome reconstruction from short-read HTS data can be challenging when viral populations contain multiple closely related variants, potentially leading to artefactual recombinants. This limitation can be overcome by third-generation long-read sequencing approaches. Accordingly, Oxford Nanopore technology (ONT) has been shown to provide cost-effective and high-throughput solutions for accurate reconstruction of virus variants [21]. The major limitation of ONT is its relatively high per-read error rate. However, this can be substantially reduced for circular DNA viruses by performing rolling circle amplification (RCA) prior to sequencing, followed by reconstruction of high-quality consensus genome sequences using established bioinformatic tools [56,21].
To elucidate the evolutionary dynamics of ToLCNDV following its emergence in southern France, we applied a high-throughput sequencing workflow based on ONT sequencing and the online tool Genome Detective Platform [5,21] to perform whole-genome analyses of both ToLCNDV DNA-A and DNA-B components. These analyses allowed us to investigate phylogenetic relationships, infer introduction patterns, identify potential recombination events, and detect signatures of natural selection, thereby providing new insights into the mechanisms shaping the molecular evolution of ToLCNDV.
2. Results
2.1. Nanopore sequencing accurately reconstructs full-length French ToLCNDV genomes, as confirmed by Sanger sequencing
ToLCNDV was first identified in autumn 2020 in six infected zucchini plants, sampled at three different sites located in the Bouches-du-Rhône and Gard départements. The virus re-emerged in 2022 at three sites in Bouches-du-Rhône, near the original 2020 outbreak area. An increase in viral detection was observed in 2023 and 2024, during which 41 additional samples were collected within a 60 km radius across the départements of Gard, Bouches-du-Rhône, and Vaucluse (Fig 1). Most infected samples were obtained from zucchini, whereas a limited number were collected from melon, cucumber and the weed species Ecballium elaterium and Datura stramonium (S1 Table). No ToLCNDV infections were detected in watermelon by ELISA and routine PCR [47,7].
Administrative boundaries are based on Natural Earth country and cultural vector data, with France and subnational boundaries used for spatial reference. Basemap data were obtained from the Natural Earth dataset All spatial analyses and figure generation were conducted in R 4.5.3 using the packages sf, ggplot2, dplyr, ggrepel, rnaturalearth, and rnaturalearthdata. Natural Earth datasets are public domain and freely available for use under the terms described at: https://www.naturalearthdata.com/about/terms-of-use/. Base layer data (country boundaries and cultural vectors) are available here: https://www.naturalearthdata.com/downloads/10m-cultural-vectors/.
Nanopore sequencing data were generated for 48 French isolates and for the Spanish reference isolate ES13–35 using the pipeline described in [21]. The number of full-length contigs per isolate ranged from 282 to 92,789 (mean = 14,517). The DNA-A/DNA-B ratio among full-length contigs varied from 0.03 to 1.26. For each sample, the most abundant haplotype for each DNA represented, on average, 76% of the reads for DNA-A and 70% of the reads for DNA-B. However, in six samples for DNA-A and 12 samples for DNA-B, the most frequent haplotype accounted for less than 50% of the reads, while one or two additional haplotypes each represented more than 20% of the reads (S1 Fig).
Within each sample, haplotypes could differ by only one or two single nucleotides polymorphisms (SNPs), or by as many as 30 mutations for DNA-A and more than 70 mutations for DNA-B, indicating that they represent different molecular subgroups. The four most abundant haplotypes in each sample together represented 96.3% of the reads for DNA-A and 90% of the reads for DNA-B (S1 Fig).
Samples from the weeds Datura stramonium and Ecballium elaterium tended to harbor more complex viral populations (S1 Fig), but the number of samples available for each species was too limited to assess whether this represents a general trend associated with long-term infection in these hosts.
Among the major haplotypes in each sample, three isolates for DNA-A and seven isolates for DNA-B contained variants that belonged to molecular subgroups different from that of the major haplotype. One isolate for DNA-A and six isolates for DNA-B were considered to clearly represent multiple infections. In the three remaining cases, the total number of reads obtained for the sample was limited, resulting in only 30–50 reads being assigned to the minor haplotype (S1 Fig). Including negative controls in all the ONT sequencing pools and/or performing technical replicas of the sequences would have allowed to be more accurate in the determination of threshold for very minor variants, even if it had no strong impact on the overall results of the analyses.
For isolate VI231143, low-frequency variants were detected for both DNA-A and DNA-B. For isolate EV231242, the evidence for mixed infection was clear for DNA-B but ambiguous for DNA-A. Thus, in both cases, mixed infections are likely to be genuine for both genomic components, but the sequences of the minor variants were not included in the phylogenetic analyses.
Within each sample, no attempt was made to associate specific DNA-A and DNA-B haplotypes based on their relative frequencies in the ONT sequencing dataset.
The Sanger sequences were 100% identical to the major haplotype obtained by Nanopore sequencing for the reference isolate ES13–35 for both ToLCNDV DNA-A and DNA-B components. Sanger sequences derived from cloned amplicons were either identical to the dominant haplotype identified by Nanopore sequencing of the original uncloned sample or differed by one or two SNPs. Most of these SNPs were unique among the Mediterranean isolates, suggesting that they likely resulted from amplification or cloning artefacts. Partial Sanger sequencing of uncloned amplicons further confirmed identity with the major Nanopore-derived haplotypes or revealed ambiguous positions corresponding to SNPs distinguishing major and minor haplotypes, thereby reflecting the underlying viral population structure within the host plant (S2 Table). Additionally, for the mixed-infected isolates VI231145 (for both DNA components) and VI231242 (only for DNA-B) (S1 Table, S1 Fig), partial Sanger sequencing chromatograms showed multiple peaks at the positions corresponding to the expected SNPs, as well as a frameshift corresponding to the expected indel in DNA-B of VI231242 (S1 Fig).
Genome organization of the French ToLCNDV isolates was consistent with that described for ToLCNDV, comprising seven ORFs in DNA-A and two in DNA-B. Predicted protein lengths were overall conserved relative to other ToLCNDV-ES isolates, with only a few deviations that were validated by Sanger resequencing. For the AC1 (Rep) protein, four isolates (FR416-1, FR240005, VI231143 and VI231240) contained an additional initiation codon at positions 2671–2673, located 90 nt upstream of the canonical AC1 start codon at positions 2581–2583. This feature would potentially result in a 30-amino-acid N-terminal extension. The same upstream ATG was detected in one Spanish isolate (accession MH577727).
Two French isolates originating from the same site (EV240029 and EV240030) contained a premature stop codon within the AC1 coding region, resulting in a Rep protein truncated by nine amino acids at its C-terminal end. In isolate EV240030, a minor haplotype carrying the premature stop codon (EV240030-0-4, accession PX842865) co-occurred with a dominant haplotype lacking the stop codon (EV240030-0-3, accession PX842864) (S2 Table). The biological significance of these AC1 length variations remains unknown.
Within the putative AC5 ORF, whose function remains unknown for ToLCNDV [43], several isolates contained a premature stop codon at position 26. This mutation suggests that a truncated protein of 120 amino acids, initiated from a downstream start codon, may be expressed instead of the full-length 161-amino-acid form.
2.2. Limited recombination and evidence of site-specific positive selection in Mediterranean ToLCNDV populations
Using GARD to analyse recombination in the Mediterranean ToLCNDV isolates, a single breakpoint was detected in DNA-B (position 2684, corresponding to Rep cleavage site), although with weak support (Δc-AIC = 290.007 compared with the null model, and Δc-AIC = 774.977 compared with a single-tree multiple-partition model). RDP4 identified a limited number of putative recombination events in both DNA-A and DNA-B. For DNA-A, event 1 involved MF967015.1 (Tunisia) as the major parent and VI240544 (France) as the minor parent (p-values: GENECONV = 0.1202, BootScan = 0.00606, SiScan = 0.0196). For DNA-B, event 1 involved MH577630.1 (Morocco) and MF967022.1 (Tunisia) (p-values: BootScan = 0.02485, SiScan = 0.000595). However, parental assignments were inconsistent across methods, and statistical support remained weak.
Split decomposition analysis using SplitsTree showed no strong evidence of recombination. Phi tests yielded non-significant results for DNA-A (p = 0.066; p = 0.47 for French isolates only) and DNA-B (p = 0.17; p = 0.49 for French isolates) as all p-values were above the 0.05 threshold. In the split decomposition DNA-A network, sequence VI240545 appeared as a potential outlier (S2 Fig). When RDP4 was applied specifically to VI240545 DNA-A together with a subset of 16 French isolates from subgroups A1 and A2, a recombination signal was detected by the MaxChi, Chimaera, Siscan and 3Seq methods (p = 1.9 10-3 to p = 1.6 10-2). Although VI240545 may therefore represent a recombinant, the limited overall genetic diversity of DNA-A reduces the strength and interpretability of this signal.
FUBAR analysis detected evidence of positive selection at a single site among the 256 codons of the coat protein (CP) in French and Mediterranean ToLCNDV isolates (posterior probability 0.960). This site corresponds to amino acid position 147, where most Mediterranean isolates carried a glycine, whereas all French isolates encoded a serine, specified by either AGC (subgroup A1) or TCA (subgroup A2) (Fig 2). Codon 147 was also identified as positively selected by MEME (p = 0.002). Among non-French Mediterranean isolates, only a single Italian isolate (ZU-Le18, accession OQ262954) encoded a serine at this position. Notably, this mutation also affects the partially overlapping putative AC5 ORF, which is encoded in antisense orientation relative to CP. In subgroup A2 isolates and in the Italian isolate ZU-Le18, a TGA stop codon (corresponding to the TCA in the CP frame) occurs at codon 26 of AC5. This likely results in expression of a truncated 120-amino-acid protein initiated from a downstream start codon at position 42, instead of the full-length 161-amino-acid form, as previously reported for ZU-Le18 [37].
The scale bar represents genetic distances inferred under a GTR substitution model. Branch colors indicate selection regimes, with green and orange denoting evidence of positive selection (dN/dS > 1) and blue indicating dN/dS ≤ 1. The color gradient was quantified using Empirical Bayers Factors (EBF) as estimated by the MEME (Mixed Effects Model of Evolution) method (www.datamonkey.org).
2.3. Distinct diversification of DNA-A and DNA-B components reveals reassortment dynamics in French ToLCNDV
Phylogenetic analysis of French ToLCNDV isolates from 2020 to 2024 confirmed that all sequences clustered within the Mediterranean clade for both DNA-A and DNA-B (Fig 3). DNA-A sequences showed limited genetic variation, sharing 98–99% nucleotide identity compared with other Mediterranean isolates. DNA-B exhibited slightly greater divergence, with 97–98% identity relative to Mediterranean ToLCNDV sequences.
The Italian sub-clade is highlighted in green. French lineages are shown in distinct colors, with two subgroups identified for DNA-A (A1 in purple and A2 in dark red) and four subgroups for DNA-B (B1 in purple, B2 in dark red, B3 in light blue, and B4 in red). Bootstrap values (>60%; n = 1,000 bootstraps) are indicated at the corresponding nodes. Spanish isolates are represented by yellow branches, other Italian isolates by light green branches, and isolates from North Africa and Middle East (Algeria, Tunisia, Morocco and Jordan) by orange branches.
Within the French dataset, the two genomic components displayed distinct phylogenetic structuring. DNA-A sequences segregated into two well-supported subgroups, A1 and A2, with low intragroup diversity (0.0026 for A1 and 0.0017 for A2) and moderate intergroup divergence (0.0103). In contrast, DNA-B sequences were more heterogeneous and partitioned into four distinct phylogenetic subgroups (B1-B4). Intragroup diversity ranged from 0.0015 (B4) to 0.0050 (B1), whereas intergroup divergence ranged from 0.020 to 0.024.
Subgroups A1 and A2, as well as B1 and B2, were already present in 2020. Subgroup B3 was first detected in 2022, and B4 in 2023, indicating ongoing diversification of the DNA-B component. Multiple combinations of DNA-A and DNA-B subgroups were observed. The most frequent associations in plants infected with only one molecular subgroup for each DNA were A1 + B1, A1 + B3, A1 + B4 and A2 + B2, although additional combinations occurred sporadically (Fig 4). Among the six (or seven with VI231143 isolate) mixed infections identified in 2023–2024, B3 was detected in association with B1, B2 or B4, and with A1, A2 or both (S1 Table).
Bootstrap values (n = 1000 bootstraps) are indicated at the corresponding nodes. Coloured lines indicate the associations between DNA-A and DNA-B component sequences for each isolate: yellow (A1 + B1), grey (A1 + B2), blue (A1 + B3), purple (A1 + B4), black (A2 + B1), green (A2 + B2), red (A2 + B3). For the six isolates presenting mixed infections, all viral components detected within the same isolate were considered to be associated.
2.4. Time-structured phylogeographic analysis reveals the introduction timeline of ToLCNDV in southern France
The combination of a Bayesian skyline tree prior and an uncorrelated relaxed clock model provided the best fit for the Mediterranean ToLCNDV dataset for both DNA-A and DNA-B. The mean substitution rate was estimated at 3.8 x 10-4 substitutions/site/year for DNA-A (95% highest posterior density (HPD): 2.9 x 10-4 – 4.5 x 10-4) and 8 x 10-4 substitutions/site/year for DNA-B (95% HPD: 6.3 x 10-4 - 9.5 x 10-4).
Time-scaled maximum clade credibility trees for both the DNA-A and the DNA-B genomic component revealed two Mediterranean clades: a minor clade containing six isolates from southern Italy and a major clade comprising more than 100 isolates from across the Mediterranean region, including all French sequences (Fig 5). The divergence between these two clades was dated to 1997–2001 for both segments, with overlapping 95% HPD (1988–2006 for DNA-A and 1994–2008 for DNA-B), i.e., more than a decade before the first reported detection of ToLCNDV in the Mediterranean Basin.
Analyses were performed using an uncorrelated relaxed molecular clock model with a Bayesian skyline tree prior. Nodes are shaded in grayscale according to posterior probability of clade support values. Branches are color-coded by geographic origin. French subgroups are highlighted with boxes. Estimated node ages and their associated 95% highest posterior density (HPD) intervals are indicated at the corresponding nodes.
Within the major Mediterranean clade, diversification was estimated to have begun around 2007–2008 (95% HPD 2006–2009 for DNA-A and 2006–2010 for DNA-B), whereas it was dated to 2011–2012 for the Italian minor clade (95% HPD 2009–2014 for DNA-A and 2010–2013 for DNA-B).
French isolates did not form a monophyletic group, supporting multiple introductions from the major Mediterranean clade. Divergence among French subgroups was dated to 2018–2019, for both DNA-A and DNA-B, except for subgroup B4, first detected in 2023, for which diversification was estimated at 2021–2022, albeit based on a limited number of samples. The ancestral nodes at the base of the French subgroups were dated between 2010 and 2013 (with overlapping 95% HPD intervals) close to the initial detection of ToLCNDV in the Mediterranean Basin (Spain, 2012), but long before its first detection in France.
Using a partial dataset that either completely excluded the Spanish isolates, or retained only 20 of them, had no impact on the clustering of the remaining isolates, and only a very limited effect on node dating (S3 Fig). The mean substitution rate remained essentially unchanged for DNA-A (4.1 x 10-4 substitutions/site/year, 95% HPD 2.9 x 10-4 – 4.5 x 10-4), but appeared lower for DNA-B when the Spanish isolates were completely excluded (6.9 x 10-4) or when only 20 Spanish isolates were retained (6.7 x 10-4 substitutions/site/year) compared with the estimate obtained using the full dataset. The corresponding 95% HPD intervals were 5.2 x 10-4 – 8.5 x 10-4 and 5.3 x 10-4 – 8.0 x 10-4, respectively. Using two different subsets of 20 Spanish isolates did not modify the result of the “partial” dataset. Nevertheless, the estimated substitution rate for DNA-B remained higher than that of DNA-A. Thus, the overrepresentation of Spanish isolates, including 53 isolates sampled in 2016, moderately affected the estimation of the substitution rate for DNA-B, but had little to no effect on that for DNA-A.
3. Discussion
This study applied a high-throughput Nanopore based circular DNA virus sequencing pipeline [21] to characterize ToLCNDV field isolates and investigate the evolution of Mediterranean ToLCNDV populations, since the first detection of the virus in France in 2020. Virus genome sequences reconstructed using the pipeline were fully consistent with Sanger data for the validated isolates, confirming the robustness and accuracy of the approach. The pipeline enabled reconstruction of full-length sequence variants for both ToLCNDV genomic components, providing direct insight into complex intra-host viral populations without the need for laborious subcloning, while reducing the risk of artefactual recombination associated with short read technologies [38,21].
Our results revealed substantial diversity among French ToLCNDV isolates, with two subgroups identified for DNA-A and four for DNA-B. The lack of monophyly among French subgroups supports multiple independent virus introductions into France [11]. Nevertheless, several subgroups were consistently detected between 2020 and 2024 in the same geographical areas, suggesting local persistence of the virus, possibly sustained by weed reservoirs. Indeed, ToLCNDV was detected in weed species growing around infected crops, including previously reported hosts [25], such as jimsonweed (Datura) and squirting cucumber (Ecballium) [25]. Viral populations characterized from weed samples clustered within the same subgroups as those infecting cultivated hosts. As a perennial cucurbit widely distributed across the Mediterranean Basin, Ecballium could represent a potential overwintering reservoir for ToLCNDV. However, ToLCNDV transmission by Bemisia tabaci between Ecballium and crops has been reported to be inefficient in Spain [15]. One potential bias of the study is that most sampling, except for weeds, was symptom-based. As a result, early infections and mildly symptomatic plants may have been overlooked, potentially leading to an underestimation of ToLCNDV prevalence and a biased view of its genetic diversity [51]. However, at least one systematic sampling was conducted independently of symptom expression. Both symptomatic and asymptomatic plants, the latter likely corresponding to recent infections, were sequenced, and no obvious differences were observed in the viral populations they contained.
Molecular analysis estimated divergence among French ToLCNDV subgroups between 2010 and 2013 for both DNA-A and DNA-B, i.e., at least six years prior to the first detection in France, and coinciding with ToLCNDV emergence in Spain. Diversification within the French subgroups was estimated at 2018–2019, likely corresponding to the timing of their introduction into France. Notably, subgroup B4 appears to have diverged more recently (2021–2022) and, although based on a limited sampling, we suggest that it may represent an additional, more recent introduction.
Available sequence data suggest a Spanish origin for the French isolates. However, this interpretation should be treated with caution, given the strong bias toward Spanish isolates detected between 2012 and 2016, which constitute the majority of publicly available ToLCNDV-ES genomic sequences (Fig 5; S3 Table). French isolates clearly differed from the minor Mediterranean clade containing isolates from southern Italy [42]. Notably, this Italian clade appears to have diverged in the early 2000s, more than a decade before the first detection of ToLCNDV in the Mediterranean Basin. This pattern suggests that isolates from Southern Italy have a distinct evolutionary trajectory and possibly an independent introduction.
For the major Mediterranean clade of ToLCNDV-ES, evolutionary analyses support a single introduction event around 2007–2008, with a convergent evolutionary history for DNA-A and DNA-B. The pathways by which ToLCNDV-ES was introduced into the Mediterranean Basin and subsequently into France remain unresolved. Seed transmission has been demonstrated in zucchini for the Italian subclade isolates [27], whereas Spanish isolates have been reported to be seedborne but not seed-transmitted in several cucurbit species [17,50].
Although DNA-A and DNA-B of ToLCNDV-ES appear to share a similar timing of introduction into the Mediterranean Basin, DNA-B exhibited greater genetic diversity with an estimated substitution rate approximately twice that of DNA-A (3.8 x 10-4 substitutions/site/year for DNA-A vs. 8 x 10-4 substitutions/site/year for DNA-B) even though the estimation for DNA-B was likely affected by the imbalanced dataset. These evolutionary rates fall within the range reported for other begomoviruses [12,33].
Recent estimates of ToLCNDV nucleotide substitution rates were approximately 7 x 10‒4 substitutions/nt/yr for DNA-A and 8.1 to 8.7 x 10-4 substitutions/nt/yr for DNA-B [23]. The higher substitution rate reported for DNA-A compared to our estimates may reflect the consideration of both Asian and Mediterranean ToLCNDV sequences in that study. Given that several ToLCNDV isolates and strains, including ToLCNDV-ES, have been described as intra- or interspecific recombinants for DNA-A [39], incorporation of these recombinant sequences may inflate substitution rate estimates. Notably, evolutionary rate estimates for DNA-B were consistent across studies.
For ToLCNDV-ES, the lower substitution rate observed for DNA-A relative to DNA-B may be explained by the presence of overlapping ORFs in DNA-A. Such genomic constraints increase the likelihood of non-synonymous mutations and are therefore subject to stronger purifying selection [36]. However, at the intra-host level, higher mutation frequencies in DNA-A than in DNA-B have been reported for another begomovirus [44], suggesting that selective and evolutionary dynamics may differ across biological scales.
In our study, overlapping regions among DNA-A ORFs were not explicitly accounted for in the estimation of dN/dS ratios. A single codon in the CP was identified as being under positive selection, with French isolates differing from most other ToLCNDV-ES sequences at this position. This pattern likely reflects, at least in part, the strong geographic structuring of genetic variation rather than selection alone, as begomovirus populations are often geographically structured [46], and recombination and demographic history can confound codon-level selection signals [12], but the fact that the S147 in the CP of French isolates is encoded by different codons AGC or TCA suggests a convergent evolution related, at least in part, to selection. The mutation also affects the putative AC5 protein in antisense orientation, resulting in a stop codon yielding a shorter protein for isolates presenting a TCA in the CP. Since all French isolates have the same S147 in their CP, the S147 CP mutation could explain potential functional or biological differences between the French isolates and the other ToLCNDV-ES, but differences among French isolates would probably be related to an impact of the mutation on AC5 and/or DNA structure rather than on the CP. Preliminary studies have not revealed any notable differences in host range or vector transmission between the French and Spanish ToLCNDV isolates tested so far. Further studies are needed to determine the possible functional or biological effects of the mutation, and whether they are mostly due to the changes in CP or in AC5. Therefore, its interpretation should be approached with caution given the constraints imposed by overlapping reading frames.
The high-throughput sequencing and analytical framework employed in this study enabled the characterization of mixed infections involving multiple DNA-A and/or DNA-B subgroups within individual plants, including both crops and weeds. Such complex intra-host virus populations may facilitate reassortment and recombination. Although several reassortment events were detected, it remains unclear whether these events occurred prior to or following introduction into France.
In contrast, recombination appeared rare among French ToLCNDV isolates. Only a single potential recombination event was detected in DNA-A, and its support was limited due to the low phylogenetic signal in the dataset. No clear recombination events were identified in DNA-B, despite sufficient intergroup diversity to allow their detection. Even in mixed-infected plants, long-read sequencing did not reveal low-frequency recombinant variants arising during virus replication.
This recombination inactivity contrasts with the reported recombination-prone nature of begomoviruses [18,2]. One possible explanation is that co-infections in France are relatively recent, limiting the time available for recombinant forms to emerge and establish. In perennial weed hosts such as Ecballium, where viruses may likely persist for multiple years, long-term viral co-infections may lead to the emergence of recombinants with enhanced fitness, as reported for other begomoviruses [1]. It is likely that recombination detection was influenced by methodological limitations, such as the number and geographic distribution of the sampled fields, rather than reflecting a true biological absence of recombination.
The multiyear persistence of distinct molecular subgroups of ToLCNDV-ES in France suggests that the virus is now established in the region, potentially facilitated by its maintenance in perennial weed reservoirs. A deeper understanding of the long-term evolutionary dynamics of ToLCNDV-ES in weed reservoirs and cultivated hosts, its potential to generate novel recombinants, and the relative fitness of its subgroups, alone or in mixed infections, will be critical for preventing further geographic expansion of the virus.
4. Materials and methods
4.1. Plant samples collection
French ToLCNDV isolates were obtained through field sampling in collaboration with private breeders, farm advisers, agricultural extension services, and official quarantine authorities. Samples were collected between 2020 and 2024 from cultivated plants, including zucchini, melon, and cucumber, as well as from weed species such as Datura stramonium (jimsonweed) and Ecballium elaterium (squirting cucumber), across the French départements (counties) of Gard, Bouches-du-Rhône, and Vaucluse (Fig 1; S1 Table). Collection was based on symptomatic plants, complemented by random sampling of asymptomatic crops and weed species located in the vicinity of infecting fields. In 2024, a systematic sampling of 30 zucchini plants was conducted within a field trial independently of symptom expression. The samples were tested by ELISA for the presence of ToLCNDV and other viruses, and five ToLCNDV-positive samples were subsequently selected for ONT sequencing (isolates VI240303-VI240328; S1 Table).
4.2. DNA extraction and ONT sequencing
The presence of ToLCNDV in the collected samples was confirmed by DAS-ELISA with a commercial ToLCNDV antiserum (Agdia EMEA, Soisy sur Seine, FR), following the manufacturer’s recommendations.
Total DNA was extracted as previously described [20]. Isolated DNA was amplified and submitted to long read ONT sequencing following the established protocol [38,21]. Briefly, extracted total DNA was submitted to circular DNA enrichment via RCA, debranched using NEBNext FFPE DNA Repair and NEBNext Ultra II And Repair/dA-Tailing modules. DNA clean-up was subsequently performed using AMPure XP Beads (Beckman Coulter, USA) before ligating adapters with the Nanopore Ligation Sequencing Kit v14, following the manufacturer’s protocol. Next, the DNA was used to prepare libraries for ONT long read sequencing. Ten to twelve samples, without negative controls, were barcoded using the Oxford Nanopore Native Barcoding kit v14, pooled, adapter ligated with the Nanopore Ligation Sequencing Kit v14 and loaded into a Nanopore flow-cell (R10.4.1). Sequencing runs were performed for 48-72h. Raw sequencing data were basecalled using Dorado (model: dna_r10.4.1_e8.2_400bps_sup@v5.2.0) (https://github.com/nanoporetech/dorado), followed by demultiplexing and trimming using Porechop (https://github.com/rrwick/Porechop). Processed data were uploaded to the Genome Detective platform [56] for virus variant analysis, providing for each sample a list of haplotypes represented by at least 10 reads. Variants supported by fewer than ten sequences were excluded from the analysis, as indels occur frequently in homopolymeric regions in these sequences. Minor haplotype sequences supported by at least ten sequences were manually inspected for obvious sequencing errors in homopolymeric regions, which were corrected when identified.
4.3. Virus cloning and Sanger sequencing
Four isolates from 2020 and four isolates from 2022 (S1 Table) were cloned using RCA products digested with BamHI and ligated into a BamHI-linearized pBluescript (KS)+ vector (Strategene, La Jolla, CA). After transformation of Escherichia coli (strain DH5α), the cloned viral DNAs were amplified with universal primers (M13-21 and M13rev) and submitted for sanger sequencing (GenoScreen, France) with M13-21, M13rev, and internal ToLCNDV-A-1000F and ToLCNDV-B-1700-F primers [7; S4 Table]. The chromatograms were analysed using Chromas (https://technelysium.com.au/wp/chromas/), and viral full-length sequences were reconstructed. Eight samples collected in 2023 were partially amplified with virus-specific primers (S4 Table) and subsequently sequenced (GenoScreen, France). These sequences were compared with those obtained by ONT long-read sequencing.
4.4. Full-length sequence dataset and Multiple Sequence Alignment workflow
Complete sequences of 50 French ToLCNDV isolates collected between 2020 and 2024 (S1 Table), together with 110 Mediterranean ToLCNDV isolates retrieved from GenBank (S3 Table), were included in the analyses. 48 French isolates, as well the Spanish reference isolate ES13–35, were sequenced using ONT long-read sequencing.
In addition to 48 isolates sequenced by ONT, two French isolates (22V0044 and CD20001) were included in the analyses. Isolate 22V0044 was sequenced by Sanger sequencing only, while isolate CD20001, corresponding to the first ToLCNDV isolate detected in France [7] had previously been sequenced by both Sanger sequencing (accessions MW310624-MW310625 for DNA-A and DNA-B, respectively) and Illumina short read sequencing.
Hereafter, the datasets comprising 50 and 160 isolates will be referred to as the “French” and “complete” datasets, respectively.
For French isolates with mixed infections including different molecular subgroups, the major haplotype representing each subgroup was retained for analysis. Separate analyses such as genetic selection and recombination were performed using both the “complete” and “French” datasets.
For large datasets, full-length DNA-A (2,740 nt) and DNA-B (2,684 nt) sequences were aligned using MAFFT (Multiple Alignment using Fast Fourier Transform algorithm) with default parameters implemented in Job Dispatcher platform (https://ebi.ac.uk/jdispatcher/) [34]. Smaller datasets were aligned using the MUSCLE algorithm [13].
4.5. Recombination and selection analysis
Preliminary screening for recombination events was done on the complete (160 sequences) and French (50 isolates, yielding 48 different sequences for DNA-A and 51 different sequences for DNA-B) datasets using the Genetic Algorithm for Recombination Detection (GARD), implemented on the Datamonkey server (https://www.datamonkey.org/). The same datasets were subsequently analysed with RDP4 [35], using seven independent detection methods: RDP, GENECONV, Bootscan, Siscan, 3Seq, Chimaera, and Maxchi, with p < 0.01 significance threshold. Recombination breakpoints identified by GARD and confirmed by at least three independent RDP4 analyses were retained. Potential recombinations among ToLCNDV sequences were independently assessed using SplitsTree (version 6.5.1) [22] to visualize phylogenetic conflicts [4]. The significance of recombination signals was validated using the Phi test [3].
Seven open reading frames (ORFs) for DNA-A and two for DNA-B were obtained from all ToLCNDV sequences and validated using the NCBI ORFfinder (https://www.ncbi.nlm.nih.gov/orffinder/). The coding sequences were aligned with the MUSCLE algorithm implemented in MEGA 7 [28] and subsequently analysed for signatures of positive selection, defined by a higher rate of nonsynonymous (dN) than synonymous (dS) substitutions, using the HyPhy-based DataMonkey platform (https://www.datamonkey.org/). Episodic diversifying selection was assessed with MEME (Mixed Effects Model of Evolution) [41]. Pervasive selection was evaluated using FUBAR (Fast, Unconstrained Bayesian AppRoximation) [40]. The thresholds were set at p < 0.01 for MEME and posterior probability > 95% for FUBAR.
4.6. Phylogenetic analysis
Maximum Likelihood phylogenetic trees were constructed with MEGA 7 after selecting the appropriate substitution model [28] for the aligned sequences. For DNA-A sequences, the Tamura 3-parameter model with a discrete Gamma distribution (T92 + G) and 1000 bootstrap replicates was applied. For DNA-B sequences, the same model was used with the additional assumption of a proportion of invariant sites (T92 + G + I).
4.7. Time-structured phylogeographic analysis
Phylogeographic analysis was conducted using the BEAST v1.10.4 software package to infer the spatiotemporal dynamics of viral spread [52]. Full-length ToLCNDV genomic sequences were aligned with MAFFT and subsequently processed with BEAUTi [10], with sampling dates and geographic origins incorporated as discrete traits. The HKY nucleotide substitution model was selected with estimated base frequencies, and a gamma distribution (4 categories) was applied. The temporal signal was evaluated with TempEst [45]. Two molecular clock parameters (strict molecular clock vs. uncorrelated relaxed clock) and three tree priors (coalescent constant size, exponential growth and Bayesian skyline) were tested. MCMC chains were run for 10 million iterations, and outputs were assessed with Tracer to ensure proper convergence. If the convergence was not obtained after 10M runs, the analysis was performed again with 20M runs, and 100M runs for DNA-B. After discarding 10% burn-in, a maximum clade credibility (MCC) tree was generated using TreeAnnotator and then visualized using FigTree v1.4.4 (https://tree.bio.ed.ac.uk/software/figtree/).
To assess the impact of the strong overrepresentation of Spanish isolates in the dataset, BEAST analyses were performed on two partial DNA-A and DNA-B datasets: (1) dataset completely excluding sequences from Spain 2012–2016, and (2) a dataset retaining 20 Spanish sequences, limited to 3–4 randomly selected sequences per year (2012–2016). In both cases, analyses were conducted using an uncorrelated relaxed molecular clock, a Bayesian skyline prior, and 10 million MCMC iterations.
Supporting information
S1 Table. French ToLCNDV isolates sequenced in the study, indicating their major haplotypes and corresponding molecular subgroups.
Abbreviations: B. du Rhône, Bouches-du-Rhône; St Remy de P, Saint-Rémy-de-Provence.
https://doi.org/10.1371/journal.ppat.1014463.s001
(XLSX)
S2 Table. Comparison of Sanger and Nanopore sequencing results for French ToLCNDV isolates.
https://doi.org/10.1371/journal.ppat.1014463.s002
(XLSX)
S3 Table. Mediterranean ToLCNDV isolates (excluding French isolates) used for comparative and phylogenetic analyses in the study.
https://doi.org/10.1371/journal.ppat.1014463.s003
(XLSX)
S4 Table. Primers used for amplification and sequencing of ToLCNDV isolates.
https://doi.org/10.1371/journal.ppat.1014463.s004
(XLSX)
S1 Fig. Relative frequencies (%) of the major haplotypes obtained for DNA-A (top) and DNA-B (middle) in 48 French ToLCNDV isolates after ONT sequencing.
For each sample, the four most abundant haplotypes are displayed according to their relative frequencies, while all remaining minor haplotypes are grouped in white. Histograms outlined in red correspond to haplotypes belonging to a molecular subgroup different from the major haplotype detected in the sample. Asterisks above the histograms indicate samples considered to show strong evidence (asterisks without brackets) or limited evidence (asterisks in bracket) of mixed infection involving different molecular subgroups. Samples collected from weeds are indicated by a triangle for Datura stramonium and a square for Ecballium elaterium. Bottom: chromatogram obtained after Sanger sequencing of DNA-B of isolate VI231242 with primer ToLCNDV-B-1700-F, showing double peaks corresponding to a mixed infection with molecular subgroups B2 and B3, and the double sequence after position 510 (2235 in the genome) due to the deletion of one “G” in sequences from subgroup B3.
https://doi.org/10.1371/journal.ppat.1014463.s005
(TIF)
S2 Fig. Split decomposition network produced with SplitsTree for DNA-A (top) and DNA-B (bottom) sequences of French ToLCNDV isolates.
Molecular subgroups (A1-A2 for DNA-A; B1-B4 for DNA-B) are indicated. The boxed alignment shows parsimony-informative sites distinguishing DNA-A sequence VI240545-0-1 from isolates belonging to subgroups A1 (brown) and A2 (green). The positions of nucleotide substitutions within the DNA-A genome are indicated above the alignment. Dots denote nucleotides identical to VI240545-0-1.
https://doi.org/10.1371/journal.ppat.1014463.s006
(TIF)
S3 Fig. Time-scaled phylogenetic trees produced with BEAST for full-length DNA-A (top) and DNA-B (bottom) sequences of ToLCNDV-ES isolates collected in the Mediterranean Basin between 2012 and 2024.
Left: dataset excluding the Spanish isolates from 2012-2016. Right: dataset including 20 Spanish isolates (3–4 per year each between 2012 and 2016). Analyses were performed using an uncorrelated relaxed molecular clock model with a Bayesian skyline tree prior. The viral genome sequences used in this study have been deposited in the NCBI GenBank database under accession numbers PX842845-PX842948 (S1 Table). Oxford Nanopore sequencing data generated in this study have been deposited in the European Nucleotide Archive (ENA) under study accession PRJEB114816.
https://doi.org/10.1371/journal.ppat.1014463.s007
(TIF)
Acknowledgments
We thank Dr. Pascal Gentit (Head of Bacteriology, Virology & GMO unit, Plant Health Laboratory, ANSES) for providing viral material from 2020 and 2022 batches as well as partial CP sequences from 2020 isolates.
References
- 1. Belabess Z, Peterschmitt M, Granier M, Tahiri A, Blenzar A, Urbino C. The non-canonical tomato yellow leaf curl virus recombinant that displaced its parental viruses in southern Morocco exhibits a high selective advantage in experimental conditions. J Gen Virol. 2016;97(12):3433–45. pmid:27902403
- 2. Bonnamy M, Blanc S, Michalakis Y. Replication mechanisms of circular ssDNA plant viruses and their potential implication in viral gene expression regulation. mBio. 2023;14(5):e0169223. pmid:37695133
- 3. Bruen TC, Philippe H, Bryant D. A simple and robust statistical test for detecting the presence of recombination. Genetics. 2006;172(4):2665–81. pmid:16489234
- 4. Bryant D, Huson DH. NeighborNet: improved algorithms and implementation. Front Bioinform. 2023;3:1178600. pmid:37799982
- 5. Cleemput S, Dumon W, Fonseca V, Abdool Karim W, Giovanetti M, Alcantara LC, et al. Genome Detective Coronavirus Typing Tool for rapid identification and characterization of novel coronavirus genomes. Bioinformatics. 2020;36(11):3552–5. pmid:32108862
- 6. Desbiez C. The never-ending story of cucurbits and viruses. Acta Hortic. 2020;(1294):173–92.
- 7. Desbiez C, Gentit P, Cousseau‐Suhard P. First report of Tomato leaf curl New Delhi virus infecting courgette in France. New Disease Reports. 2021;43.
- 8. Desbiez C, Moury B, Lecoq H. The hallmarks of “green” viruses: do plant viruses evolve differently from the others?. Infect Genet Evol. 2011;11(5):812–24. pmid:21382520
- 9. Desbiez C, Wipf-Scheibel C, Millot P, Berthier K, Girardot G, Gognalons P, et al. Distribution and evolution of the major viruses infecting cucurbitaceous and solanaceous crops in the French Mediterranean area. Virus Res. 2020;286:198042. pmid:32504705
- 10. Drummond AJ, Suchard MA, Xie D, Rambaut A. Bayesian phylogenetics with BEAUti and the BEAST 1.7. Mol Biol Evol. 2012;29(8):1969–73. pmid:22367748
- 11. Duffy S, Seah YM. 98% identical, 100% wrong: per cent nucleotide identity can lead plant virus epidemiology astray. Philos Trans R Soc Lond B Biol Sci. 2010;365(1548):1891–7. pmid:20478884
- 12. Duffy S, Shackelton LA, Holmes EC. Rates of evolutionary change in viruses: patterns and determinants. Nat Rev Genet. 2008;9(4):267–76. pmid:18319742
- 13. Edgar RC. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 2004;32(5):1792–7. pmid:15034147
- 14. European Food Safety Authority (EFSA), van Gemert J, Schenk M, Candresse T, Bottex B, Delbianco A, et al. Pest survey card on tomato leaf curl New Delhi virus. EFS3. 2020;17(7).
- 15. Farina A, Rapisarda C, Fiallo-Olivé E, Navas-Castillo J. Tomato leaf curl New Delhi virus spain strain is not transmitted by Trialeurodes vaporariorum and is inefficiently transmitted by Bemisia tabaci mediterranean between zucchini and the wild cucurbit Ecballium elaterium. Insects. 2023;14(4):384. pmid:37103199
- 16. Fortes IM, Sánchez-Campos S, Fiallo-Olivé E, Díaz-Pendón JA, Navas-Castillo J, Moriones E. A novel strain of tomato leaf curl New Delhi virus has spread to the mediterranean basin. Viruses. 2016;8(11):307. pmid:27834936
- 17. Fortes IM, Pérez-Padilla V, Romero-Rodríguez B, et al. Begomovirus tomato leaf curl new delhi virus is seedborne but not seed transmitted in melon. Plant Disease. 2023;107:473–9.
- 18. García-Andrés S, Accotto GP, Navas-Castillo J, Moriones E. Founder effect, plant host, and recombination shape the emergent population of begomoviruses that cause the tomato yellow leaf curl disease in the Mediterranean basin. Virology. 2007;359(2):302–12. pmid:17070885
- 19. Gaudin J, Piry S, Wipf-Scheibel C. Outbreak of Cucumber Mosaic Virus Subgroup IB in Pepper from the Espelette Area (Basque Country, Southwestern France) and First Report of Five Taxa as Natural Hosts of CMV. Plant Disease. 2025;109:983–7.
- 20. Gilbertson RL, Rojas MR, Russell DR, Maxwell DP. Use of the asymmetric polymerase chain reaction and DNA sequencing to determine genetic variability of bean golden mosaic geminivirus in the Dominican Republic. J Gen Virol. 1991;72 (Pt 11):2843–8. pmid:1940873
- 21. Golyaev V, Dierickx S, Deforche K, Dumon W, Vanderschuren H. A method for in-depth analysis of circular DNA virus populations by unambiguously profiling the low abundant virus variants and partial genomic components. Nucleic Acids Res. 2025;53(6):gkaf221. pmid:40173013
- 22. Huson DH, Bryant D. Application of phylogenetic networks in evolutionary studies. Mol Biol Evol. 2006;23(2):254–67. pmid:16221896
- 23. Iqbal Z, Alshoaibi A, Al Hashedi SA, Sattar MN, Ramadan KMA. Deciphering the genetic landscape of tomato leaf curl New Delhi virus: Dynamic and region-specific diversity revealed by comprehensive sequence analyses. PLoS One. 2025;20(7):e0326349. pmid:40674325
- 24. Joannon B, Lavigne C, Lecoq H, Desbiez C. Barriers to gene flow between emerging populations of Watermelon mosaic virus in Southeastern France. Phytopathology. 2010;100(12):1373–9. pmid:20879843
- 25. Juárez M, Rabadán MP, Martínez LD, et al. Natural hosts and genetic diversity of the emerging tomato leaf curl new delhi virus in spain. Frontiers in Microbiology. 2019;10:140.
- 26. Juárez M, Tovar R, Fiallo-Olivé E, Aranda MA, Gosálvez B, Castillo P, et al. First detection of tomato leaf curl New Delhi virus Infecting Zucchini in Spain. Plant Dis. 2014;98(6):857. pmid:30708660
- 27. Kil EJ, Vo TTB, Fadhila C. Seed transmission of tomato leaf curl New Delhi virus from zucchini squash in italy. Plants. 2020;9(5):563.
- 28. Kumar S, Stecher G, Tamura K. MEGA7: Molecular evolutionary genetics analysis version 7.0 for bigger datasets. Mol Biol Evol. 2016;33(7):1870–4. pmid:27004904
- 29. Kutnjak D, Tamisier L, Adams I, Boonham N, Candresse T, Chiumenti M, et al. A primer on the analysis of high-throughput sequencing data for detection of plant viruses. Microorganisms. 2021;9(4):841. pmid:33920047
- 30. Lecoq H, Fabre F, Joannon B, Wipf-Scheibel C, Chandeysson C, Schoeny A, et al. Search for factors involved in the rapid shift in Watermelon mosaic virus (WMV) populations in South-eastern France. Virus Res. 2011;159(2):115–23. pmid:21605606
- 31. Lecoq H, Wipf-Scheibel C, Nozeran K. Comparative molecular epidemiology provides new insights into Zucchini yellow mosaic virus occurrence in France. Virus Research. 2014;186:135–43.
- 32. Lefeuvre P, Martin DP, Elena SF, Shepherd DN, Roumagnac P, Varsani A. Evolution and ecology of plant viruses. Nat Rev Microbiol. 2019;17(10):632–44. pmid:31312033
- 33. Long X, Zhang S, Shen J, Du Z, Gao F. Phylogeography and evolutionary dynamics of tobacco curly shoot virus. Viruses. 2024;16(12):1850. pmid:39772160
- 34. Madeira F, Madhusoodanan N, Lee J, Eusebi A, Niewielska A, Tivey ARN, et al. The EMBL-EBI Job Dispatcher sequence analysis tools framework in 2024. Nucleic Acids Res. 2024;52(W1):W521–5. pmid:38597606
- 35. Martin DP, Murrell B, Golden M, Khoosal A, Muhire B. RDP4: Detection and analysis of recombination patterns in virus genomes. Virus Evol. 2015;1(1):vev003. pmid:27774277
- 36. Martín-Hernández I, Pagán I. Gene overlapping as a modulator of begomovirus evolution. Microorganisms. 2022;10(2):366. pmid:35208820
- 37. Mastrochirico M, Spanò R, De Miccolis Angelini RM, Mascia T. Molecular characterization of a recombinant isolate of tomato leaf curl New Delhi virus associated with severe outbreaks in Zucchini Squash in Southern Italy. Plants (Basel). 2023;12(13):2399. pmid:37446959
- 38. Mehta D, Cornet L, Hirsch-Hoffmann M, Zaidi SS-E-A, Vanderschuren H. Full-length sequencing of circular DNA viruses and extrachromosomal circular DNA using CIDER-Seq. Nat Protoc. 2020;15(5):1673–89. pmid:32246135
- 39. Moriones E, Praveen S, Chakraborty S. Tomato leaf curl New Delhi virus: An emerging virus complex threatening vegetable and fiber crops. Viruses. 2017;9(10):264.
- 40. Murrell B, Moola S, Mabona A, Weighill T, Sheward D, Kosakovsky Pond SL, et al. FUBAR: a fast, unconstrained bayesian approximation for inferring selection. Mol Biol Evol. 2013;30(5):1196–205. pmid:23420840
- 41. Murrell B, Wertheim JO, Moola S, Weighill T, Scheffler K, Kosakovsky Pond SL. Detecting individual sites subject to episodic diversifying selection. PLoS Genet. 2012;8(7):e1002764. pmid:22807683
- 42. Panno S, Caruso AG, Troiano E, Luigi M, Manglli A, Vatrano T, et al. Emergence of tomato leaf curl New Delhi virus in Italy: estimation of incidence and genetic diversity. Plant Pathology. 2019;68(3):601–8.
- 43. Patthamapornsirikul A, Desbiez C, Siriwan W, Verdin E. A brief review on the emerging tomato leaf curl New Delhi virus (ToLCNDV) infecting cucurbits in the Mediterranean Basin. Eur J Plant Pathol. 2025;174(4):579–94.
- 44. Pinto VB, Quadros AFF, Godinho MT, Silva JC, Alfenas-Zerbini P, Zerbini FM. Intra-host evolution of the ssDNA virus tomato severe rugose virus (ToSRV). Virus Res. 2021;292:198234. pmid:33232784
- 45. Rambaut A, Lam TT, Max Carvalho L, Pybus OG. Exploring the temporal structure of heterochronous sequences using TempEst (formerly Path-O-Gen). Virus Evol. 2016;2(1):vew007. pmid:27774300
- 46. Rocha CS, Castillo-Urquiza GP, Lima ATM, et al. Brazilian begomovirus populations are highly recombinant, rapidly evolving, and segregated based on geographical location. Journal of Virology. 2013;87:105784–99.
- 47. Romay G, Pitrat M, Lecoq H, et al. Resistance against melon chlorotic mosaic virus and tomato leaf curl New Delhi virus in melon. Plant Disease. 2019;103:2913–9.
- 48. Rubio L, Galipienso L, Ferriol I. Detection of plant viruses and disease management: relevance of genetic diversity and evolution. Front Plant Sci. 2020;11:1092. pmid:32765569
- 49. Ruiz L, Simon A, Velasco L, Janssen D. Biological characterization of Tomato leaf curl New Delhi virus from Spain. Plant Pathology. 2016;66(3):376–82.
- 50. Sáez C, Kheireddine A, García A, Sifres A, Moreno A, Font-San-Ambrosio MI, et al. Further molecular diagnosis determines lack of evidence for real seed transmission of tomato leaf curl New Delhi virus in cucurbits. Plants (Basel). 2023;12(21):3773. pmid:37960129
- 51.
Stobbe A, Roossinck MJ. Plant virus diversity and evolution. Current Research Topics in Plant Virology. Springer International Publishing. 2016. 197–215. https://doi.org/10.1007/978-3-319-32919-2_8
- 52. Suchard MA, Lemey P, Baele G, Ayres DL, Drummond AJ, Rambaut A. Bayesian phylogenetic and phylodynamic data integration using BEAST 1.10. Virus Evol. 2018;4(1):vey016. pmid:29942656
- 53. Tepfer M, Girardot G, Fénéant L, Ben Tamarzizt H, Verdin E, Moury B, et al. A genetically novel, narrow-host-range isolate of cucumber mosaic virus (CMV) from rosemary. Arch Virol. 2016;161(7):2013–7. pmid:27138549
- 54. Torralba B, Blanc S, Michalakis Y. Reassortments in single-stranded DNA multipartite viruses: confronting expectations based on molecular constraints with field observations. Virus Evol. 2024;10(1):veae010. pmid:38384786
- 55. Troiano E, Parrella G. First report of tomato leaf curl New Delhi virus in Lagenaria siceraria var. longissima in Italy. Phytopathol Mediterr. 2023;62(1):17–24.
- 56. Vilsker M, Moosa Y, Nooij S, Fonseca V, Ghysens Y, Dumon K, et al. Genome Detective: an automated system for virus identification from high-throughput sequencing data. Bioinformatics. 2019;35(5):871–3. pmid:30124794
- 57. Vo TTB, Lal A, Ho PT. Different infectivity of Mediterranean and Southern Asian tomato leaf curl New Delhi virus isolates in cucurbit crops. Plants. 2022;11:704.
- 58. Zaidi SS-E-A, Martin DP, Amin I, Farooq M, Mansoor S. Tomato leaf curl New Delhi virus: a widespread bipartite begomovirus in the territory of monopartite begomoviruses. Mol Plant Pathol. 2017;18(7):901–11. pmid:27553982