Figures
Abstract
The naming of HIV-1 circulating recombinant forms (CRFs)—descendent viruses from the same intersubtype recombination events, is along with the designation of ‘subtypes’ and ‘groups’, routinely used to track HIV-1 diversity. However, we argue that continuing to designate all detected CRFs as distinct entities is biologically unjustified, as many represent recombinants of limited epidemiological significance. Indeed, the mechanistic underpinning of HIV-1 recombination highlights the arbitrary nature of naming these incidental recombinants, the majority of which are rarely detected again. This underlines the need to prioritise taxonomically meaningful clades, with a focus on biological significance such as emergence events associated with significant epidemiological spread, phenotypic properties or transmission advantage.
Citation: Grant HE, Olabode AS, Seiler Vellame D, Poon AFY, Brown AJL, Robertson DL (2026) Pervasive HIV recombination limits the utility of circulating recombinant form nomenclature. PLoS Pathog 22(7): e1014286. https://doi.org/10.1371/journal.ppat.1014286
Editor: Welkin E. Johnson, Boston College, UNITED STATES OF AMERICA
Published: July 13, 2026
Copyright: © 2026 Grant et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Funding: HEG was supported by the MRC Precision Medicine Doctoral Training Programme. ASO and AFYP were supported by a project grant from the Canadian Institutes of Health Research (PJT-183832). DSV acknowledges support from a BBSRC Research Experience Placement while a student. ALB was supported by NIH (GM110749). DLR acknowledges funding from the Medical Research Council (MRC, MC_UU_12014/12 and MC_UU_00034/5). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.
Introduction
When HIV-1 was first sequenced in 1985 [1], it had already undergone over 60 years of evolution [2] before causing a global pandemic. It was not until sequences from central Africa subsequently became available that the enormous scale [3,4] and age [5] of HIV-1 genetic diversity became apparent [6]. Founder effects linked to spread of the main pandemic group (M) from west central Africa [7] have resulted in distinct clades (around 15% divergence; [8]) that are named subtypes A-D and F-H, J and K, with A and F further partitioned into sub-subtypes A1-A6 and F1-F2. In general, the subtypes do not differ clinically in presentation or in response to antiretroviral therapy, but rather reflect their demographic histories. For example, subtype B has been predominantly associated with men who have sex with men and persons who inject drugs in North America and western Europe [9], while subtype C is predominantly found in heterosexual populations in southern Africa [10].
Although direct PCR-sequencing of viral RNA was developed in 1991 [11] lack of automation meant procedures were laborious, and subtype designations were based on sequences of partial regions of the genome or individual genes. For instance, in 1992, the first set of HIV-1 subtypes were designated A–E based on the env gene alone—however, it was later noted that subtype E viruses were classified as subtype A based on gag [12] and thus most likely not a ‘pure’ subtype. This evidence was substantiated with the emergence of cheaper and more rapid sequencing technologies, where analysis of the whole subtype E genomes established that the gag and pol regions appeared more like subtype A1 [14,13]. Therefore, this lineage was later designated as the first circulating recombinant form, CRF01_AE, a recombinant of subtypes A and an unknown subtype E [15].
In 1999, an HIV-1 Nomenclature Workshop discussed ambiguities in the HIV nomenclature and formalised the naming of HIV groups, subtypes, sub-subtypes and established circulating recombinant forms (CRF) to describe clades descended from recombination events involving multiple subtypes [15]. To limit identification of incidental recombinants, three or more epidemiologically unlinked genomes with a shared recombination pattern were recommended to designate a CRF. Currently, there are 177 identified HIV-1 CRFs (https://www.hiv.lanl.gov/components/sequence/HIV/crfdb/crfs.comp, accessed 01 June 2026). The genetic diversity within and between CRFs ranges widely: some exhibiting divergence equivalent to subtypes, others being highly related to each other, some CRFs have large numbers of sequences, but many are never sequenced again.
Here, we make the case that while monitoring for clades with novel properties remains important, many CRFs require no special designation. Future designation of HIV-1 clades of interest, whether they are a CRF or not, should focus on evidence for epidemiological significance, as they do for other pathogens.
Recombination is a continuous process
Mutation and recombination are not entirely independent processes for HIV [16], since both are a consequence of the reverse transcription process [17]. Template switching is an obligate part of reverse transcription [18], and additional template switches can happen multiple times per replication event [19], making recombination more frequent than mutation at this level [20]. While this process serves to repair damaged genomes from the same parental population [21], recombination in individuals with multiple infections can generate unique recombinant forms (URFs). Since the detectability of these events is proportional to the divergence between the parental material, inter-subtype recombination is well documented, while intra-subtype recombination is not.
Signatures of recombination can be detected at both long and short evolutionary time frames. In the more distant past, the SIV from chimpanzee host populations (SIVcpz) from which HIV-1 M descended, is a chimera of two SIVs (SIVrcm and SIVgsn) from different monkey prey species of chimpanzees [22] and has adaptive significance linked to counteracting the antiviral host molecule tetherin [23]. A subsequent deep recombination event within HIV-1 M is supported by differences in the estimated times to the most recent common ancestor across the genome [24].
Global CRF lineages represent a spectrum of older and contemporary recombination events [25], and breakpoints within the same genome can have different ages as sequential recombination events take place [26]. For example, CRF135 is composed of breakpoints between CRF01 and 07, and breakpoints between subtypes C and B in the CRF07 genomic region.
‘Pure’ HIV-1 subtypes
All HIV strains are descendants of recombinants to some extent, including the ancestors of subtypes [27], but the ability to detect recombination patterns depends on the divergence between the contributing strains, the time since the event, and the availability of parental reference sequences. This is exemplified by the missing ‘Subtype E’ parent of CRF01_AE [14,13], or by CFR02_AG, which has since been inferred to be the parent of subtype G, not the other way around [28].
Using an unsupervised clustering method that did not rely on pre-defined reference subtypes, Olabode et al. [29] found evidence of extensive recombination in the global diversity of HIV-1 genomes by constructing networks for a series of sliding windows along the alignment, and applying a community detection method (dynamic stochastic block modeling; [30]) to group them according to the relatedness for each window. Recombination was then identified by samples that changed membership between communities. Using this method, the genomic diversity of the M group of HIV-1 was partitioned into 25 clusters, with only 5% of genomes remaining in the same cluster along their entire length. Thus, 95% of genomes across the entire M group were found to be identifiably recombinant.
Recombination patterns along the genome
Recombination can happen anywhere along the genome, determined by several factors, such as sequence similarity [31] and RNA structure [32,33]. Whilst there will be biases in breakpoint frequency along the genome because of these underlying mechanistic processes, selection in the form of viability of a recombinant once generated, will also shape this distribution. A consistent pattern has been observed in multiple studies of breakpoint location, indicating moderate levels of recombination in gag and pol, very low levels within env, and much higher levels at the accessory genes vif, vpr, tat, and nef on either side of env [29,34–36], which can be observed with phylogenetic and non-phylogenetic based detection tools, between and even within subtypes (Fig 1). The HIV-1 envelope protein is essential for cellular entry with an extremely complex trimer structure, providing a significant functional constraint against disruption by recombination with dissimilar sequence [37–39]. Hence, inter-subtype daughter viruses with recombination breakpoints within env may be less successful.
Patterns of recombination within the same subtypes follow similar constraints as seen (c) within subtype B, and (d) within subtype C.
Selection for and against certain breakpoints results in clear hotspots (and coldspots) of recombination that mean similar recombination patterns can arise independently by chance. In Uganda, where 50% of recent genomes are unique recombinants, the envelope gene was found to have recombined intact either as subtype A1 on a subtype D background, or the converse, multiple times [36].
CRF frequencies
In the Los Alamos HIV Database (accessed 01 June 2026), of all CRF classifications, 82 (46%) had ten or fewer representative sequences, and a large majority (n = 146, 82%), had fewer than a hundred. As partial gene sequences are much more common in the database, when restricted to full genomes, 151 (85%) have ten or fewer representatives.
There has been speculation that recombination between subtypes might bring together favourable mutations [40]. But it appears that, in general, the majority of CRFs are not reported after their initial designation. We see a log-linear relationship between the number of representative CRF sequences and the time since their introduction (Fig 2) suggesting that these variants are lost through genetic drift more often than they expand due to any selective advantage.
The number of sequences follows a linear relationship, where the frequency increases 15.6% every year, (linear model, log number of sequences against year, intercept 149, slope = 0.07, ).
The three main exceptions are CRF01_AE, CRF02_AG, and CRF07_BC with over 110K, 32K, and 32K sequences, respectively. Together, they alone make up 86% of the CRF sequences available, and each of these CRFs has more numerous representation than the subtypes F, G, H, J, and K combined (28K). This is in part explained by their age; CRF01_AE originated in Africa, before seeding epidemics in Thailand and China, and is a large enough clade that there have been multiple proposals to further stratify it into sub-lineages [41–43]. Similarly, CRF02_AG (composed of subtypes A and G) is a very old and diverse lineage found widely in West and Central West Africa [44,45], while CRF07_BC is a dominant clade in China, presumably linked to its emergence there. Therefore, the origins of these clades are more likely to be similar to those of the subtypes themselves, i.e., reflecting founder effects during the history of the HIV-1 pandemic [46].
Potential for misleading classification
Different automated subtyping tools include different references, since it quickly becomes cumbersome to include new CRFs as they arise. For example, REGA [47] includes references up to CRF47, and SCUEAL [48] up to CRF51 whereas JPHMM (a hidden Markov Model-based tool) [49] does not include any CRFs in its reference set. In addition, the inclusion of CRF references in the subtyping process can lead to misleading subtyping results as it is difficult to decipher new CRF sections from parts of the genome that resemble the parent.
Looking at KT276261 as subtyped by SCUEAL (Fig 3), a strict interpretation might suggest this genome is an inter-subtype recombinant with three parents: G, CRF25, and CRF43. However, the section of the genome predicted to be CRF25 has parent subtypes A and G, and the section subtyped CRF43 has parents CRF02_AG and subtype G. Therefore, it could also be the case that KT276261 more resembles a subtype G genome throughout, where parts of the genome are closer to certain reference sections.
The bottom row in the figure shows an example of a genome from Spain in 2014 [KT276261] with the recombination pattern G/CRF25/G/CRF43 as subtyped by the tool SCUEAL.
Importantly, while subtype classification is straightforward where there is strong geographic structure, in regions where there are multiple subtypes newly circulating, myriad recombinant forms will arise, as is the case in London (UK) [50]. Moreover, where subtypes have co-existed in the same population for several decades, as in Uganda, URFs become the most common form of the virus [36].
Conclusion
Recombination is a pervasive, ongoing, and even predictable evolutionary process in HIV-1 evolution, such that naming every cluster of circulating recombinants has limited utility. The HIV-1 subtype classification system, whilst imperfect, is widely understood by the community because it reflects founder events that happened before the recorded pandemic [2]. The use of highly curated ‘non-recombinant’ subtype references means that our understanding of HIV diversity is fixed relative to a single point in time.
We do not propose revising the designation of widely established subtypes or CRFs with well-recognised epidemiological relevance. Rather, we suggest that the naming of all new CRFs based on being identified on only three occasions is no longer justified, because many do not have large enough numbers of sequences available to be confident they are of any special functional or epidemiological significance more than incidental transmitted non-recombinant viruses. In addition, tracking and naming every transmitted recombinant overemphasises the readily detectable inter-subtype recombination events, obfuscating the importance of recombination in HIV’s evolutionary history.
Practically, we suggest that use of unsupervised clustering methods, such as the DSBM [29], be used for finding suitable reference sequences. Analyses that require large numbers of sequences to explore the diversity and evolution of HIV (like phylodynamic or clustering studies) would certainly benefit from this approach, as it would include the most appropriate references, whilst also reducing the potentially confounding reticulate evolutionary histories due to inter and intra-subtype recombination. Furthermore, classification of diversity can be updated as more data becomes available, and reflect the frequency of new sequences and the dynamic process of recombination.
Acknowledgments
We thank all data producers for making their HIV-1 sequences freely available. For the purpose of open access, the authors have applied a Creative Commons Attribution (CC BY) licence (where permitted by UKRI ‘Open Government Licence’ or ‘Creative Commons Attribution. No-derivatives (CC BY-ND) licence’ may be stated instead) to any Author Accepted Manuscript version arising.
References
- 1. Ratner L, Haseltine W, Patarca R, Livak KJ, Starcich B, Josephs SF, et al. Complete nucleotide sequence of the AIDS virus, HTLV-III. Nature. 1985;313(6000):277–84. pmid:2578615
- 2. Worobey M, Gemmel M, Teuwen DE, Haselkorn T, Kunstman K, Bunce M, et al. Direct evidence of extensive diversity of HIV-1 in Kinshasa by 1960. Nature. 2008;455(7213):661–4. pmid:18833279
- 3. Alizon M, Wain-Hobson S, Montagnier L, Sonigo P. Genetic variability of the AIDS virus: nucleotide sequence analysis of two isolates from African patients. Cell. 1986;46(1):63–74. pmid:2424612
- 4. Potts KE, Kalish ML, Bandea CI, Orloff GM, St Louis M, Brown C, et al. Genetic diversity of human immunodeficiency virus type 1 strains in Kinshasa, Zaire. AIDS Res Hum Retroviruses. 1993;9(7):613–8. pmid:8369166
- 5. Zhu T, Korber BT, Nahmias AJ, Hooper E, Sharp PM, Ho DD. An African HIV-1 sequence from 1959 and implications for the origin of the epidemic. Nature. 1998;391(6667):594–7. pmid:9468138
- 6. Rambaut A, Robertson DL, Pybus OG, Peeters M, Holmes EC. Phylogeny and the origin of HIV-1. Nature. 2001;410(6832):1047–8.
- 7. Faria NR, Rambaut A, Suchard MA, Baele G, Bedford T, Ward MJ, et al. HIV epidemiology. The early spread and epidemic ignition of HIV-1 in human populations. Science. 2014;346(6205):56–61. pmid:25278604
- 8. Li G, Piampongsant S, Faria NR, Voet A, Pineda-Peña A-C, Khouri R, et al. An integrated map of HIV genome-wide variation from a population perspective. Retrovirology. 2015;12:18. pmid:25808207
- 9. Vermund SH, Leigh-Brown AJ. The HIV epidemic: high-income countries. Cold Spring Harb Perspect Med. 2012;2(5):a007195. pmid:22553497
- 10. Wilkinson E, Engelbrecht S, de Oliveira T. History and origin of the HIV-1 subtype C epidemic in South Africa and the greater southern African region. Sci Rep. 2015;5:16897. pmid:26574165
- 11. Zhang LQ, Simmonds P, Ludlam CA, Brown AJL. Detection, quantification and sequencing of HIV-1 from the plasma of seropositive individuals and from factor viii concentrates. AIDS. 1991;5:675–82.
- 12. Robertson DL, Hahn BH, Sharp PM. Recombination in AIDS viruses. J Mol Evol. 1995;40(3):249–59. pmid:7723052
- 13. Gao F, Robertson DL, Morrison SG, Hui H, Craig S, Decker J, et al. The heterosexual human immunodeficiency virus type 1 epidemic in Thailand is caused by an intersubtype (A/E) recombinant of African origin. J Virol. 1996;70(10):7013–29. pmid:8794346
- 14. Carr JK, Salminen MO, Koch C, Gotte D, Artenstein AW, Hegerich PA, et al. Full-length sequence and mosaic structure of a human immunodeficiency virus type 1 isolate from Thailand. J Virol. 1996;70(9):5935–43. pmid:8709215
- 15. Robertson D, Anderson J, Bradac J, Carr J, Foley B, Funkhouser R, et al. HIV-1 nomenclature proposal. Science. 2000;288(5463):55.
- 16. Schlub TE, Grimm AJ, Smyth RP, Cromer D, Chopra A, Mallal S, et al. Fifteen to twenty percent of HIV substitution mutations are associated with recombination. J Virol. 2014;88(7):3837–49. pmid:24453357
- 17.
Coffin J, Hughes S, Varmus H. Reverse transcription of the viral genome in vivo. Retroviruses. New York: Cold Spring Harbor Laboratory Press; 1997.
- 18. Temin HM. Retrovirus variation and reverse transcription: abnormal strand transfers result in retrovirus genetic variation. Proc Natl Acad Sci U S A. 1993;90(15):6900–3. pmid:7688465
- 19. Cromer D, Grimm AJ, Schlub TE, Mak J, Davenport MP. Estimating the in-vivo HIV template switching and recombination rate. AIDS. 2016;30(2):185–92. pmid:26691546
- 20. Hu WS, Temin HM. Genetic consequences of packaging two RNA genomes in one retroviral particle: pseudodiploidy and high rate of genetic recombination. Proc Natl Acad Sci U S A. 1990;87(4):1556–60. pmid:2304918
- 21. Rawson JMO, Nikolaitchik OA, Keele BF, Pathak VK, Hu W-S. Recombination is required for efficient HIV-1 replication and the maintenance of viral genome integrity. Nucleic Acids Res. 2018;46(20):10535–45. pmid:30307534
- 22. Bailes E, Gao F, Bibollet-Ruche F, Courgnaud V, Peeters M, Marx PA, et al. Hybrid origin of SIV in chimpanzees. Science. 2003;300(5626):1713. pmid:12805540
- 23. Sauter D, Schindler M, Specht A, Landford WN, Münch J, Kim K-A, et al. Tetherin-driven adaptation of Vpu and Nef function and the evolution of pandemic and nonpandemic HIV-1 strains. Cell Host Microbe. 2009;6(5):409–21. pmid:19917496
- 24. Olabode AS, Avino M, Ng GT, Abu-Sardanah F, Dick DW, Poon AFY. Evidence for a recombinant origin of HIV-1 Group M from genomic variation. Virus Evolution. 2019;5(1):1–8.
- 25. Zhang M, Foley B, Schultz A-K, Macke JP, Bulla I, Stanke M, et al. The role of recombination in the emergence of a complex and dynamic HIV epidemic. Retrovirology. 2010;7:25. pmid:20331894
- 26. Gao Y, He S, Tian W, Li D, An M, Zhao B, et al. First complete-genome documentation of HIV-1 intersubtype superinfection with transmissions of diverse recombinants over time to five recipients. PLoS Pathog. 2021;17(2):e1009258. pmid:33577588
- 27. Kalish ML, Robbins KE, Pieniazek D, Schaefer A, Nzilambi N, Quinn TC, et al. Recombinant viruses and early global HIV-1 epidemic. Emerg Infect Dis. 2004;10(7):1227–34. pmid:15324542
- 28. Abecasis AB, Lemey P, Vidal N, de Oliveira T, Peeters M, Camacho R, et al. Recombination confounds the early evolutionary history of human immunodeficiency virus type 1: subtype G is a circulating recombinant form. J Virol. 2007;81(16):8543–51. pmid:17553886
- 29. Olabode AS, Ng GT, Wade KE, Salnikov M, Grant HE, Dick DW, et al. Revisiting the recombinant history of HIV-1 group M with dynamic network community detection. Proc Natl Acad Sci U S A. 2022;119(19).
- 30. Matias C, Miele V. Statistical clustering of temporal networks through a dynamic stochastic block model. J R Stat Soc B Stat Methodol. 2016;79(4):1119–41.
- 31. Baird HA, Gao Y, Galetto R, Lalonde M, Anthony RM, Giacomoni V, et al. Influence of sequence identity and unique breakpoints on the frequency of intersubtype HIV-1 recombination. Retrovirology. 2006;3:91. pmid:17164002
- 32. Simon-Loriere E, Galetto R, Hamoudi M, Archer J, Lefeuvre P, Martin DP, et al. Molecular mechanisms of recombination restriction in the envelope gene of the human immunodeficiency virus. PLoS Pathog. 2009;5(5):e1000418. pmid:19424420
- 33. Simon-Loriere E, Martin DP, Weeks KM, Negroni M. RNA structures facilitate recombination-mediated gene swapping in HIV-1. J Virol. 2010;84(24):12675–82. pmid:20881047
- 34. Fan J, Negroni M, Robertson DL. The distribution of HIV-1 recombination breakpoints. Infect Genet Evol. 2007;7(6):717–23. pmid:17851137
- 35. Archer J, Pinney JW, Fan J, Simon-Loriere E, Arts EJ, Negroni M, et al. Identifying the important HIV-1 recombination breakpoints. PLoS Comput Biol. 2008;4(9):e1000178. pmid:18787691
- 36. Grant HE, Hodcroft EB, Ssemwanga D, Kitayimbwa JM, Yebra G, Roger L, et al. Pervasive and non-random recombination in near full-length HIV genomes from Uganda. Virus Evolution. 2020;6(1):1–12.
- 37. Bagaya BS, Vega JF, Tian M, Nickel GC, Li Y, Krebs KC, et al. Functional bottlenecks for generation of HIV-1 intersubtype Env recombinants. Retrovirology. 2015;12:44. pmid:25997955
- 38. Woo J, Robertson DL, Lovell SC. Constraints from protein structure and intra-molecular coevolution influence the fitness of HIV-1 recombinants. Virology. 2014;454–455(1):34–9.
- 39. Golden M, Muhire BM, Semegni Y, Martin DP. Patterns of recombination in HIV-1M are influenced by selection disfavouring the survival of recombinants with disrupted genomic RNA and protein structures. PLoS One. 2014;9(6):e100400. pmid:24936864
- 40. Turk G, Carobene MG. Deciphering how HIV-1 intersubtype recombination shapes viral fitness and disease progression. EBioMedicine. 2015;2(3):188–9.
- 41. An M, Han X, Zhao B, English S, Frost SDW, Zhang H, et al. Cross-continental dispersal of major HIV-1 CRF01_AE clusters in China. Front Microbiol. 2020;11:61. pmid:32082287
- 42. Feng Y, He X, Hsi JH, Li F, Li X, Wang Q, et al. The rapidly expanding CRF01_AE epidemic in China is driven by multiple lineages of HIV-1 viruses introduced in the 1990s. AIDS. 2013;27(11):1793–802. pmid:23807275
- 43. Li X, Liu H, Liu L, Feng Y, Kalish ML, Ho SYW, et al. Tracing the epidemic history of HIV-1 CRF01_AE clusters using near-complete genome sequences. Sci Rep. 2017;7(1):4024. pmid:28642469
- 44. Faria NR, Suchard MA, Abecasis A, Sousa JD, Ndembi N, Bonfim I, et al. Phylodynamics of the HIV-1 CRF02_AG clade in Cameroon. Infection, Genetics and Evolution. 2012;12(2):453–60.
- 45. Mir D, Jung M, Delatorre E, Vidal N, Peeters M, Bello G. Phylodynamics of the major HIV-1 CRF02_AG African lineages and its global dissemination. Infect Genet Evol. 2016;46:190–9. pmid:27180893
- 46. Archer J, Robertson DL. Understanding the diversification of HIV-1 groups M and O. AIDS. 2007;21(13):1693–700. pmid:17690566
- 47. de Oliveira T, Deforche K, Cassol S, Salminen M, Paraskevis D, Seebregts C, et al. An automated genotyping system for analysis of hiv-1 and other microbial sequences. Bioinformatics. 2005;21:3797–800.
- 48. Kosakovsky Pond SL, Posada D, Stawiski E, Chappey C, Poon AFY, Hughes G, et al. An evolutionary model-based algorithm for accurate phylogenetic breakpoint mapping and subtype prediction in HIV-1. PLoS Comput Biol. 2009;5(11):e1000581. pmid:19956739
- 49. Zhang M, Schultz A-K, Calef C, Kuiken C, Leitner T, Korber B, et al. jpHMM at GOBICS: a web server to detect genomic recombinations in HIV-1. Nucleic Acids Res. 2006;34(Web Server issue):W463-5. pmid:16845050
- 50. Yebra G, Frampton D, Gallo Cassarino T, Raffle J, Hubb J, Ferns RB, et al. A high HIV-1 strain variability in London, UK, revealed by full-genome analysis: Results from the ICONIC project. PLoS One. 2018;13(2):e0192081. pmid:29389981