Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Analysing genome sequences and associated metadata during the COVID-19 pandemic in Iraq revealed points to be improved: An observational retrospective study

  • Ali Hadi Abbas ,

    Contributed equally to this work with: Ali Hadi Abbas, Aoula Al-Zebeeby

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing

    alih.abbas@uokufa.edu.iq

    Affiliation Department of Microbiology, Faculty of Veterinary Medicine, University of Kufa, Al-Najaf, Iraq

  • Aoula Al-Zebeeby ,

    Contributed equally to this work with: Ali Hadi Abbas, Aoula Al-Zebeeby

    Roles Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Resources, Validation, Visualization, Writing – original draft, Writing – review & editing

    Affiliation Department of Pathology, Faculty of Veterinary Medicine, University of Kufa, Al-Najaf, Iraq

  • Mohammed Al-Saadi,

    Roles Investigation, Visualization, Writing – original draft, Writing – review & editing

    Affiliation Department of Internal and Preventive Medicine, College of Veterinary Medicine, University of Al-Qadisiyah, Al-Qadisiyah, Iraq

  • Ahmed Jasim Neamah

    Roles Writing – original draft, Writing – review & editing

    Affiliation Department of Microbiology, College of Veterinary Medicine, University of Al-Qadisiyah, Al-Qadisiyah, Iraq

Abstract

The COVID-19 pandemic started in Wuhan China and rapidly transmitted worldwide, the illness is characterised by respiratory manifestations like coughing, breathing difficulties and pneumonia that could lead to death. Real-time whole genome sequencing of severe acute respiratory syndrome corona virus 2 (SARS-CoV-2) was adopted in many countries to track the infection dynamics and evolution of the virus. In parallel with the global efforts, genome sequencing trials were established in Iraq during the COVID-19 pandemic, however, this new approach has not been assessed yet. Therefore, for better readiness and improvement for future pandemics, here we obtained all genomes of SARS-CoV-2 virus from Iraq (182) that were deposited in National Center for Biotechnology Information (NCBI) during the period (2020–2023). Statistical analyses of sample size, distribution and other epidemiological parameters from associated metadata, as well as the quality of genome sequences were assessed. Our data analyses highlighted some drawbacks that could be improved, namely, that most genomic sequences (62%) were collected from only two cities, a low sample size was noticed and sequencing quality was inconsistent. There was a shortage and impairment of sequencing facilities especially those of the Ministry of Health. Consequently, genome sequencing should be achieved in centres that produce the best quality. The results revealed the importance of well-documented and high-quality sequences that represent many important cities in the country, which is crucial to draw a clear projection for health officials on infection dynamics and tracking viral evolution to help in taking successful steps towards infection control.

Introduction

COVID-19 defined as an acute respiratory disease caused by a novel coronavirus of Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2). This illness was first detected in Wuhan Province, China [1,2].

The latest update from WHO on COVID-19 global cases until 27th April 2025, there are over 777,745,434 detected cases from 240 countries and more than 7 million deaths [3]. When infected individuals cough, sneeze, talking, or breathe, tinny droplets will be generated containing the virus, which will be discharged into the surrounding air. Additionally, the virus can be transmitted from person to person via direct contact or indirect contamination of shared objects [4]. The clinical signs of the illness vary from mild or even no symptoms to a fatal condition [5]. The most predominant clinical symptoms are fever, headache, coughing, difficulty breathing, shortness of breath, diarrhoea, nausea, vomiting, chills, body aches, sore throat, congestion/runny nose, and loss of smell or taste. Ultimately, the infection could lead to pneumonia or respiratory failure, septic shock, heart attacks, liver problems, and death.

At the onset of the newly emerged disease, in Wuhan, China in late 2019, the causative agent was not yet known. The first genome sequence was generated in January 2020 from bronchial lavage. This genome sequence revealed a coronaviridae virus belongs to the genus Betacoronavirus [6], with a genome sized ~ 30 kb of single-stranded positive sense RNA, containing five Open Reading Frames (ORFs) encoding viral proteins named as: RNA-dependent RNA polymerase ORF1a and b, spike protein (S), envelope (E), membrane (M) and an ORF for the nucleocapsid (N) [6,7].

COVID-19 pandemic showed high vulnerability of health systems worldwide, especially in the low and middle-income countries, some of these issues are lack of Personal Protective Equipment (PPE), testing kits, inadequate transport chains, trained individuals, lack of enough funding to cope with a pandemic at such scale [8,9].

Right after the start of the pandemic, real-time whole genome sequencing consortiums like the first COVID-19 Genomic UK Consortium (COG-UK) started in March 2020, an integrated national scale of the SARS-CoV-2 genomic surveillance network were established in different countries to track viral evolution and mutations lead to increase pathogenicity or transmissibility [10]. Tracking viral evolution is essential as viruses are constantly evolving and changing pathogenicity to hosts [11,12].

It has been proven that whole-genome sequencing (WGS) is the most precise and accurate tool to follow up on the transmission dynamics of epidemics and give the required information for the outbreak control to decision makers [13]. Real-time WGS was applied during the emergence of the West African Ebola outbreak between 2014–2016 [14,15]. Since then, real-time analysis and sharing of sequencing data during outbreaks has become routine in public health assets as a fundamental measure of outbreak response [14]. However, inconsistent efforts and incomplete metadata with lack of important parameters have a huge impact on downstream data analyses [16,17].

During the COVID-19 pandemic new approaches were developed to track infection and monitoring newly emerged variants this was through test waste water for SARS-CoV-2. This strategy offered fast and affordable choice to track emergence and evolution of the virus as it easy to get access to the samples and proven successful [18,19].

In Iraq, the first genome sequence of SARS-CoV-2 was released in 2020 [20], followed by other attempts to sequence COVID-19 cases in certain regions of the country to detect possible emerging lineages such as alpha variant and delta variant [21,22]. In this study, an observational retrospective study was adopted, we analysed the genomic sequences and linked metadata deposited in NCBI public repositories in order to assess the sequencing efforts for tracking SARS-CoV-2 virus in Iraq and to inform health policy makers to improve practices for future outbreaks. The achieved analyses revealed aspects to be improved in order to upgrade this capability in Iraq.

Methods

SARS-CoV-2 genome sequences

The genome sequences used in this study were obtained from the NCBI website that dedicated to the SARS-CoV-2 data hub (https://www.ncbi.nlm.nih.gov/labs/virus/vssi/#/virus?SeqType_s=Nucleotide&Country_s=Iraq&HostLineage_ss=Homo%20sapiens%20(human),%20taxid:9606&SLen_i=27000%20TO%2030000&CreateDate_dt=2020-04-28T00:00:00.00Z%20TO%202023-09-06T23:59:59.00Z), the following search filtering was applied; sequence length were set to (27000–30000), the option of release date was set from (April 2020–7th September 2023) and Iraq was selected for geographical region option, all other options were left on default settings. This search involved whole genome sequences from Iraq covering the time from the start of the pandemic cases in Iraq, i.e., (April 2020–7th September 2023). A multi-FASTA file of 182 whole genome sequences was downloaded (see S1 Table), and a metadata file contains all listed information about these samples is provided (see S2 Table).

All data analyses were achieved using the R platform via RStudio [23]. The complete R code used in this study is provided in (S1 File).

Geographical distribution of genomes across Iraq map

Special R packages implemented in the RStudio interface version (2023.09.1 + 494) [23].

SARS-CoV-2 variants assignment

Genome sequences of the SARS-CoV-2 virus were assigned to the variants by pangolin software [24] for variant detection, and they are included in the metadata file downloaded from NCBI (S2 Table).

Statistical analyses and plot generation

All statistical analyses were performed on R platform version 4.3.2 (2023-10-31) [25]. The Kruskal-Walis test of one-way ANOVA for skewed data implemented in “rstatix” package [26] were applied to test if there were significant differences between the analysed groups of genomes. A Pairwise Wilcox test within the “rstatix” package was achieved to detect which group had a significant statistical difference from other groups.

Logarithm base 10 and square root normalisation were applied to achieve more representative data visualisation plots using R version 4.3.2. In addition to several R packages used in generating plots were listed in (S3 Table).

Results

We analysed these genomic sequences along with the provided metadata collected from samples of COVID-19 cases from Iraq. It is worth mentioning that many aspects are missing from the metadata linked to this genomic dataset, for example gender, age, occupation, travel status, nationality, underlying diseases, and purpose of study.

Distribution of genomes from Iraq

A total of 182 whole genome sequences of SARS-CoV-2 were taken from cases around the country from the time given (i.e., from April 2020- September 2023), in comparison to WGS sequences from developed countries such as the UK and USA were 3, 218, 629 and 3,403, 246, respectively using the same research criteria in the NCBI database. This reveals a huge difference in the number of WGS of SARS-CoV-2 between Iraq and pioneered countries. Iraq these genomes were collected from seven different Iraqi cities, and 28 (15.4%) genomes were taken from unspecified locations in Iraq. Most samples were sequenced from one city (Samawa) with 75 (41.2%) genomes, followed by Erbil 38 (20.9%) followed by the capital city of Iraq (Baghdad), and the rest four cities were represented by a very small number of genomes (Fig 1).

thumbnail
Fig 1. Distribution of SARS-CoV-2 genomes collected from Iraq according to the NCBI repositories. A bar plot represents the actual number of genomes collected from each city; “unknown” refers to genomes from Iraq with no location assigned. The plot generated using the R platform, and the package ggplot2.

https://doi.org/10.1371/journal.pone.0326750.g001

Assessment of genome quality

Genome quality evaluation is an important step for ensuring the quality of genomic data and its liability for further analyses. The completeness of the genome is crucial for perfection of comparative genomics and whole genome SNP mapping, which consequently affects the final variant assignment of SARS-CoV-2. The presence of gaps in the genome sequences reduces the efficacy of the a forementioned analyses. Usually, gaps in the genome are missing DNA bases, which are represented by the “N” character instead of normal DNA characters (“A”, “C”, “T”, and “G”).

We calculated the number of “N” characters in genomes collected from each assigned sequencing institutions or location as recorded in NCBI meta data. Analysis of variance showed that genomes from Baghdad, followed by ThiQar, have a significant increase in gaps (p < 0.05) in comparison to all other locations. Further analysis for gnomes from Baghdad that were collected or sequenced by 3 different institutions revealed that genomes from the Iraqi Ministry of Health had a significant level of gaps (Fig 2B), p-value < 0.0001.

thumbnail
Fig 2. Genome quality assessment.

A) Box plot of genome quality assessment was analysed according to the number of ambiguous bases in the genome sequence represented by the character “N”. The Y axis represents the logarithmic normalisation of the numbers of “N” characters in the genomes from each city. Coloured by the country of the laboratory where the sequencing was achieved. B) The same as in A but focused only on genomes from Baghdad, which showed the highest number of ambiguous bases; Y axis showed the actual number of Ns. Box plots are coloured according to the organisation of the sequencing laboratory. Plots were generated using the R statistical platform and library ggplot2.

https://doi.org/10.1371/journal.pone.0326750.g002

The genomes with the best sequence quality were those from Erbil and Duhok with the least number of “Ns”, while the highest number of gaps can be noticed in genomes sequenced from Baghdad and from ThiQar (Fig 2A).

SARS-CoV-2 genomes across time and location

The collection dates covered the time from the start of COVID-19 cases in Iraq in April 2020 to the first of December 2022, the samples were collected unevenly (Fig 3). The broader time coverage was noticed in samples collected from Erbil as it covered the period from December 2020 to October 2022 followed by samples from Baghdad and Najaf; however, the collection date of samples from Baghdad showed biases towards the second half of 2022, while samples from Najaf were only 8 samples that were collected sporadically, mainly during year the 2021. Significantly, samples from Samawa were confined within a relatively narrow timeframe, which was in the year 2021. The rest of the samples from an unspecified location, ThiQar, Duhok and Basrah were collected at a specific time (Fig 3).

thumbnail
Fig 3. Distribution of genomes from the NCBI database from Iraqi cities across time.

This plot represents the collection date of genomes (X axis) as a timeline for each city (Y axis), colour-coded according to the sample location.

https://doi.org/10.1371/journal.pone.0326750.g003

SARS-CoV-2 variants across location and time

According to the genomic data, the SARS-CoV-2 variants were plotted against time and location (Fig 4). Remarkably, samples isolated from patients in Baghdad were represented by a broader number of variants (11 variants) most variants were the BAx variants and the newly emerged hybrid variant XBBx, which were positioned in the latest collection dates. While samples from Erbil were represented by 10 variants, the B.1.617.2 lineage was the most lasting one over time; others were clustered within AYx variants and BA.1x variants. Although samples from Samawa formed the majority of genomes from Iraq, they only fall into 8 variants of Bx lineages. Notably, 28 samples from an unknown location in Iraq fall into only one variant B.1.428.1.

thumbnail
Fig 4. Shows the distribution of variants across time and location.

A) The Y axis represents the variants of SARS-CoV-2 according to pangolin classification. Points were coloured according to the location of genomes (cases). Plots were generated by the R platform and library ggplot2. B) Variants that were isolated from each city are coloured according to pangolin classification.

https://doi.org/10.1371/journal.pone.0326750.g004

Discussion

Globally, COVID-19 pandemic was the first to apply such an unpresented large-scale real-time genomic sequencing approach worldwide, which allowed researchers and health care policymakers both locally and internationally to draw plans for disease transmission, control and national and international preparedness by unravelling disease transmission dynamics and SARS-CoV-2 virus evolution, and the emerging of new variants in a real-time manner [27,28]. In Iraq, the first and only application of relatively large-scale whole genome sequencing on infectious agents was on the biggening of COVID-19 pandemic in a determination to meet global efforts to track SARS-CoV-2 virus infection and evolution. It is an important first step to use such an approach to provide valuable information during outbreaks to inform official health care entities [29]. External quality assessment of whole genome sequencing projects is essential to ensure better quality and improvement for future challenges [30,31].

In Iraq, we noticed that the genomic sequencing efforts are not centralised or nationally organised, it rather depends on individual research driven trials. In this study, we assess this new experience in Iraq to improve the overall process, increase awareness among authorities and health care specialists on points to improve, avoid misleading information and elevate the quality of the produced genomic data to meet global efforts on controlling COVID-19 pandemic and future outbreaks.

Most of the analyses in this study relied on state-of- the-art data visualisation software, that simplified and summarise huge and complex data in clear pinpoint figures. The use of illustrations and informative figures is highly recommended in the studies related to public health, to show global public health information as it simplified complex data, enhances communication, make it readable by non-technical individuals, easy to be viewed and transferred, as well as, showing trends in the data [32,33].

The associated meta-data of the genomic data set deposited in NCBI repositories showed that the whole genomic sequencing effort was not centrally organised, as different institutions (academic, health care officials, local and international organisations) were involved in sample collection, from a few locations and not evenly distributed across Iraq. Sample size and distribution are important factors that affect whether WGS studies are epidemiologically informative [34]. Although some locations, like Erbil, showed the best distribution of sample collection across time, the overall trend clearly presents an inconsistent timeframe for the sample collection date, which might reflect an unclear view of the evolution and dynamics of the pandemic [35]. Real-time sample collection and sequencing can facilitate not only evolution dynamics but also reveal the potential source of infection, portal of entry, location of the first emerging of a given SARS-CoV-2 variant, which has a big impact on how to choose and apply effective control measures that could help halt the transmission, to the resolution of a given outbreak even within a given health setting [16]. In this study, the distribution of SARS-CoV-2 variants across locations revealed a different behaviour regardless of the number of collected samples, with Baghdad having the highest number of variants, followed by Erbil and Samawa, respectively. These differences might be evoked due to the social specifications of each city: Baghdad as the capital of Iraq, which hosts the main international airport in the country, and Erbil, with similar specifications as the capitol of KRG (Kurdistan Regional Government), even though the number of genomes is much less than those from Samawa, this might suggest increasing the sampling and sequencing from these cities and other cities like Najaf, Basrah and Karbala as the latest cities have also been considered portal of entry to the country and places for annual mass gatherings for both national and international visitors for big religious events. The socio-spatial structure and people movement across cities or communities have major effect on disease transmission and causative agent evolution [36], population density, and geographical and environmental differences have an effect on disease dissemination [37]. The only location in Iraq that showed a unified variant (B.1.428.1) prevailed in a location that has not been specified (termed “unknown” in this study). Analyses of SARS-CoV-2 variant assignment and the sample collection time frame suggested a local, confined outbreak in this undefined location.

At the time of this study, the most recent variant identified was XBB variants, from samples isolated from cases in Baghdad. This is probably because Baghdad is the only location that has samples collected from the time of this variant emergence.

Genome quality is one of the most important parameters for a whole genomic study. The assessment of the quality of genomes from each location or institution was done by calculating the number of sequence gaps. The analysis suggests that the best complete genomes were those from the Kurdistan region (represented by genomes from Erbil and Duhok), while the worst case was the genomes from Baghdad. Further investigation on genomes from this city revealed that three sequencing institutions were involved in the generation of whole genome sequences. These institutions were the Ministry of Health, University of Baghdad, and US air force school of medicine. This analysis revealed that the main source of bad-quality genomes was those sequenced or processed in the Ministry of Health facility. This showed a major issue in the procedure of generating complete genome sequences from an important health institution. Hence, further training or scrutiny is needed over this facility on equipment and personal, to avoid the generation of low-quality genomes in future outbreaks or relying on better performing sequencing capacities like the ones in the KRG region.

We also noticed that there is many information missing from the metadata linked to this dataset, for example, age, gender, travel status, occupation, nationality, underlying diseases, and purpose of study (i.e., reasons beyond sampling). Incomplete metadata has a major effect on data-driven analyses [38]. It is worth mentioning that linking this metadata to the genome sequences in open-source global repositories has a major benefit to downstream data analyses and best possible piece of information [34]. Many metadata policies that can be adopted were proposed [39]. It is important to indicate that the result of this study is largely affected by the database resources and the metadata policy adapted by this data repository. Missing essential information such as temporal, demographic, geographical, social and even environmental metadata hindering the usability of shared WGSs, which reducing the benefits from data analyses to gain more informative inferences that help public health authorities to build a comprehensive picture of an outbreak [28,40,41].

In this study many issues were extrapolated from the available stored sequences in open access repository, most of these issues can be fixed to elevate the future work. This emphasizes on the importance of sharing pathogen sequence data and a complete set of metadata. While it highlighted common drawbacks like the lack of infrastructure, lack of experienced individuals, load on health sector, shortage in funding and cross contamination from different laboratories [8,42], it can also add up to mentioned previous issues in low and middle-income countries.

Some of noticed issues came from unparalleled centralised efforts with impaired funding, lack of training and lack of infrastructure hindered such nation-wide efforts, such issues were noticed in a number of low- income countries [43,44].

This study has relied on an open access repository provided by GenBank that stores genomes from all courtiers including Iraq, while this is a free to download, analyse and populate data, it is limited to what has been deposited in this datahub.

Conclusion

The analyses highlighted many issues that could be largely improved. Specifically, sampling issues in the context of time and location, the lack of the appropriate patient’s records and the importance of adapting a unified effort to give a better outcome to the real-time-whole genome sequencing efforts. This analysis provides important information to the health policy makers in the country on how to resolve these issues to improve health practices to confront potential future outbreaks. We largely recommend devoting more fund and resource on genomic sequencing facilities to upgrade it and to deal with the pandemic pressure with a centralised programme to unit all efforts nationwide.

Supporting information

S1 Table. Whole genome sequences of 182 COVID-19 samples and corresponding accession numbers that were used in this study.

https://doi.org/10.1371/journal.pone.0326750.s001

(CSV)

S2 Table. A metadata information has all listed information about the samples as it stored in GenBank.

https://doi.org/10.1371/journal.pone.0326750.s002

(CSV)

S3 Table. R packages with corresponding references used for analysing data in this study.

https://doi.org/10.1371/journal.pone.0326750.s003

(XLSX)

S1 File. The complete R code was used in this study to generate figures and to do statistical analysis.

https://doi.org/10.1371/journal.pone.0326750.s004

(TXT)

Acknowledgments

The authors are deeply appreciating all efforts that have been done to generate whole genome sequencing of SARS-CoV-2 genomes in Iraq. The work described in this paper would not have been possible without the many data providers who have shared their data openly.

References

  1. 1. Huang C, Wang Y, Li X, Ren L, Zhao J, Hu Y, et al. Clinical features of patients infected with 2019 novel coronavirus in Wuhan, China. Lancet. 2020;395(10223):497–506. pmid:31986264
  2. 2. Li Q, Guan X, Wu P, Wang X, Zhou L, Tong Y, et al. Early transmission dynamics in Wuhan, China, of novel coronavirus-infected pneumonia. N Engl J Med. 2020;382(13):1199–207. pmid:31995857
  3. 3. WHO. Web Page. 2025 [cited 2025 May 15]. Number of COVID-19 cases reported to WHO (cumulative total). Available from: https://data.who.int/dashboards/covid19/cases?n=c
  4. 4. Kalin Ünüvar G, Doğanay M, Alp E. Current infection prevention and control strategies of COVID-19 in hospitals. Turk J Med Sci. 2021;51(SI-1):3215–20. pmid:34289652
  5. 5. Phan LT, Nguyen TV, Luong QC, Nguyen TV, Nguyen HT, Le HQ, et al. Importation and human-to-human transmission of a novel coronavirus in Vietnam. N Engl J Med. 2020;382(9):872–4. pmid:31991079
  6. 6. Wu F, Zhao S, Yu B, Chen Y-M, Wang W, Song Z-G, et al. A new coronavirus associated with human respiratory disease in China. Nature. 2020;579(7798):265–9. pmid:32015508
  7. 7. Al-Saadi MHA, Hoidy WH. SARS-CoV-2: literature review focusing on structure, diagnosis and vaccine development. Nano BioMed Eng. 2022;14(4).
  8. 8. Albano PM. Cross-contamination in molecular diagnostic laboratories in low- and middle-income countries: a challenge to COVID-19 testing. Philipp J Pathol. 2020;5(2):7–11.
  9. 9. Marcotullio PJ, Schmeltz M. COVID-19 in three global cities: comparing impacts and outcomes. J Extr Even. 2021;08(02).
  10. 10. Aldeer M, Hilli AA, Ismail IS. Projecting the short-term trend of COVID-19 in Iraq. Digit Gov Res Pract. 2020;2(1):1–7.
  11. 11. Abbas AH, Al Saegh HA, ALaraji FS. Sequence diversity and evolution of infectious bursal disease virus in Iraq. F1000Research. 2021;10:293.
  12. 12. Al-Zebeeby A, Abbas AH, Alsaegh HA, Alaraji FS. The first record of an aggressive form of ocular tumour enhanced by Marek’s disease virus infection in layer flock in Al-Najaf, Iraq. Vet Med Int. 2024;2024:1793189. pmid:39376215
  13. 13. Faria NR, Azevedo R, Kraemer MUG, Souza R, Cunha MS, Hill SC, et al. Zika virus in the Americas: Early epidemiological and genetic findings. Science. 2016;352(6283):345–9. pmid:27013429
  14. 14. Khoury MJ, Bowen MS, Clyne M, Dotson WD, Gwinn ML, Green RF, et al. From public health genomics to precision public health: a 20-year journey. Genet Med. 2018;20(6):574–82. pmid:29240076
  15. 15. Killough N, Patterson L, The Covid-Genomics Uk Cog-Uk Consortium, Peacock SJ, Bradley DT. How public health authorities can use pathogen genomics in health protection practice: a consensus-building Delphi study conducted in the United Kingdom. Microb Genom. 2023;9(2):mgen000912. pmid:36745548
  16. 16. Francis RV, Billam H, Clarke M, Yates C, Tsoleridis T, Berry L, et al. The impact of real-time whole-genome sequencing in controlling healthcare-associated SARS-CoV-2 outbreaks. J Infect Dis. 2022;225(1):10–8. pmid:34555152
  17. 17. Francis RV, Billam H, Clarke M, Yates C, Tsoleridis T, Berry L, et al. The impact of real-time whole-genome sequencing in controlling healthcare-associated SARS-CoV-2 Outbreaks. J Infect Dis. 2022;225(1):10–8. pmid:34555152
  18. 18. Pilapil JD, Notarte KI, Yeung KL. The dominance of co-circulating SARS-CoV-2 variants in wastewater. Int J Hyg Environ Health. 2023;253:114224. pmid:37523818
  19. 19. Mohring J, Leithäuser N, Wlazło J, Schulte M, Pilz M, Münch J, et al. Estimating the COVID-19 prevalence from wastewater. Sci Rep. 2024;14(1):14384. pmid:38909097
  20. 20. Al-Rashedi NAM, Licastro D, Rajasekharan S, Dal Monego S, Marcello A, Munahi MG, et al. Genome sequencing of a novel coronavirus SARS-CoV-2 isolate from Iraq. Microbiol Resour Announc. 2021;10(4):e01316-20. pmid:33509990
  21. 21. Al-Rashedi NAM, Alburkat H, Hadi AO, Munahi MG, Jasim A, Hameed A, et al. High prevalence of an alpha variant lineage with a premature stop codon in ORF7a in Iraq, winter 2020-2021. PLoS One. 2022;17(5):e0267295. pmid:35617193
  22. 22. Al-Rashedi NAM, Alburkat H, Munahi MG, Jasim AH, Salman BK, Oda BS, et al. Genome sequence of an early imported case of SARS-CoV-2 delta variant (B.1.617.2 AY.122) in Iraq in April 2021. Microbiol Resour Announc. 2022;11(11):e0097722. pmid:36250864
  23. 23. RStudio T. RStudio: Integrated Development for R. Boston, MA: RStudio, Inc; 2016. Available from: www.rstudio.com; https://doi.org/10.1007/978-81-322-2340-5
  24. 24. Rambaut A, Holmes EC, O’Toole Á, Hill V, McCrone JT, Ruis C, et al. A dynamic nomenclature proposal for SARS-CoV-2 lineages to assist genomic epidemiology. Nat Microbiol. 2020;5(11):1403–7. pmid:32669681
  25. 25. Team RC. R: A Language and Environment for Statistical Computing. R A Lang Environ Stat Comput. 2023.
  26. 26. Kassambra A. Package “rstatix”: Pipe-Friendly Framework for Basic Statistical Tests. 2023;0.7.2:100.
  27. 27. Porter AF, Sherry N, Andersson P, Johnson SA, Duchene S, Howden BP. New rules for genomics-informed COVID-19 responses-Lessons learned from the first waves of the Omicron variant in Australia. PLoS Genet. 2022;18(10):e1010415. pmid:36227810
  28. 28. Zufan SE, Lau KA, Donald A, Hoang T, Foster CSP, Sikazwe C, et al. Bioinformatic investigation of discordant sequence data for SARS-CoV-2: insights for robust genomic analysis during pandemic surveillance. Microb Genom. 2023;9(11):001146. pmid:38019123
  29. 29. Saravanan KA, Panigrahi M, Kumar H, Rajawat D, Nayak SS, Bhushan B, et al. Role of genomics in combating COVID-19 pandemic. Gene. 2022;823:146387. pmid:35248659
  30. 30. Wegner F, Roloff T, Huber M, Cordey S, Ramette A, Gerth Y, et al. External quality assessment of SARS-CoV-2 sequencing: an ESGMD-SSM pilot trial across 15 European laboratories. J Clin Microbiol. 2022;60(1):e0169821. pmid:34757834
  31. 31. Foster CSP, Stelzer-Braid S, Deveson IW, Bull RA, Yeang M, Au J-P, et al. Assessment of inter-laboratory differences in SARS-CoV-2 consensus genome assemblies between public health laboratories in Australia. Viruses. 2022;14(2):185. pmid:35215779
  32. 32. Nattestad M, Chin CS, Schatz MC. Ribbon: Visualizing complex genome alignments and structural variation. bioRxiv. 2016;0344:82123.
  33. 33. Roosan D, Del Fiol G, Butler J, Livnat Y, Mayer J, Samore M, et al. Feasibility of population health analytics and data visualization for decision support in the infectious diseases domain: a pilot study. Appl Clin Inform. 2016;7(2):604–23. pmid:27437065
  34. 34. Oude Munnink BB, Nieuwenhuijse DF, Stein M, O’Toole Á, Haverkate M, Mollers M, et al. Rapid SARS-CoV-2 whole-genome sequencing and analysis for informed public health decision-making in the Netherlands. Nat Med. 2020;26(9):1405–10. pmid:32678356
  35. 35. Kraemer HC. Epidemiological methods: about time. Int J Environ Res Public Health. 2010;7(1):29–45. pmid:20195431
  36. 36. Aguilar J, Bassolas A, Ghoshal G, Hazarie S, Kirkley A, Mazzoli M, et al. Impact of urban structure on infectious disease spreading. Sci Rep. 2022;12(1):3816. pmid:35264587
  37. 37. Martínez L, Short JR. The pandemic city: urban issues in the time of COVID-19. Sustainability. 2021;13(6):3295.
  38. 38. Gozashti L, Corbett-Detig R. Shortcomings of SARS-CoV-2 genomic metadata. BMC Res Notes. 2021;14(1):189. pmid:34001211
  39. 39. Carey ME, Dyson ZA, Ingle DJ, Amir A, Aworh MK, Chattaway MA, et al. Global diversity and antimicrobial resistance of typhoid fever pathogens: Insights from a meta-analysis of 13,000 Salmonella Typhi genomes. Elife. 2023;12:e85867. pmid:37697804
  40. 40. Black A, MacCannell DR, Sibley TR, Bedford T. Ten recommendations for supporting open pathogen genomic analysis in public health. Nat Med. 2020;26(6):832–41. pmid:32528156
  41. 41. Oakeson KF, Wagner JM, Mendenhall M, Rohrwasser A, Atkinson-Dunn R. Bioinformatic analyses of whole-genome sequence data in a public health laboratory. Emerg Infect Dis. 2017;23(9):1441–5. pmid:28820135
  42. 42. Furuse Y. Genomic sequencing effort for SARS-CoV-2 by country during the pandemic. Int J Infect Dis. 2021;103:305–7. pmid:33333251
  43. 43. Pronyk PM, de Alwis R, Rockett R, Basile K, Boucher YF, Pang V, et al. Advancing pathogen genomics in resource-limited settings. Cell Genom. 2023;3(12):100443. pmid:38116115
  44. 44. Struelens MJ, Ludden C, Werner G, Sintchenko V, Jokelainen P, Ip M. Real-time genomic surveillance for enhanced control of infectious diseases and antimicrobial resistance. Front Sci. 2024;2.