Skip to main content
Advertisement
  • Loading metrics

A protocol for biodiversity-informed wildlife disease surveillance

  • Michael D. Catchen ,

    Roles Methodology, Software, Visualization, Writing – original draft, Writing – review & editing

    michael.catchen@umontreal.ca

    Affiliation Département de sciences biologiques, Université de Montréal, Montréal, Québec, Canada

  • Francis Banville,

    Roles Conceptualization, Writing – original draft, Writing – review & editing

    Affiliation Département de sciences biologiques, Université de Montréal, Montréal, Québec, Canada

  • Amélie C. Boutin,

    Roles Conceptualization, Writing – original draft

    Affiliation Department of Biology, Carleton University, Ottawa, Ontario, Canada

  • Cole B. Brookson,

    Roles Conceptualization, Writing – original draft, Writing – review & editing

    Affiliation Department of Epidemiology of Microbial Diseases, Yale University School of Public Health, New Haven, Connecticut, United States of America

  • Colin J. Carlson,

    Roles Writing – review & editing

    Affiliation Department of Epidemiology of Microbial Diseases, Yale University School of Public Health, New Haven, Connecticut, United States of America

  • Gabriel Dansereau,

    Roles Conceptualization, Methodology

    Affiliation Département de sciences biologiques, Université de Montréal, Montréal, Québec, Canada

  • Rory Gibb,

    Roles Writing – review & editing

    Affiliation Department of Genetics, Evolution and Environment, University College London, London, United Kingdom

  • Marianne Houle,

    Roles Conceptualization, Writing – original draft

    Affiliation Département de sciences biologiques, Université de Montréal, Montréal, Québec, Canada

  • Benjamin Kaza,

    Roles Conceptualization, Writing – original draft

    Affiliations Département de sciences biologiques, Université de Montréal, Montréal, Québec, Canada, Department of Public & Ecosystem Health, College of Veterinary Medicine, Cornell University, Ithaca, New York, United States of America

  • Hailey Robertson,

    Roles Conceptualization, Writing – original draft, Writing – review & editing

    Affiliation Department of Epidemiology of Microbial Diseases, Yale University School of Public Health, New Haven, Connecticut, United States of America

  • David Simons,

    Roles Conceptualization, Data curation, Writing – original draft, Writing – review & editing

    Affiliation Department of Anthropology & Center for Infectious Disease Dynamics, Pennsylvania State University, State College, Pennsylvania, United States of America

  • Stephanie N. Seifert,

    Roles Conceptualization, Writing – original draft, Writing – review & editing

    Affiliation College of Veterinary Medicine, Washington State University, Pullman, Washington, United States of America

  • Timothée Poisot

    Roles Conceptualization, Methodology, Supervision, Writing – original draft, Writing – review & editing

    Affiliation Département de sciences biologiques, Université de Montréal, Montréal, Québec, Canada

Abstract

Land use and climate change are increasing the risk of spillover of zoonotic disease into human populations. However, we lack actionable information about the prevalence of pathogens in wildlife populations for most of the globe, challenging our ability to implement strategies to prevent zoonoses. Even when this data exists, it has historically been sampled opportunistically and without guidance based on known geographic distributions of hosts of zoonotic pathogens. Biosurveillance is essential to mitigating zoonotic spillover risk, but given the expensive nature of monitoring pathogens in wildlife, we need to be strategic about deciding where and what to sample to obtain as much useful information as possible — particularly to identify regions where spillover is most probable and implement preventative measures to mitigate outbreak risk. The field of biodiversity monitoring has established many practices that can directly inform optimal biosurveillance efforts. One such concept is the Biodiversity Observation Network (or BON), which aims to select monitoring locations that most effectively and efficiently capture the status and trends of biodiversity. We present a protocol for integrating data on host biodiversity into sampling priority for wildlife disease surveillance based on host species distribution models, with optional potential to integrate pathogen prevalence data (if available). This protocol has the flexibility to target different forms of sampling (collecting host occurrence vs. pathogen prevalence data) to adapt to different levels of data availability, but still makes adaptive sampling recommendations based on a principled understanding of host distribution and pathogen biology. We illustrate this flexibility with two case studies, prioritizing sampling for Hanta- and Arenaviridae in rodents and shrews in India and South Korea, respectively representing data poor and data rich contexts. We view this framework as a basis for integrating long-term biosurveillance and biodiversity monitoring programs, and maximizing the useful information available for public health decision making.

Author summary

The majority of infectious diseases that infect humans are zoonotic, meaning they circulate in wildlife populations and human infections arise from contact with wildlife, leading to transmission of disease in humans (called zoonotic spillover). Zoonotic disease presents a major challenge for public health, and in order to target preventative measures and outbreak response, public health practitioners need a robust understanding of what and where pathogens circulate in wildlife to anticipate spillover risk. However, given the limited budgets of wildlife disease surveillance programs, public health researchers must be strategic in choosing where and when to sample to best understand the status of zoonotic disease in wildlife populations. Biodiversity plays a major role in structuring the geographic structure of where zoonotic pathogens circulate, and the last few decades have seen a massive increase in the amount and types of biodiversity data that is available to researchers. Biodiversity has the potential to play a vital role in guiding sampling effort for wildlife disease, but is largely underutilized in biosurveillance programs. In this manuscript, we introduce a standardized protocol for incorporating data on the hosts of zoonotic pathogens to better target sampling effort to improve our understanding of zoonotic pathogens in wildlife and better map spillover risk to predict and prevent future outbreaks.

Introduction

Despite the well-established links between biodiversity and the potential for zoonotic pathogen spillover [13], wildlife disease surveillance programs rarely integrate biodiversity data to guide sampling efforts. The relationship between biodiversity and disease risk is complex, as both the amount of biodiversity and biodiversity change can influence the presence of pathogens in context-dependent ways [46]. Host and pathogen diversity are entwined through ecological and evolutionary processes [2,68], which together shape zoonotic spillover risk by affecting the prevalence of pathogens among wildlife that humans may encounter and from which infections may arise. Climate and land use change exacerbate this risk by modifying the geographic ranges of hosts of zoonotic pathogens [9], potentially leading to increased cross-species transmission [10], including spillover into human populations.

As a result, monitoring infectious diseases in wildlife is an important goal of several international agreements, which all fall under the One Health approach. The interlinkages between biodiversity and health are recognized in the Global Action Plan on Biodiversity and Health from the Convention on Biological Diversity (CBD) and the Pandemic Agreement of the World Health Organization (WHO). Specifically, the CBD encourages Parties to “[reinforce] planning and surveillance of biodiversity, including for wildlife habitats and zoonotic pathogen spillover risk, to better assess and address health and disease risks in order to manage wild species sustainably” [11] and each party to the WHO agreement must “[coordinate] multisectoral surveillance to detect and conduct risk assessment of emerging or re-emerging pathogens with pandemic potential” [12]. To meet this goal, countries must design monitoring systems that efficiently allocate resources to maximize their knowledge of zoonotic disease prevalence and spillover risk. Improving spillover risk estimation therefore requires a multifaceted approach that incorporates wildlife disease surveillance with data on landscape change and host ecology — including shifts in reservoir species’ habitats, community composition, and geographic distributions [8,13,14].

Many surveillance programs for zoonotic pathogens with the potential for widespread harm have been developed: the Global Outbreak Alert and Response Network (GOARN), Global Emerging Infections Surveillance (GEIS), and the Global Early Warning System (GLEWS), to name just a few [15]. However, these programs rarely utilize biodiversity data to guide surveillance effort. The field of biodiversity monitoring has established many practices and tools that can directly inform biosurveillance efforts [16], including many long-term biodiversity monitoring programs, for example: the Long Term Ecological Research Network (LTER), the National Ecological Observatory Network (NEON), Terrestrial Ecosystem Research Network (TERN), and community driven projects like the North American Breeding Bird Survey (BBS). In addition, open databases like the Global Biodiversity Information Facility (GBIF) have made large-scale data far more accessible to researchers, and data standards like Darwin Core have made data from disparate sources interoperable to enable analyses previously not possible. Further integration of long-term biodiversity monitoring programs into biosurveillance has the potential to substantially improve the spatial design of wildlife disease sampling. A core concept in ecological monitoring is the Biodiversity Observation Network (or BON): a set of monitoring locations designed to best capture the status and trends of biodiversity [17]. Ideally, BONs establish a set of spatial locations at which biodiversity data are collected in standardized forms (e.g., Essential Biodiversity Variables (EBVs; [18]) or Essential Ecosystem Service Variables (EESVs; [19])), which can then be easily aggregated for the detection and attribution of biodiversity change [20] and be used to inform decision-making at local, regional, and ultimately global scales [21]. Here, we explore how practices from biodiversity monitoring can enable better biosurveillance, and how the BON perspective enables prioritization of locations where wildlife disease surveillance could be maximally informative, particularly when disease data are scarce.

Given the cost of effective biosurveillance, strategic allocation of resources toward where, when, what, and how much to sample is imperative to ensure sampling effort yields new and useful information. Efforts have been made to think about biosurveillance sampling prioritization from both a statistical [22,23] and spatially explicit perspective [24,25]. Here, we develop a context-dependent protocol for adaptive sampling to maximize the reduction of our uncertainty of the status of a given pathogen or host. The most informative data for a particular context depends on what existing data is available (e.g., host occurrence records, pathogen prevalence data) and the local drivers of pathogen prevalence. Improving understanding of both host presence and pathogen prevalence within host species is essential to assess spillover risk. However, the locations for efficient sampling for these two goals are intrinsically different: to improve host presence prediction we should go where we are most unsure about the host’s presence, whereas to improve prevalence estimation we should rather sample where we are already sure hosts are.

Here, we establish a set of guidelines to use biodiversity data to prioritize sampling locations for wildlife disease surveillance. Notably, we select these sampling locations by accounting both for the distribution of biodiversity and current uncertainty about this distribution, as well as utilizing existing prevalence data if available. This enables sampling prioritization of both where to sample, but also what type of data is most informative in that location (i.e., host occurrence vs. pathogen prevalence). As applying this framework requires numerous decisions, we provide guidance about where it can be adapted to reflect local priorities, sampling constraints, and the availability of existing data on species presence and the prevalence of pathogens among host species. We present case studies illustrative of this framework in both data-poor (no prevalence data) and data-rich (spatially explicit estimates of prevalence) contexts on Hanta- and Arenaviridae in small mammals.

The role of biodiversity data in biosurveillance

Mitigating zoonotic spillover requires distinguishing hazard — the active circulation of a pathogen in a host reservoir — from risk, which is the probability of that hazard being realized via spillover into humans [2628]. The intensity of this hazard fluctuates with shifts in host abundance, infection prevalence, and pathogen biology, all of which occur independently of human presence, and thus the proximity of reservoir hosts to human populations does not alone constitute a significant risk.

Spillover risk is actualized when a susceptible human is exposed to a zoonotic disease hazard, through processes governed by sociological, behavioural, and economic drivers at the human-wildlife interface [29,30], all of which fluctuate through time. For instance, domestic or agricultural activities in environments contaminated by animal excreta elevate risk, whereas mitigations like improved sanitation or hygiene can reduce it towards zero. Monitoring the underlying hazard is critical because the conditions for exposure are not static. Climate-driven range shifts and land-use change can rapidly intensify human-wildlife contact, transforming a latent hazard into an immediate risk [1,10,3133].

Biodiversity data is also critical — but underutilized — for predicting probable hosts and geographic ranges of as-yet-undescribed pathogens which may pose future threats to health. Shifts in biodiversity and biogeography are invariably intertwined with the risk of zoonoses [3,34], and many pathogens may be moving into regions with higher potential for contact with humans [35]. Only a small proportion of mammalian viruses have been identified [36] and most mammalian species have been poorly sampled with respect to viruses, with current patterns of viral diversity largely reflecting historical priorities in discovery effort [37]. This means that the zoonotic hazards posed by unknown viruses that circulate in wild animal populations cannot be directly estimated. However, as we describe below, using models to predict the likely hosts of zoonotic pathogens prior to their emergence [3841] helps to alleviate the uncertainty created by this viral “dark matter.”

When is biodiversity data relevant to guiding disease surveillance?

The diversity of host communities can be a central determinant of pathogen dynamics and prevalence [4,31,4245]. One of the leading ecological determinants of spillover risk is the density of infected hosts in a location proximate to human populations [2]. Unfortunately, this itself is often not actionable because it relies on spatially explicit data about abundance of — and pathogen prevalence in — all relevant hosts, which is rarely (if ever) available. Therefore, estimates of spillover risk are limited both by lack of knowledge on disease dynamics and the distribution and abundance of host species.

We address this data limitation by using weighted estimates of host composition to prioritize sampling locations. Further, this determines which type of sampling (species occurrence vs. pathogens prevalence) is most informative at a given location. While predicted presence of host species does not fully capture the risk of disease transmission, it hints at the potential for pathogen presence. This is especially the case when different species can be weighted by their relative importance as reservoir hosts, i.e., accounting for interspecific variation in host competence, population density, and probable role in pathogen maintenance.

By contrast with a situation where all species are assumed to be equally important, prior knowledge about the relevance of different hosts for pathogen transmission can be incorporated in these models as a relative weight of host sampling priority. This data can come from phylogenetic similarity to known hosts (potentially incorporating phylogenetic distance to humans, a known predictor of zoonotic potential [41]), existing records of infection, or expert knowledge on the system in question [46]. In the extreme case where there is no prior information to establish weights for host species, this approach lends itself to an adaptation of the well-established practice of estimating local host diversity by stacking predicted species distributions [47], which is particularly appropriate when using statistical predictions of species presence [48]. Although stacking species distribution models (SDMs) is often biased towards higher estimates of local species richness [49], in our context this bias is unobjectionable as it leads towards more cautious recommendations that in the worst case overestimate the ranges of reservoirs, and the stacking of species-specific SDMs has been shown to produce better estimates of richness compared to joint (multi-species) SDMs [50]. For a full set of guidelines for interpreting the outputs of SDMs, see Box 1.

Box 1. Guidelines for the interpretation of species distribution models

Species distribution models (SDMs) are widely used tools in ecology that relate species occurrences with environmental variables to identify how species are geographically constrained by different ecological factors [51]. SDMs can then be used to estimate the realized ecological niche of a species — the environmental space (rather than geographic space) that is suitable for the species to occupy. Because occurrence data is often presence-only, meaning there are few/no records of the verified absence of a species, the maps generated from SDMs should be interpreted as the relative environmental suitability for a species in a given location compared to other sites, rather than a probability of occurrence at a particular site — though there are methods that allow for this translation [52].

In disease ecology, SDMs are often used to understand the current (and future) distribution of wildlife hosts and vectors (e.g., [5355]) and determine which areas have the potential to carry spillover risk. While SDMs can help in planning and management decisions (e.g., sampling prioritization), their interpretation can prove challenging for end users [56]. For example, predictions for each species (e.g., in a multi-host system) are frequently stacked (i.e., the suitability scores for each species are summed at each location) to understand their overlapping distribution. However, SDMs do not typically incorporate biotic interactions (for discussion of joint species distribution models, which incorporate co-occurrence data and are generally used to infer relationships rather than create spatially-explicit predictions, see [57,58]), which means they cannot predict species abundance or the definitive interactions between species within a pixel where species are predicted to co-occur [59]. Without the influence of other species and other latent variables, SDMs likely overestimate the “true” suitable area (false positives) and overall number of co-occurring species. Still, stacking SDMs with continuous suitability scores tends to produce better species richness predictions than Joint Species Distribution Models (JSDMs) [50]. Additionally, SDMs typically use predictors that are treated as functionally static in time (e.g., WorldClim bioclimatic variables), which bakes in the assumption that species are in equilibrium with their environment [60]. In reality, species distribution patterns likely shift through time, varying with seasonal patterns, ecosystem disturbance, and resource availability [60]. Thus, when incorporating SDMs into a decision-making process, users must consider the plausibility of predictions and inherent uncertainty in the context of their particular expertise.

Further complicating the interpretation of SDMs are the many decisions made during data selection, pseudo-absence generation, model specification, and evaluation that proliferate through the outputs and generate errors [61,62]. Even a poorly designed model can successfully be trained and generate predictions that seem plausible, but are ultimately as unreliable as the underlying model [63]. Thus, there’s no replacement for collaboration between modelers with technical expertise and local decision-makers, with a clearly defined purpose and standardized reporting and documentation [56,64].

What are host distribution models telling us about pathogens?

The circulation of pathogens (such as viruses or obligate intracellular bacterial infections) through an animal population can vary significantly over time and is influenced by host species behavior, immunology, and environmental factors, among other drivers. Over long time periods, pathogens are thought to circulate stably within their primary host populations throughout a landscape according to the contacts of individual members of a population. These pathogens can also circulate through other susceptible animal populations if the two competent animal populations overlap in the landscape, e.g. “bridge hosts” [65] can enable the persistence of pathogens in “maintenance hosts.” For example, possums (Trichosurus vulpecula) are the primary host of bovine tuberculosis (Mycobacterium bovis), but infection in red deer and ferrets and subsequent spillback enables pathogen persistence at larger spatial scales [66].

In places where no susceptible or potentially susceptible hosts are present, it is not possible to find infected individuals. Thus, the extent of a pathogen can be thought of as the aggregation of ranges for each host species, and a rough estimate of pathogen habitat suitability across space can be obtained by weighting host susceptibility to infection [67,68]. The ensemble distribution then represents the landscape that the pathogen faces, which is not directly formed by the physical environment, but is instead determined by the distribution of hosts in a landscape.

A protocol for biodiversity-informed biosurveillance

Here we present a protocol for adaptive sampling of wildlife disease that adjusts sampling priority, both across space and the type of data collected (host occurrence vs. disease prevalence), based on the availability of existing data. Our protocol provides a standardized framework for practitioners to decide where to sample, reevaluate priorities after data is collected, and design further samples to most efficiently inform the predicted status of zoonotic disease in wildlife.

Biodiversity dose, host weighting, and uncertainty

Throughout this manuscript, we refer to the estimate of local host composition that reflects both host richness and a priori estimates of host importance as the “biodiversity dose” — the ecological capacity for a location to maintain the pathogen. This is related to the existing concept of community competence, but distinct as the biodiversity dose is calculated as a weighted average of SDM habitat suitability scores (we provide a glossary of terms and their definitions in S3 Table). Dose therefore captures both the likely presence of host reservoirs and their relative importance in disease transmission. Biodiversity dose is defined for a group of hosts, ideally encompassing the full set of hosts for a given pathogen, but also accounting for the subset of hosts of interest for a given monitoring program.

This still leaves the task of assigning species weights. If no information is available about species importance as a host for a pathogen, a naive assumption of equal weights can be applied. If the capacity of a species to serve as a reservoir for a given pathogen is predicted with a model [39], the relative confidence in the prediction can be used as a weight.

The method used to identify infection can also be used as a basis for species weighting: for example, PCR and serology provide qualitatively different information about the host-pathogen relationship. PCR can only identify an ongoing infection during viral shedding, whereas serology can determine if a pathogen was ever present in a host. However, positive serological tests only indicate that a host was exposed to the pathogen, and does not definitively indicate the host can or did transmit the pathogen. The form of collected data should be considered, along with these factors, when assigning weights. When there is uncertainty about the capacity of a host to be infected by a pathogen, methods that assign weights based on phylogenetic proximity to known hosts [69] can also provide usable information, particularly as distance to humans is a strong predictor of zoonotic potential [70].

At some level, weights will reflect expert understanding of host-pathogen relationships and judgement of risk. While explicit recommendations of formulae to construct weights from phylogenies, predicted host-pathogen association networks, and serological evidence remain outside the scope of this paper, future work could apply comparative methods to test the efficacy of different weight-construction methods on simulated data under different levels of data availability to assess the robustness of weight construction in practice.

Assessing the robustness of sampling priority to both different weight selection and host inclusion is essential to build confidence that this protocol provides the most informative data possible. We discuss the use of sensitivity analyses to better understand the consequences of these choices on the recommended configuration of the monitoring network in the final subsection of the protocol. If there is considerable epistemic uncertainty regarding assigned host weights, or inclusion of specific hosts, priority maps from multiple candidate sets of weights can be created, and the regions where sampling priority is high under all candidate priority maps can be used as the region to target sampling.

Another crucial element for guiding sampling is the uncertainty about the distribution of the biodiversity dose, i.e., the weighted average of the uncertainty associated with the SDM for each host. There are several different ways to derive uncertainty from an SDM, and they tend to reflect different aspects of model fit — see Box 2 for more information on the various forms of SDM uncertainty that can be used. High uncertainty points to an increased priority for sampling host occurrence to better refine the prediction of host presence and overall biodiversity dose.

The information about biodiversity dose, dose uncertainty, and prevalence data (if available) can be used to make a map of sampling priority to decide where and what to sample — see Fig 1. It is critical to emphasize that while the diagram in Fig 1 is useful for planning the type of data to sample when prioritizing data collection, it cannot be used to infer risk or guide outbreak response. Instead, this protocol helps identify what type of sampling is most appropriate at a given stage of surveillance, based on host biodiversity distributions and assumptions about host importance. When the sampling objective is population surveillance or estimating pathogen prevalence within hosts, areas of high biodiversity dose (Fig 1B, Quadrants A & B) should be prioritized. However, when the goal is to improve host occurrence data, sampling should focus on areas of high dose uncertainty (Fig 1B, Quadrants B & D).

thumbnail
Fig 1. A protocol for biodiversity-informed wildlife disease surveillance.

(A): The conceptual framework for optimizing sampling priority using host species distribution models, and pathogen prevalence data if available (left). (B): The type of sampling that is most informative (pathogen prevalence vs. host occurrence), depending on a location’s level of biodiversity dose (weighted host habitat suitability) and uncertainty about that dose.

https://doi.org/10.1371/journal.pesy.0000023.g001

Based on the goal of sampling, we can combine dose and uncertainty predictions into an overall priority score at each location x as

where is the SDM score of species i at location x, is the uncertainty of this score, is the weight for the biodiversity dose, is the weight for the uncertainty such that , and is the weight for species i, such that .

Integration of prevalence data into monitoring network design

When available, data on pathogen prevalence can refine sampling priorities, but the best way to use this information for sampling is contingent on the spatial and temporal scale of the prevalence data and the goals of sampling. For example, continued surveillance of populations known to have high prevalence may be useful if the goal is to determine if this population serves as a persistent reservoir for a given pathogen, or if it was undergoing a transient outbreak when the existing data was collected.

There are numerous factors that impact how prevalence data can be used to inform spatial sampling prioritization. Pathogen prevalence data is often reported at coarse spatial scales associated with regional administrative boundaries. Alternatively, if existing prevalence data is precise with high enough spatial coverage that it enables spatially explicit estimates of prevalence, sampling can most effectively improve spatial predictions of prevalence by targeting regions where the degree of prevalence is uncertain, but for which we are confident there is a high biodiversity dose (as we do in the second case study). This is useful for locating existing reservoirs across space to better map zoonotic hazard and spillover risk. Another consideration is the sample size from which existing prevalence estimates are derived. For example, prevalence estimates derived from 10 total tested individuals are far more uncertain than 500 total individuals, and this uncertainty can be used to inform sampling priority. For a full overview of potential methods to use depending on the scale of existing prevalence data see S1 Table, and for guidelines on how to target different forms of sampling regions based on their predicted dose, dose uncertainty, and predicted prevalence, see S2 Table.

Another important consideration is the variability in what existing prevalence data is saying. For example, serological diagnostic tests indicate exposure to a pathogen at any point in the past, and do not directly indicate the individual’s ability to transmit that pathogen, whereas PCR can detect both active infection and viral load as a measure of transmissibility. Together, these can provide complementary forms of information about both historical infection dynamics and active prevalence. Fully representative sampling of wildlife populations is rarely feasible, so estimates of prevalence are obtained from a small subsample of the entire population of unknown size. Both of these factors induce error in the resulting prevalence estimates; however, there are numerous methods from population ecology for estimating abundance with imperfect detection (e.g., with N-Mixture Models, [71]) and to infer population trends over time [72]. These can naturally be extended to the context of imperfect detection of both host abundance and pathogens within hosts in infectious disease modeling [73].

Sampling point selection

Once we have a sampling priority map, we want to use it to guide sampling site selection. There are many options for point selection algorithms rooted in sampling theory, which itself has a well-developed theory of spatial sampling. A full review of spatial sampling theory is outside of the scope of this paper, but see [74] and [25], and for a comprehensive review of the general theory of sampling, see [75].

A crucial point here is that because we are interested in targeting areas of high priority, we are constrained to a set of point-selection methods for unequal probability sampling, meaning each location in space can have a distinct probability of being included in the sample. A naive approach would be to draw samples with probability directly proportional to priority value, but due to the autocorrelation associated with biodiversity dose and uncertainty, this would result in many sampling points clumped close together, which would likely provide redundant information, and be an inefficient use of sampling effort. This is a well-established issue in spatial sampling, and as a result there are many algorithms for selecting spatially balanced samples — for example: Generalized Random Tessellation Sampling (GRTS; [76]), the Pivotal Method [77], and Balanced-Acceptance Sampling (BAS; [78]). These methods are all included in the BiodiversityObservationNetworks.jl package in Julia, which we utilize in the case studies.

Recent results suggest that most site selection algorithms achieve equivalent performance for the same problem [79], which allows users to pick an algorithm based on specific features, such as the support for auxiliary environmental data. For our purposes, we find BAS the most flexible method because it is both the most computationally efficient and achieves better spatial balance than the alternatives [78].

In practice, algorithms for generating spatially balanced samples targeted toward regions of high priority may not be skewed “enough” toward high priority sites for a given use case, so we also tend to tilt the distribution of priority scores to make the values of high priority even higher. This is done using an exponential transform, , where is the original priority at a given location x, and the parameter controls how much to tilt the adjusted map toward areas of high priority.

Another common method for designing samples is stratification. In the context of spatial sampling, stratification consists of splitting the spatial domain into discrete regions (corresponding with specific strata), each of which contains a user-determined number of samples. In our case, stratification can be used to divide space into the different data collection regimes (sampling for host presences, prevalence or both; Fig 1B) and distribute sampling effort across these strata based on the goals of the monitoring program, with spatially balanced point selection methods being used within each stratum. For further guidelines for interpretation and practical considerations when using selected sampling points, see Box 3.

Sensitivity/robustness analysis

This protocol relies on several choices of weights based on the goals of sampling and a priori estimates of host relevance for the pathogen(s) of interest. Assessing the sensitivity of the priority map to weight choices is important to ensure the priority map is robust to small tweaks to the weights that do not reflect meaningful biological differences in host relevance.

For a given set of weights, , a natural way to assess the sensitivity of the overall priority is to examine the mean absolute change in the priority map, P, when a small random perturbation is applied to the weights. This gives us an interpretable metric of how sensitive the resulting priority map is at a chosen weight value. We create a “nudged” version of the weights, , by adding noise , which is a vector of i.i.d. of the same length as , and then renormalized (so also sums to 1). We then construct the priority map, P, for the original weights, and the priority map from the “nudged” weights. The overall sensitivity, S, of the priority map at a given weight value can be computed as the sum of the absolute difference across all locations x:

If the selected weights for a given use case are close to high sensitivity regions in the space of possible weights, end users should consider the additional caveat that the selected weights are very sensitive to small changes, and ensure the selected weight values are reflective of meaningful biological differences in host relevance. We provide an example of this analysis in S1 Fig, assessing the sensitivity of weights for the first case study.

Box 2. Guidelines for the refinement of various methodological steps

The workflow described in this manuscript involves many steps, and most of them call for decisions made by end users. In this box, we go through the different steps, and highlight key methodological considerations.

List of hosts: the list of potentially competent hosts can be assembled from biodiversity data (IUCN range maps, in-country checklist, GBIF data), or from past wildlife disease data (e.g., from VIRION [80], literature surveys, or previous surveillance programs). The list of hosts may be arbitrarily filtered or expanded to reflect local priorities.

Host weights: weights of hosts used in the biodiversity dose calculation should essentially serve as a quantification of their expected contribution to disease transmission, and will almost always be a compromise between strength of evidence (e.g., serology, PCR, and pathogen isolation provide different levels of confidence in host species competence), and transmission risk (potentially estimated using phylogenetic similarity to well-known hosts). Weights can also ultimately attempt to capture the risk of transmission to human populations (e.g., by weighting synanthropic species higher).

Host SDMs: the usual recommendations about the training and validation for SDMs apply throughout this pipeline (see [81]), and in particular the choice of predictor data, spatial extent, the quality control of occurrences, and the selection of predictive variables are important.

Uncertainty: There are many ways to quantify SDM uncertainty. Some models (like Generalized Linear Models and the form of a Boosted Regression Tree we use in the case studies) have estimates of variance built into the model structure. Absent this, a common alternative is measuring the variance of the predicted suitability score at each site across many cross-validation folds, or methods using conformal prediction [82]. An alternative is using Bayesian Additive Regression Trees (BART; [83]), which directly obtains samples from parameters, and which can thereby be aggregated into uncertainty in predicted suitability.

Prevalence: Relevant factors for choosing what prevalence data to incorporate are the minimum number of individuals sampled, the time since data was collected, and the form of test used. These can each impact the reliability of the prevalence estimate, and therefore utility of prevalence data.

BON design: Many algorithms exist for spatially balanced sampling with unequal inclusion probabilities, and the ability of these algorithms to handle various forms of auxiliary data should be considered (see [79]). It is necessary to make the typical considerations when planning sampling: the accessibility of a given site, the cost-effectiveness of sampling given how much time is required to reach a location, the ability to access private land, the sovereignty of indigenous land — these can all be used to further adjust the inclusion probability.

Sensitivity analysis: Assessing the sensitivity of the overall priority map to both selected species weights (as discussed in the section on sensitivity), as well as the weights toward uncertainty and dose are all key factors to ensure the priority map is robust to small changes in weighting. Further checks on the response of the overall priority map to the inclusion/exclusion of host species (particularly those for which there is no direct evidence they can host a particular pathogen) should also be considered.

Case studies

To illustrate this approach, we consider two case studies on Arena- and Hantaviridae and their mammalian hosts, using the global database compiled by [84]. The first case study is on rodent and shrew hosts of Arenaviridae in India, and reflects the scenario where there is no prevalence data available. The second is for Hantaviridae in South Korea, where we have sufficient prevalence data to make spatially explicit predictions of prevalence and use this to guide sampling.

Creating dose and uncertainty maps

For both case studies, we use the same SDM methodology outlined here. We emphasize that the specific approach to building SDMs is not the focus — sampling prioritization is only going to be as good as the SDM, and we present guidelines for the interpretation of SDM outputs in Box 1.

We obtain occurrence records for each host species from GBIF (https://doi.org/10.15468/dl.b3z33h for India, https://doi.org/10.15468/dl.wzcxp9 for South Korea). As predictor variables, we use the 19 bioclimatic variables from CHELSA [85] at 1 km x 1 km resolution for South Korea, and from WorldClim [86] at 2.5 arcminute resolution (approximately 5 km x 5 km at the equator) for India (to keep the raster size reasonable). We generate pseudoabsences [87] using background thickening [88], with a buffer around each occurrence record where no pseudoabsences can go (25 km for India, 8 km for South Korea).

Models were trained using SpeciesDistributionToolkit.jl [89] and EvoTrees.jl [90] in Julia. Specifically, we use a Boosted Regression Tree (BRT; [91]) with a Gaussian loss metric, meaning the value of each node in the tree is fit to a Gaussian using maximum-likelihood estimation. Therefore, the BRT provides both a suitability prediction score, and an uncertainty value associated with it, for each pixel.

In order to standardize the dose values for each species, we compute the empirical cumulative distribution function (ECDF) on raw output predictions, ensuring all prediction layers are on the same scale.

Case Study 1: Designing a network without prevalence data

Our first case study considers the situation where we are primarily limited by a lack of prevalence data. This case study was inspired by [92], published 9 years after Lassa virus (Mammarenavirus lassaense) was first identified in Nigeria. In the immediate aftermath of the discovery of a virus of severe public health importance, [92] were the (self-identified) first to sample for Lassa virus outside of Africa. Lassa virus circulates in hosts with widespread distributions, and at the time, Lassa virus was difficult to distinguish in clinical settings from other infectious diseases present in India (e.g., typhoid fever, malaria). Therefore it was not unreasonable to suspect it may be present in rodent and shrew populations in India.

Today we know that Lassa virus is confined to Western Africa, but this case study serves as an example for how we use biodiversity data to prioritize sampling in the context of a recently identified virus without existing reliable information about its geographic extent, prevalence, or known reservoir species. We construct a priority map for sampling using the host species tested by [92]: Funambulus tristriatus, Hystrix indica, Rattus rattus, and Suncus murinus. Note that in a modern context, other forms of data could be used to supplement the list of potential hosts that were not available to [92], e.g., phylogenies constructed with modern sequencing methods, and crowdsourced occurrence data from GBIF of other rodents and shrews in the region. For the sake of example, we choose to stick with the hosts considered in the original study.

To generate a priority map for sampling, we follow the SDM procedure outlined above, and for the sake of example assign species weights at random (resulting in 0.11 for Suncus murinus, 0.34 for Rattus rattus, 0.2 for Hystrix indica, and 0.35 for Funambulus tristriatus).

Fig 2A shows a bivariate map of both the dose and uncertainty aggregated across host species. The map of India highlights three large-scale regions of interest: the south, where the model is confident that many host species are present, with increasing uncertainty moving northward; the northwest, where the model is uncertain about many or all species; and the northeast, where the model is also relatively confident that host species are present. We say a region in the top 50% of uncertainty for a given species is for “discovery sampling”, and a region in the top 50% of biodiversity dose is for “prevalence sampling”. In Fig 2B we see the total number of species that fall into each of these categories across space. Fig 2C shows our priority map, with both and equal to 0.5 (reflecting equal priority toward discovery and surveillance). The white points reflect sampling locations generated using Balanced Acceptance Sampling applied to the tilted priority map with (see Sampling Point Selection subsection above). This tilting value was chosen arbitrarily to result in a set of sampling locations concentrated at high priority values. As the tilting value increases, the median priority score at selected sites increases, and the variance of the priority score at selected sites decreases (S2 Fig). In practice, the tilting value should be adjusted to reflect confidence and uncertainty in the decisions of host inclusion and weighting made during the prioritization process — if there is little uncertainty associated with the hosts and weighting, high tilting values will ensure all sites are very high priority, where as more uncertainty may warrant a diversity of mid-to-high priority values, which will be selected by lower tilting values.. The teal markers reflect the four historical sampling locations from [92]. Finally, in Fig 2D, we see the host species that contributed the most to the priority score across space (computed as the maximum value of at location x across species i).

thumbnail
Fig 2. The case study for India.

(A) Biodiversity dose + uncertainty bivariate plot. (B) Number of hosts in prevalence-regime (top 50% of habitat suitability for that species) vs. discovery (top 50% of uncertainty for that species), (C) Sampling priority map and BON sites. X’s are historical sampling locations from [92]. (D) The largest host contributor to priority. Administrative boundaries come from GADM.

https://doi.org/10.1371/journal.pesy.0000023.g002

Case Study 2: Accounting for prevalence data in network design

Our second case study revolves around incorporating prevalence data into sampling prioritization. In contrast to the prior example where very little is known about the prevalence of a given pathogen, we now demonstrate the protocol for the surveillance of two hantaviruses in South Korea (Orthohantavirus hantanense & Orthohantavirus seoulense), which are two of the main causative agents of hemorrhagic fever with renal syndrome (HFRS; [93]), and the recently characterized Imjin virus (Thottimvirus imjinense, [94]), which, although the pathogenicity in humans is unknown, has been shown to replicate in human macrophages and endothelial cells [95]. This is reflective of biosurveillance targeting multiple objectives: both known zoonoses and a recently discovered virus with limited knowledge about the potential to infect humans. Typically, the hosts of these viruses are rodents and shrews, and in our study we focus on three small-mammal hosts: Rattus norvegicus, Apodemus agrarius, and Crocidura lasiura. These host-virus pairs were selected as the database contained prevalence records across the country at a broad enough spatial scale to enable spatially explicit prediction of prevalence.

The data comes from Project ArHa, a database of Arena- and Hantaviridae host-virus associations [84]; specifically the following studies from South Korea with explicit prevalence estimates: [94,96107].

Prevalence information is incorporated into the prioritization by spatially explicitly estimating prevalence across the extent using Gaussian Process regression (also known as kriging) for each host-virus pair with at least 5 populations sampled across space. This is done using a Squared-Exponential covariance kernel with a fixed length scale parameter using GaussianProcesses.jl [108], and the hyperparameters for the model are optimized using the L-BFGS algorithm in Optim.jl [109]. This provides both spatial estimates of prevalence, but crucially, spatially explicit uncertainty maps associated with this prevalence estimate.

Fig 3 shows both the bivariate plot of biodiversity dose vs. dose uncertainty (Fig 3A) and dose vs. prevalence uncertainty (Fig 3B). This identifies a clear geographic trend in dose and uncertainty. Dose is highly concentrated in the southeast of the Korean peninsula, with dose uncertainty (uncertainty regarding the presence of hosts) concentrated in the northwest, near Seoul and many other large cities. In contrast, prevalence uncertainty is primarily concentrated in the region of highest dose, the southeast. This indicates that southeastern Korea is a priority region for for prevalence sampling, as it contains many hosts with high confidence, but lacks existing data on the prevalence of these pathogens.

thumbnail
Fig 3. The case study for South Korea.

(A): Bivariate plot of dose (increasing green) and dose uncertainty (increasing blue). (B): Bivariate plot of dose (increasing green) and prevalence uncertainty (increasing orange; measured as the sum of estimated variance from kriging across each host-virus pair). (C): The spatial strata for which occurrence, prevalence, and both forms of sampling should occur (see main text for how these are computed). (D): The priority scores for regions where discovery sampling is a priority (increasingly blue) and prevalence sampling is the priority (increasingly purple) in their respective strata. (E): The total priority map with sites selected using Balanced Acceptance Sampling. The marker shape of each site corresponds to the type of sampling that should occur there, based on the stratum in (C) that each point falls within. Administrative boundaries come from GADM.

https://doi.org/10.1371/journal.pesy.0000023.g003

To better isolate regions where we are confident hosts are present for prevalence sampling, and regions of high host presence uncertainty for discovery sampling, we stratify the spatial extent into discrete regions for prevalence sampling and occurrence sampling as follows: the region that is both in the top 50% of biodiversity dose and top 50% of prevalence uncertainty is the stratum for prevalence sampling, and the region in the top 50% of dose uncertainty is the stratum for discovery sampling (Fig 3C). Spatial regions that fall into both categories are then targets for both forms of sampling. The choice of 50% quantiles is arbitrary for this example. The sampling priority within each of these strata can be seen in Fig 3D. This allows explicit quantification of the regions of highest priority for each type of sampling, and highlights that more occurrence data for hosts is the highest priority in the northwest, and that prevalence sampling is likely to be most fruitful in the southeast.

The overall sampling priority map can be seen in Fig 3E, along with the points sampled using Balanced Acceptance Sampling with a tilting value of (see previous case study for discussion of tilting value selection), with the marker indicating the type of sampling (prevalence, occurrence, or both) that should take place at each location.

This case study demonstrates how we can integrate host SDMs with existing prevalence information to design sampling strategies and split geographic space into discrete regions where each type of sampling (occurrence and prevalence) should take place.

Box 3. Guidelines for the interpretation of recommended sampling locations

We want to emphasize that sampling points, though given as exact (longitude, latitude) coordinates, are not to be interpreted without critical thought. Since the grain at which the underlying SDMs are operating is often significantly larger than the scale at which one may have to decide where to place a sampling trap, selected points should be thought of as most useful as the centroids of “local” regions (where “local” corresponds to the pixel size of the SDMs used for prioritization) wherein sampling effort could best improve knowledge about the ecological settings of viral transmission. To make this point explicit, if the model suggests sampling for Mus musculus in an urban area and the coordinate itself is in the middle of the center of a highway, the model’s output should be interpreted as indicating the general area where sampling should be done (see below figure).

For example, Fig 4 shows three selected sampling sites for the second case study on South Korea, with OpenStreetMap data corresponding to the cell in the priority map that was selected. As the original climate layers have pixels that are approximately 1 km2, there is considerable heterogeneity within these pixels. Experts on particular taxa should assess the scope of the local region near each sampling point, and use it to guide the practical logistics of where effort (e.g., traps) are most likely to result in useful information.

thumbnail
Fig 4. Selected sites.

Three selected sites for the second case study (colored markers, left) overlaid on a map of South Korea with the 10 most populous cities shown in dark grey dots, and the major highway network shown in light grey (to get a sense of the accessibility of these sites). Right: Land cover data, derived from OpenStreetMap, for each site. Administrative boundaries come from GADM.

https://doi.org/10.1371/journal.pesy.0000023.g004

Beyond just the spatial scale at which recommendations are made, temporal variation in both prevalence and host occurrence is a well-known driver of disease dynamics, often driven by environmental variation tied to seasonality [110,111]. Factors like migration and birth rates can affect these within-year cycles [112]. Between-year fluctuations, particularly in rodents, are both dramatic and can also counter-intuitively have an inverse effect on populations’ prevalence [113]. Here, we do not consider temporal variation in host abundances in our sampling recommendation protocol, since to do so would require orders of magnitude more data than we consider available, and would likely require the use of a different class of models/approaches.

Discussion

The biogeography of hosts of zoonotic disease imposes geographic structure on the risk of spillover and where public health intervention is most judiciously targeted [13]. Despite the numerous efforts to conduct systematized biosurveillance [15] and scale-up biodiversity monitoring [21], the benefit of incorporating biodiversity data into wildlife disease surveillance has yet to be fully realized [16]. Here, we have presented a protocol for adaptive sampling of wildlife disease based on weighted predictions of host richness and uncertainty, and guidelines for integrating prevalence data into sampling prioritization if available. The strength of this approach is prioritizing sampling locations for wildlife disease surveillance by accounting for varying availability of prevalence and host competence data, allowing it to be tailored to any particular stage of developing biosurveillance programs.

When enough prevalence data are known to allow the spatially explicit modeling of prevalence [114], relying on biodiversity data becomes a lower priority. Still, because spatially explicit modeling of prevalence is mostly feasible at small spatial scales, and because it is important to keep track of larger-scale changes in species distributions [10,115], we anticipate that the design of a biosurveillance network that is aligned with biodiversity data collection should become a standard practice. A direction for future work is to explicitly assess how good current biodiversity observation networks are for jointly supporting pathogen surveillance and biodiversity monitoring, both to assess their existing coverage and its relevance for wildlife disease monitoring, and to make recommendations for cost effective extensions to improve their utility for both biosurveillance and monitoring biodiversity change.

While this protocol relies on the spatial distribution of suitable habitat for pathogen reservoirs, our approach is not designed to capture temporal variance in virus prevalence or transmission dynamics, nor is it particularly suited for guiding time-series based optimization of surveillance efforts. Instead, this framework identifies areas with suitable environmental conditions for pathogen hosts, but does not predict when virus prevalence, or spillover risk, is increased within those areas. To fully maximize the effectiveness of sampling, users must integrate domain expertise regarding viral ecology, seasonal transmission patterns, and host-pathogen dynamics to interpret model outputs appropriately. Additionally, the static nature of SDMs may not adequately reflect rapidly changing environmental conditions, anthropogenic landscape modifications, or evolving host community structure that could influence reservoir distributions or pathogen prevalence in time and space.

There are numerous future directions for incorporating more sophisticated forms of SDMs to help alleviate these challenges: temporally-explicit SDMs are possible but require much more data at a high temporal resolution, which are typically not available for most species [116]. Similarly, mechanistic SDMs, which directly integrate biological processes into habitat suitability predictions, require far more data than is typically available. Another limitation of conventional correlative SDMs is that habitat suitability does not directly translate to species occupancy. Species may not be present in suitable habitat due to competition or dispersal limitation, and this can impact viral transmission if a host species is replaced by a less competent host (e.g., Lassa fever is reported less in cities where Mus musculus has displaced primary reservoir species). Joint SDMs (JSDMs; [57]), which model species composition across an entire species pool (thereby accounting for species interactions), are another possible frontier, but JSDMs require explicit co-occurrence data, typically not attainable from platforms like GBIF. Generally, these more informative SDMs are not possible without extensive, long-term species monitoring, which is why the goal of this protocol is to provide an efficient standardized protocol for sampling in the low data context, and to integrate biosurveillance with existing biodiversity monitoring programs.

Operationalizing this protocol will require clearing numerous logistical hurdles; the primary constraint on any biosurveillance or biodiversity monitoring program being funding. Efforts to develop a Global Biodiversity Observing System (GBiOS; [21]), analogous to the Global Climate Observing System (GCOS), propose long-term funding through a UN coalition fund, similar to the Systematic Observations Financing Facility (SOFF) that funds GCOS. Although many national bodies have developed biosurveillance programs [15], ensuring they are working in tandem is essential to maximizing their utility, and this funding model may prove more effective in addressing the trans-national nature of One Health goals, with the aim of creating a confederated system for disease surveillance that builds on existing long-term biodiversity monitoring and biosurveillance programs. Given the considerable overlap between the goals of the Kunming-Montreal Global Biodiversity Framework (GBF) and One Health objectives [117], long-term funding of GBiOS should also support the integration of biosurveillance programs. Such an integration relies on standards to ensure data is available in interoperable formats [16] to ensure monitoring has the maximum possible benefits [118]. An increasing amount of pathogen prevalence data is only actionable when individual test outcomes, both positive and negative, are shared in a way that is both standardized and interoperable with the commonly used representations of biodiversity data [119].

Earth’s shifting biogeography is driving changes in the distribution of zoonotic pathogens, creating new pathways for spillover that pose a major challenge for public health [2]. Further integration of open biosurveillance and biodiversity monitoring systems is essential to meet our common goals [16] and mitigate the consequences of this change on human health. The One Health approach emphasizes the necessity of incorporating the interlinkages between biodiversity and human health, and the protocol we have presented provides a basis to integrate the increasingly available data on Earth’s biodiversity into biosurveillance design that directly accounts for the intertwined nature of zoonotic pathogens and biodiversity. Better biosurveillance itself is not sufficient to prevent spillover, but it’s a necessary pillar to make prevention possible — it’s the best we can do in the context of limited information on pathogen prevalence — but a full approach to prevention of zoonotic spillover is broader than this alone [120]. This protocol represents a necessary next step toward integrating biosurveillance with global, long-term biodiversity monitoring systems [21] to ensure a just planetary future [121].

Supporting information

S1 Fig. The sensitivity of the overall priority map to adjusted weights.

The color of each point on the left ternary diagram is proportional to the sensitivity of each point to nudging weights (see main text). Each panel on the right corresponds to the overall priority map at weight values in the corresponding color on the left.

https://doi.org/10.1371/journal.pesy.0000023.s001

(TIFF)

S2 Fig. The density of priority scores across selected sampling sites for the India case study for three levels of tilting values (purple, blue, and green respectively).

Each tilting value was used to select 10 different sets of sampling sites, and the density for each replicate is plotted.

https://doi.org/10.1371/journal.pesy.0000023.s002

(TIFF)

S1 Table. Methods for incorporating prevalence data with varying levels of spatial and temporal coverage.

https://doi.org/10.1371/journal.pesy.0000023.s003

(PDF)

S2 Table. Interpretation and Action to take based on predicted dose, prevalence and uncertainty.

https://doi.org/10.1371/journal.pesy.0000023.s004

(PDF)

S3 Table. Glossary of selected terms used throughout the manuscript.

https://doi.org/10.1371/journal.pesy.0000023.s005

(PDF)

References

  1. 1. Eby P, Peel AJ, Hoegh A, Madden W, Giles JR, Hudson PJ, et al. Pathogen spillover driven by rapid changes in bat ecology. Nature. 2023;613(7943):340–4. pmid:36384167
  2. 2. Plowright RK, Parrish CR, McCallum H, Hudson PJ, Ko AI, Graham AL, et al. Pathways to zoonotic spillover. Nat Rev Microbiol. 2017;15(8):502–10. pmid:28555073
  3. 3. Johnson CK, Hitchens PL, Pandit PS, Rushmore J, Evans TS, Young CCW, et al. Global shifts in mammalian population trends reveal key predictors of virus spillover risk. Proc Biol Sci. 2020;287(1924):20192736. pmid:32259475
  4. 4. Keesing F, Ostfeld RS. Impacts of biodiversity and biodiversity loss on zoonotic diseases. Proceedings of the National Academy of Sciences of the United States of America. 2021;118(17):e2023540118.
  5. 5. Halliday FW, Rohr JR, Laine A-L. Biodiversity loss underlies the dilution effect of biodiversity. Ecol Lett. 2020;23(11):1611–22. pmid:32808427
  6. 6. Carlson CJ, Brookson CB, Becker DJ, Cummings CA, Gibb R, Halliday FW, et al. Pathogens and Planetary Change. Nature Reviews Biodiversity. 2025;1(1):32–49.
  7. 7. Johnson PTJ, de Roode JC, Fenton A. Why infectious disease research needs community ecology. Science. 2015;349(6252):1259504.
  8. 8. Stephens PR, Altizer S, Smith KF, Alonso Aguirre A, Brown JH, Budischak SA, et al. The macroecology of infectious diseases: a new perspective on global-scale drivers of pathogen distributions and impacts. Ecol Lett. 2016;19(9):1159–71. pmid:27353433
  9. 9. Rulli MC, D’Odorico P, Galli N, John RS, Muylaert RL, Santini M. Land Use Change and Infectious Disease Emergence. Reviews of Geophysics. 2025;63(2):e2022RG000785.
  10. 10. Carlson CJ, Albery GF, Merow C, Trisos CH, Zipfel CM, Eskew EA, et al. Climate change increases cross-species viral transmission risk. Nature. 2022;607(7919):555–62. pmid:35483403
  11. 11. Convention on Biological Diversity. Global action plan on biodiversity and health. 2024.
  12. 12. WHO. WHO Pandemic Agreement. 2025.
  13. 13. Murray KA, Olivero J, Roche B, Tiedt S, Guégan J-F. Pathogeography: leveraging the biogeography of human infectious diseases for global health management. Ecography. 2018;41(9):1411–27. pmid:32313369
  14. 14. Bell D, von Agris J, Tacheva B, Brown GW. Natural spillover risk and disease outbreaks: is over-simplification putting public health at risk?. Journal of Epidemiology and Global Health. 2025;15(1):65.
  15. 15. Sharan M, Vijay D, Yadav JP, Bedi JS, Dhaka P. Surveillance and response strategies for zoonotic diseases: a comprehensive review. Sci One Health. 2023;2:100050. pmid:39077041
  16. 16. Poisot T, Becker DJ, Catchen MD, Gibb R, Shimabukuro PHF, Carlson CJ. Biodiversity science and biosurveillance are fellow travelers. Bioscience. 2025;:biaf091.
  17. 17. Scholes RJ, Walters M, Turak E, Saarenmaa H, Heip CH, Tuama ÉÓ, et al. Building a global observing system for biodiversity. Current Opinion in Environmental Sustainability. 2012;4(1):139–46.
  18. 18. Pereira HM, Ferrier S, Walters M, Geller GN, Jongman RHG, Scholes RJ. Essential Biodiversity Variables. Science. 2013;339(6117):277–8.
  19. 19. Schwantes AM, Firkowski CR, Affinito F, Rodriguez PS, Fortin MJ, Gonzalez A. Monitoring Ecosystem Services with Essential Ecosystem Service Variables. Frontiers in Ecology and the Environment. 2024;22(8):e2792.
  20. 20. Gonzalez A, Chase JM, O’Connor MI. A framework for the detection and attribution of biodiversity change. Philos Trans R Soc Lond B Biol Sci. 2023;378(1881):20220182. pmid:37246383
  21. 21. Gonzalez A, Vihervaara P, Balvanera P, Bates AE, Bayraktarov E, Bellingham PJ, et al. A global biodiversity observing system to unite monitoring and guide action. Nat Ecol Evol. 2023;7(12):1947–52. pmid:37620553
  22. 22. Farver TB, Thomas C, Edson RK. An application of sampling theory in animal disease prevalence survey design. Preventive Veterinary Medicine. 1985;3(5):463–73.
  23. 23. Nusser SM, Clark WR, Otis DL, Huang L. Sampling Considerations for Disease Surveillance in Wildlife Populations. J Wildl Manag. 2008;72(1):52–60.
  24. 24. Andrade-Pacheco R, Rerolle F, Lemoine J, Hernandez L, Meïté A, Juziwelo L, et al. Finding hotspots: development of an adaptive spatial sampling approach. Sci Rep. 2020;10(1):10939. pmid:32616757
  25. 25. Dumelle M, Higham M, Hoef JMV, Olsen AR, Madsen L. A comparison of design-based and model-based approaches for finite population spatial sampling and inference. Methods Ecol Evol. 2022;13(9):2018–29. pmid:36340863
  26. 26. Gibb R, Redding DW, Friant S, Jones KE. Towards a “people and nature” paradigm for biodiversity and infectious disease. Philos Trans R Soc Lond B Biol Sci. 2025;380(1917):20230259. pmid:39780600
  27. 27. Gibb R, Franklinos LHV, Redding DW, Jones KE. Ecosystem perspectives are needed to manage zoonotic risks in a changing climate. BMJ. 2020;371:m3389. pmid:33187958
  28. 28. Hosseini PR, Mills JN, Prieur-Richard A-H, Ezenwa VO, Bailly X, Rizzoli A, et al. Does the impact of biodiversity differ between emerging and endemic pathogens? The need to separate the concepts of hazard and risk. Philos Trans R Soc Lond B Biol Sci. 2017;372(1722):20160129. pmid:28438918
  29. 29. Gibb R, Ryan SJ, Pigott D, Fernandez M d P, Muylaert RL, Albery GF. The Anthropogenic Fingerprint on Emerging Infectious Diseases. 2024.
  30. 30. Friant S, Mistrick J, Luis AD, Harden C, Simons D, Fichet-Calvet E, et al. Reducing the threats of rodent-borne zoonoses requires an understanding and leveraging of three key pillars: disease ecology, synanthropy, and rodentation. Lancet Planet Health. 2025;9(9):101300. pmid:40882655
  31. 31. Gibb R, Redding DW, Chin KQ, Donnelly CA, Blackburn TM, Newbold T, et al. Zoonotic host diversity increases in human-dominated ecosystems. Nature. 2020;584(7821):398–402. pmid:32759999
  32. 32. Gottdenker NL, Streicker DG, Faust CL, Carroll CR. Anthropogenic land use change and infectious diseases: a review of the evidence. Ecohealth. 2014;11(4):619–32. pmid:24854248
  33. 33. Daszak P, Cunningham AA, Hyatt AD. Emerging infectious diseases of wildlife--threats to biodiversity and human health. Science. 2000;287(5452):443–9. pmid:10642539
  34. 34. Murray KA, Preston N, Allen T, Zambrana-Torrelio C, Hosseini PR, Daszak P. Global biogeography of human infectious diseases. Proc Natl Acad Sci U S A. 2015;112(41):12746–51. pmid:26417098
  35. 35. Nova N, Athni TS, Childs ML, Mandle L, Mordecai EA. Global Change and Emerging Infectious Diseases. Annual Review of Resource Economics. 2022;14(14):333–54.
  36. 36. Carlson CJ, Zipfel CM, Garnier R, Bansal S. Global estimates of mammalian viral diversity accounting for host sharing. Nat Ecol Evol. 2019;3(7):1070–5. pmid:31182813
  37. 37. Gibb R, Albery GF, Mollentze N, Eskew EA, Brierley L, Ryan SJ, et al. Mammal virus diversity estimates are unstable due to accelerating discovery effort. Biol Lett. 2022;18(1):20210427. pmid:34982955
  38. 38. Han BA, Schmidt JP, Bowden SE, Drake JM. Rodent reservoirs of future zoonotic diseases. Proc Natl Acad Sci U S A. 2015;112(22):7039–44. pmid:26038558
  39. 39. Becker DJ, Albery GF, Sjodin AR, Poisot T, Bergner LM, Chen B, et al. Optimising predictive models to prioritise viral discovery in zoonotic reservoirs. Lancet Microbe. 2022;3(8):e625–37. pmid:35036970
  40. 40. Cummings CA, Vicente-Santos A, Carlson CJ, Becker DJ. Viral epidemic potential is not uniformly distributed across the bat phylogeny. Commun Biol. 2025;8(1):1510. pmid:41168474
  41. 41. Olival KJ, Hosseini PR, Zambrana-Torrelio C, Ross N, Bogich TL, Daszak P. Host and viral traits predict zoonotic spillover from mammals. Nature. 2017;546(7660):646–50. pmid:28636590
  42. 42. Dobson A. Population dynamics of pathogens with multiple host species. Am Nat. 2004;164 Suppl 5:S64-78. pmid:15540143
  43. 43. Keesing F, Holt RD, Ostfeld RS. Effects of Species Diversity on Disease Risk. Ecology Letters. 2006;9(4):485–98.
  44. 44. Civitello DJ, Cohen J, Fatima H, Halstead NT, Liriano J, McMahon TA, et al. Biodiversity inhibits parasites: Broad evidence for the dilution effect. Proc Natl Acad Sci U S A. 2015;112(28):8667–71. pmid:26069208
  45. 45. Johnson PTJ, Preston DL, Hoverman JT, Richgels KLD. Biodiversity decreases disease through predictable changes in host community competence. Nature. 2013;494(7436):230–3. pmid:23407539
  46. 46. Tseng KK, Koehler H, Becker DJ, Gibb R, Carlson CJ, del Pilar Fernandez M, et al. Viral Genomic Features Predict Orthopoxvirus Reservoir Hosts. Communications Biology. 2025;8(1):309.
  47. 47. Thuiller W, Pollock LJ, Gueguen M, Münkemüller T. From species distributions to meta-communities. Ecology Letters. 2015;18(12):1321–8.
  48. 48. Grenié M, Violle C, Munoz F. Is prediction of species richness from stacked species distribution models biased by habitat saturation?. Ecological Indicators. 2020;111:105970.
  49. 49. Calabrese JM, Certain G, Kraan C, Dormann CF. Stacking species distribution models and adjusting bias by linking them to macroecological models. Global Ecology and Biogeography. 2013;23(1):99–112.
  50. 50. Zurell D, Zimmermann NE, Gross H, Baltensweiler A, Sattler T, Wüest RO. Testing species assemblage predictions from stacked and joint species distribution models. Journal of Biogeography. 2019;47(1):101–13.
  51. 51. Elith J, Phillips SJ, Hastie T, Dudík M, Chee YE, Yates CJ. A statistical explanation of MaxEnt for ecologists. Diversity and Distributions. 2010;17(1):43–57.
  52. 52. Smith JR, Levine JM. Linking relative suitability to probability of occurrence in presence‐only species distribution models: Implications for global change projections. Methods Ecol Evol. 2025;16(4):854–65.
  53. 53. Lippi CA, Mundis SJ, Sippy R, Flenniken JM, Chaudhary A, Hecht G, et al. Trends in mosquito species distribution modeling: insights for vector surveillance and disease control. Parasit Vectors. 2023;16(1):302. pmid:37641089
  54. 54. Kopsco HL, Smith RL, Halsey SJ. A Scoping Review of Species Distribution Modeling Methods for Tick Vectors. Front Ecol Evol. 2022;10.
  55. 55. Bussières-Fournel A, Poisot T. Climate change increases the distribution of reservoirs of the raccoon rabies virus in Quebec. 2025.
  56. 56. Guisan A, Tingley R, Baumgartner JB, Naujokaitis-Lewis I, Sutcliffe PR, Tulloch AIT, et al. Predicting species distributions for conservation decisions. Ecol Lett. 2013;16(12):1424–35. pmid:24134332
  57. 57. Pollock LJ, Tingley R, Morris WK, Golding N, O’Hara RB, Parris KM, et al. Understanding co‐occurrence by modelling species simultaneously with a Joint Species Distribution Model (JSDM). Methods Ecol Evol. 2014;5(5):397–406.
  58. 58. Wilkinson DP, Golding N, Guillera‐Arroita G, Tingley R, McCarthy MA. Defining and evaluating predictions of joint species distribution models. Methods Ecol Evol. 2020;12(3):394–404.
  59. 59. Blanchet FG, Cazelles K, Gravel D. Co-Occurrence Is Not Evidence of Ecological Interactions. Ecology Letters. 2020;23(7):1050–63.
  60. 60. Milanesi P, Della Rocca F, Robinson RA. Integrating dynamic environmental predictors and species occurrences: Toward true dynamic species distribution models. Ecol Evol. 2019;10(2):1087–92. pmid:32015866
  61. 61. Merow C, Smith MJ, Silander JA Jr. A practical guide to MaxEnt for modeling species’ distributions: what it does, and why inputs and settings matter. Ecography. 2013;36(10):1058–69.
  62. 62. Barry S, Elith J. Error and uncertainty in habitat models. Journal of Applied Ecology. 2006;43(3):413–23.
  63. 63. Fourcade Y, Besnard AG, Secondi J. Paintings predict the distribution of species, or the challenge of selecting environmental predictors and evaluation statistics. Global Ecol Biogeogr. 2017;27(2):245–56.
  64. 64. Araújo MB, Anderson RP, Márcia Barbosa A, Beale CM, Dormann CF, Early R, et al. Standards for distribution models in biodiversity assessments. Sci Adv. 2019;5(1):eaat4858. pmid:30746437
  65. 65. Caron A, Cappelle J, Cumming GS, de Garine-Wichatitsky M, Gaidet N. Bridge hosts, a missing link for disease ecology in multi-host systems. Vet Res. 2015;46(1):83. pmid:26198845
  66. 66. Nugent G. Maintenance, spillover and spillback transmission of bovine tuberculosis in multi-host wildlife complexes: a New Zealand case study. Vet Microbiol. 2011;151(1–2):34–42. pmid:21458931
  67. 67. Cao B, Bai C, Wu K, La T, Su Y, Che L, et al. Tracing the future of epidemics: Coincident niche distribution of host animals and disease incidence revealed climate-correlated risk shifts of main zoonotic diseases in China. Glob Chang Biol. 2023;29(13):3723–46. pmid:37026556
  68. 68. Redding DW, Gibb R, Jones KE. Ecological impacts of climate change will transform public health priorities for zoonotic and vector-borne disease. 2024.
  69. 69. Elmasri M, Farrell MJ, Davies TJ, Stephens DA. A hierarchical Bayesian model for predicting ecological interactions using scaled evolutionary relationships. Ann Appl Stat. 2020;14(1).
  70. 70. Guth S, Visher E, Boots M, Brook CE. Host phylogenetic distance drives trends in virus virulence and transmissibility across the animal-human interface. Philos Trans R Soc Lond B Biol Sci. 2019;374(1782):20190296. pmid:31401961
  71. 71. Royle JA. N-mixture models for estimating population size from spatially replicated counts. Biometrics. 2004;60(1):108–15. pmid:15032780
  72. 72. Kéry M, Dorazio RM, Soldaat L, Van Strien A, Zuiderwijk A, Royle JA. Trend estimation in populations with imperfect detection. Journal of Applied Ecology. 2009;46(6):1163–72.
  73. 73. DiRenzo GV, Che-Castaldo C, Saunders SP, Campbell Grant EH, Zipkin EF. Disease-structured N-mixture models: a practical guide to model disease dynamics using count data. Ecology and Evolution. 2019;9(2):899–909.
  74. 74. Wang J-F, Stein A, Gao B-B, Ge Y. A review of spatial sampling. Spatial Statistics. 2012;2:1–14.
  75. 75. Thompson SK. Sampling. 1st ed. Hoboken, NJ: Wiley. 2012.
  76. 76. Stevens DL Jr, Olsen AR. Spatially Balanced Sampling of Natural Resources. Journal of the American Statistical Association. 2004;99(465):262–78.
  77. 77. Grafström A, Lundström NLP, Schelin L. Spatially balanced sampling through the pivotal method. Biometrics. 2012;68(2):514–20. pmid:22040115
  78. 78. Robertson BL, Brown JA, McDonald T, Jaksons P. BAS: balanced acceptance sampling of natural resources. Biometrics. 2013;69(3):776–84. pmid:23844595
  79. 79. Norman KE, Poisot T. Site selection algorithms for optimal ecological monitoring design. Ecol Indic. 2025;178:114104. pmid:41623434
  80. 80. Carlson CJ, Gibb RJ, Albery GF, Brierley L, Connor RP, Dallas TA, et al. The Global Virome in One Network (VIRION): an Atlas of Vertebrate-Virus Associations. mBio. 2022;13(2):e0298521. pmid:35229639
  81. 81. Zurell D, Franklin J, König C, Bouchet PJ, Dormann CF, Elith J, et al. A standard protocol for reporting species distribution models. Ecography. 2020;43(9):1261–77.
  82. 82. Poisot T. Conformal prediction quantifies the uncertainty of species distribution models. 2024.
  83. 83. Carlson CJ. embarcadero: Species distribution modelling with Bayesian additive regression trees in r. Methods Ecol Evol. 2020;11(7):850–8.
  84. 84. Simons D, Rivero R, Guiote AM-C, Gordon HLM, Milne GC, Rickard G, et al. Protocol to produce a systematic Arenavirus and Hantavirus host-pathogen database: Project ArHa. Wellcome Open Res. 2025;10:227. pmid:40927039
  85. 85. Karger DN, Conrad O, Böhner J, Kawohl T, Kreft H, Soria-Auza RW, et al. Climatologies at high resolution for the earth’s land surface areas. Sci Data. 2017;4:170122. pmid:28872642
  86. 86. Fick SE, Hijmans RJ. WorldClim 2: new 1‐km spatial resolution climate surfaces for global land areas. Intl Journal of Climatology. 2017;37(12):4302–15.
  87. 87. Barbet-Massin M, Jiguet F, Albert CH, Thuiller W. Selecting pseudo-absences for species distribution models: How, where and how many?. Methods Ecol Evol. 2012;3(2):327–38.
  88. 88. Vollering J, Halvorsen R, Auestad I, Rydgren K. Bunching up the background betters bias in species distribution models. Ecography. 2019;42(10):1717–27.
  89. 89. Poisot T, Bussières-Fournel A, Dansereau G, Catchen MD. A Julia toolkit for species distribution data. Peer Community Journal. 2025;5.
  90. 90. Desgagne-Bouchard P, Pandey A, S J, Blaom P, amyhxqin, Müller-Widmann D. Evovest/EvoTrees.jl: v0.18.0". Zenodo. 2025.
  91. 91. Elith J, Leathwick JR, Hastie T. A working guide to boosted regression trees. J Anim Ecol. 2008;77(4):802–13. pmid:18397250
  92. 92. Rodrigues FM, Gupta NP, Pinto BD. Serological survey for the detection of antibodies to lassa virus in India. J Indian Med Assoc. 1978;70(2):25–8. pmid:659905
  93. 93. Park K, Kim J, Noh J, Kim S-G, Cho H-K, Kim K, et al. Epidemiological surveillance and phylogenetic diversity of Orthohantavirus hantanense using high-fidelity nanopore sequencing, Republic of Korea. PLoS Negl Trop Dis. 2025;19(2):e0012859. pmid:39919119
  94. 94. Song J-W, Kang HJ, Gu SH, Moon SS, Bennett SN, Song K-J, et al. Characterization of Imjin virus, a newly isolated hantavirus from the Ussuri white-toothed shrew (Crocidura lasiura). J Virol. 2009;83(12):6184–91. pmid:19357167
  95. 95. Shin OS, Yanagihara R, Song J-W. Distinct innate immune responses in human macrophages and endothelial cells infected with shrew-borne hantaviruses. Virology. 2012;434(1):43–9. pmid:22944108
  96. 96. Gu SH, Kang HJ, Baek LJ, Noh JY, Kim H-C, Klein TA, et al. Genetic diversity of Imjin virus in the Ussuri white-toothed shrew (Crocidura lasiura) in the Republic of Korea, 2004-2010. Virol J. 2011;8:56. pmid:21303516
  97. 97. Kim H-C, Kim W-K, Klein TA, Chong S-T, Nunn PV, Kim J-A, et al. Hantavirus surveillance and genetic diversity targeting small mammals at Camp Humphreys, a US military installation and new expansion site, Republic of Korea. PLoS One. 2017;12(4):e0176514. pmid:28448595
  98. 98. Kim HC, Klein TA, Chong ST, Collier BW, Usa M, Yi SC. Seroepidemiological survey of rodents collected at a U.S. military installation, Yongsan garrison, Seoul, Republic of Korea. Military Medicine. 2007;172(7):759–64.
  99. 99. Kim W-K, No JS, Lee D, Jung J, Park H, Yi Y, et al. Active Targeted Surveillance to Identify Sites of Emergence of Hantavirus. Clin Infect Dis. 2020;70(3):464–73. pmid:30891596
  100. 100. Klein TA, Kim H-C, Chong S-T, Kim J-A, Lee S-Y, Kim W-K, et al. Hantaan virus surveillance targeting small mammals at nightmare range, a high elevation military training area, Gyeonggi Province, Republic of Korea. PLoS One. 2015;10(4):e0118483. pmid:25874643
  101. 101. Lee HW, Lee PW, Johnson KM. Isolation of the etiologic agent of Korean Hemorrhagic fever. J Infect Dis. 1978;137(3):298–308. pmid:24670
  102. 102. Lee S-H, Kim W-K, No JS, Kim J-A, Kim JI, Gu SH, et al. Dynamic Circulation and Genetic Exchange of a Shrew-borne Hantavirus, Imjin virus, in the Republic of Korea. Sci Rep. 2017;7:44369. pmid:28295052
  103. 103. No JS, Kim W-K, Kim J-A, Lee S-H, Lee S-Y, Kim JH, et al. Detection of Hantaan virus RNA from anti-Hantaan virus IgG seronegative rodents in an area of high endemicity in Republic of Korea. Microbiol Immunol. 2016;60(4):268–71. pmid:26917012
  104. 104. Park K, Lee S-H, Kim J, Lee J, Lee G-Y, Cho S, et al. A Portable Diagnostic Assay, Genetic Diversity, and Isolation of Seoul Virus from Rattus norvegicus Collected in Gangwon Province, Republic of Korea. Pathogens. 2022;11(9):1047. pmid:36145479
  105. 105. Ryou J, Lee HI, Yoo YJ, Noh YT, Yun S-M, Kim SY, et al. Prevalence of hantavirus infection in wild rodents from five provinces in Korea, 2007. J Wildl Dis. 2011;47(2):427–32. pmid:21441196
  106. 106. Sames WJ, Klein TA, Kim HC, Chong ST, Lee IY, Gu SH, et al. Ecology of Hantaan Virus at Twin Bridges Training Area, Gyeonggi Province, Republic of Korea, 2005-2007. Journal of Vector Ecology: Journal of the Society for Vector Ecology. 2009;34(2):225–31.
  107. 107. Seo MH, Kim C-M, Kim D-M, Yun NR, Park JW, Chung JK. Emerging hantavirus infection in wild rodents captured in suburbs of Gwangju Metropolitan City, South Korea. PLoS Negl Trop Dis. 2022;16(6):e0010526. pmid:35737659
  108. 108. Fairbrother J, Nemeth C, Rischard M, Brea J, Pinder T. GaussianProcesses.jl: A Nonparametric Bayes Package for the Julia Language. Journal of Statistical Software. 2022;102:1–36.
  109. 109. Mogensen PK, Riseth AN. Optim: A mathematical optimization package for Julia. JOSS. 2018;3(24):615.
  110. 110. Altizer S, Hochachka WM, Dhondt AA. Seasonal dynamics of mycoplasmal conjunctivitis in eastern North American house finches. Journal of Animal Ecology. 2004;73(2):309–22.
  111. 111. Cosgrove CL, Wood MJ, Day KP, Sheldon BC. Seasonal variation in Plasmodium prevalence in a population of blue tits Cyanistes caeruleus. J Anim Ecol. 2008;77(3):540–8. pmid:18312339
  112. 112. van Dijk JGB, Hoye BJ, Verhagen JH, Nolet BA, Fouchier RAM, Klaassen M. Juveniles and migrants as drivers for seasonal epizootics of avian influenza virus. J Anim Ecol. 2014;83(1):266–75. pmid:24033258
  113. 113. Davis S, Calvet E, Leirs H. Fluctuating rodent populations and risk to humans from rodent-borne zoonoses. Vector Borne Zoonotic Dis. 2005;5(4):305–14. pmid:16417426
  114. 114. Gras P, Knuth S, Börner K, Marescot L, Benhaiem S, Aue A, et al. Landscape Structures Affect Risk of Canine Distemper in Urban Wildlife. Front Ecol Evol. 2018;6.
  115. 115. Lawlor JA, Comte L, Grenouillet G, Lenoir J, Baecher JA, Bandara RMWJ, et al. Mechanisms, detection and impacts of species redistributions under climate change. Nat Rev Earth Environ. 2024;5(5):351–68.
  116. 116. Anselmetto N, Garbarino M, Weldy MJ, Bell DM, Daly C, Epps CW, et al. Leveraging long‐term data to improve biodiversity monitoring with species distribution models. Journal of Applied Ecology. 2025;62(11):2914–29.
  117. 117. Banville F, Burnel C, Carlson CJ, Eiffener E, Acevedo GM, Velez AP. The Global Biodiversity Framework Supports Global Assessment of One Health Actions. 2026.
  118. 118. Shanbehzadeh M, Nopour R, Kazemi-Arpanahi H. Designing a standardized framework for data integration between zoonotic diseases systems: Towards one health surveillance. Informatics in Medicine Unlocked. 2022;30:100893.
  119. 119. Schwantes CJ, Sánchez CA, Stevens T, Zimmerman R, Albery G, Becker DJ, et al. A minimum data standard for wildlife disease research and surveillance. Sci Data. 2025;12(1):1054. pmid:40544158
  120. 120. Ramakrishnan N. Bio-surveillance as One Health: A Critique of Recent Definitions and Policy Initiatives. Development. 2023;66(3–4):215–25.
  121. 121. Chapman M, Goldstein BR, Schell CJ, Brashares JS, Carter NH, Ellis-Soto D, et al. Biodiversity monitoring for a just planetary future. Science. 2024;383(6678):34–6. pmid:38175872