Figures
Abstract
Neighborhood factors, encompassing social, built, and natural environments, may explain geographic differences in the impact of COVID-19 pandemic on populations. Data from pre-existing national, population-based cohorts could be leveraged to better understand how pre-existing conditions (both individual and neighborhood) contribute to risk factor development and disease progression. We catalogued spatial and neighborhood data in the Collaborative Cohort of Cohorts for COVID-19 Research (C4R), comprising 14 diverse US cohorts (>50,000 participants). The C4R sample is generally spatially and socially representative of the overall nation, with C4R’s calculated spatial coverage representing 28% of US land area and 52% of the total US population. However, C4R (vs. non C4R) areas were more urban, wealthy, with more foreign-born residents, and less car-dependent with lower proportion employed and green. Twelve cohorts collected neighborhood characteristics – most commonly social environment data on neighborhood socioeconomic status– based on participants’ addresses. The most common built environment measures were related to food access, followed by other destination-based measures such as walkability. Natural environment data were available in the fewest cohorts, with emphasis on air quality or greenspace. This work provides clarity on available neighborhood and spatial data and facilitates future harmonization of data from C4R cohorts. Ultimately, this may enable future longitudinal and comparative analyses of neighborhood influences on COVID-19.
Citation: Hirsch JA, Besser LM, Pescador Jimenez M, Dickinson ST, Cornelius T, Francisco ST, et al. (2026) Spatial and neighborhood data in the collaborative cohort of cohorts for COVID-19 Research (C4R). PLoS One 21(7): e0352170. https://doi.org/10.1371/journal.pone.0352170
Editor: Hani Amir Aouissi, Environmental Research Center (CRE), ALGERIA
Received: September 5, 2025; Accepted: June 6, 2026; Published: July 22, 2026
Copyright: © 2026 Hirsch et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: To analyze de-identified individual participant data including neighborhood measures, researchers must receive authorization for access. This authorization is required to maintain the integrity of the data and protect participant privacy. Those wishing to access C4R data will need to submit a Data Access Request (DAR) through the database of Phenotypes and Genotypes (dbGaP) (https://www.ncbi.nlm.nih.gov/gap/). Authorized researchers can then access C4R data from BioData Catalyst® (BDC) (https://biodatacatalyst.nhlbi.nih.gov/). BioData Catalyst (BDC) is a cloud-based research ecosystem that permits access to and a platform for the analysis of scientific data supported by NHLBI. BDC currently hosts C4R datasets from 11 of the 14 C4R cohorts with additional C4R Data becoming accessible in the coming months. Cohort data from many of the C4R parent studies is also accessible through a dbGaP data access request. BDC provides instructions for finding and viewing C4R and cohort data as well as the steps to request authorization to use de-identified individual participant data for scientific analysis within BDC. For additional information or data, contact c4r@cumc.columbia.edu.
Funding: This study was financially supported by the National Heart, Lung, and Blood Institute [https://www.nhlbi.nih.gov] in the form of grants received by ECO (NHLBI-CONNECTS OT2HL156812), PGL (R01HL104580 and R01HL114091), SCB (R01HL148880), LCG (RC2HL101649), AVD (R01HL071759), AMK (R01HL093009 & R01HL149809), NRK (R01HL120725), SB (R01HL148431), & DM (U10HL109086). This study was also financially supported by the National Institute of Diabetes and Digestive and Kidney Diseases [https://www.niddk.nih.gov] in the form of a grant received by LCG (R01DK106209). This study was also financially supported by the National Institute of Environmental Health Sciences [https://www.niehs.nih.gov] in the form of grants received by JDK (R01ES030994 & R01ES023500). This study was also financially supported by the National Institute on Minority Health and Health Disparities [https://www.nimhd.nih.gov] in the form of grants received by AVD (P60MD002249 & P60MD002249-05S1). This study was also financially supported by the U.S. Environmental Protection Agency Science to Achieve Results (STAR) [https://www.epa.gov/research-grants/star] research assistance agreements in the form of grants received by JDK (RD831697 & RD-83830001). This study was also financially supported by the National Institute of Neurological Disorders and Stroke [https://www.ninds.nih.gov] in the form of a grant received by TR (R01NS29993). This study was also financially supported by the National Institute on Aging [https://www.nia.nih.gov] in the form of grants received by TR & SCB (RF1AG074306), JAH (R01AG072634), & GSL (R01AG049970, R01AG049970-S1, & R56AG049970. This study was also financially supported by the National Institute of Neurological Disorders and Stroke and the National Institute on Aging cooperative agreement in the form of a grant (U01NS041588) received by SEJ. This study was also financially supported by the Commonwealth Universal Research Enhancement Program (CURE) through the Pennsylvania Department of Health [https://www.pa.gov/agencies/health/research/research/cure] in the form of a grant received by GSL (4100072543). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health, including the National Heart, Lung, and Blood Institute, National Institute on Aging, and National Institute of Environmental Health Sciences. Representatives of the NINDS were involved in the review of the manuscript but were not directly involved in the collection, management, analysis or interpretation of the data. The work supported by the U.S. Environmental Protection Agency has not been formally reviewed by the Environmental Protection Agency. The views expressed in this document are solely those of the authors, and the Environmental Protection Agency does not endorse any products or commercial services mentioned in this publication.
Competing interests: The authors have read the journal’s policy and have the following competing interests: RGB is employed by the Department of Veterans Affairs, the positions and findings in this manuscript do not necessarily reflect those of the United States government. Dr. Buhr received personal fees from Viatris/Theravance Biopharma, DynaMed, and American College of Physicians, not related to this work. This does not alter our adherence to PLOS ONE policies on sharing data and materials. There are no patents, products in development or marketed products associated with this research to declare.
Introduction
As the coronavirus-19 (COVID) pandemic unfolded across the US, the burden of COVID infection, severity, poor recovery, and mortality were disproportionately high in low-income and lower resourced communities [1–3]. Extant research on how environmental factors within communities were associated with COVID outcomes [4–7] demonstrated higher infection and mortality rates in larger metropolitan areas [8], higher severity of COVID infection with higher levels of air pollution [9–11], and reduced risk of COVID mortality with greater greenspace exposure [12–14]. Despite the importance of social, natural, and built environments as powerful determinants of health behaviors and conditions, prior efforts to disentangle the role of environments during the COVID pandemic were limited by ecologic designs that lacked individual longitudinal data. These data are necessary to disentangle existing health trajectories and health risks that affect COVID susceptibility and severity from direct environmental impacts on COVID infection, outcomes, or long-term prognosis.
Several national, population-based longitudinal studies measure complex risk factor development, progression of disease, and longitudinal, life-course predictors of health. Within the context of the COVID pandemic, these cohorts presented a unique opportunity to examine how pre-existing health conditions and personal characteristics contributed to COVID risks. They also allow the examination of potential long-term changes in behaviors or health stemming from the pandemic period (e.g., decreased or increased physical activity from lockdown). Established early in the pandemic, the Collaborative Cohort of Cohorts for Covid-19 Research (C4R) brings together 14 cohort studies (S1 Table, list in abbreviations) that have followed participants across the US for between 10 and 50 years. All cohorts are listed in abbreviations; details on each cohort and C4R published elsewhere [15]. Briefly, C4R has collected information on COVID vaccination, infection, recovery, and mortality, as well as behavior and psychosocial outcomes from over 50,000 participants combined.
Outside of the efforts of C4R, many of these cohorts have previously collected spatial and neighborhood data about their participants, often over many exams and decades. These deeply phenotyped, longitudinal, well-characterized cohorts span diverse populations, life phases, and geographies. Thus, they offer sufficient variability to better understand how environmental context impacts COVID and other health outcomes. Yet efforts to operationalize environmental data across different cohorts have been decentralized and are often led by independent investigators, resulting in variable data collection protocols and metrics [16–19]. Collaboration among cohorts and harmonization of environmental data encourages the identification of factors impacting COVID outcomes and disparities. Specifically, harmonized geospatial infrastructure could support future longitudinal and comparative analyses. The first step toward this harmonization is to identify and quantify which data exist, at what times, and for which cohorts.
This paper provides a comprehensive inventory of the available spatial data across all C4R cohorts and examines the spatial extent of C4R participants. We systematically assessed which data are available within each cohort and quantified the representativeness of these combined cohorts within the US. Our longer-term goal was to document available data and inform measure harmonization and development within C4R. We expect this work will facilitate well-powered research on the influence of spatial/neighborhood characteristics on COVID-19 outcomes and pandemic shifts across diverse geographies and participants.
Methods
Cohort geospatial inventory
In 2021, three study team members (JAH, LMB, MPJ) drafted a geospatial inventory form to capture essential geospatial methods and measures within each C4R cohort (S4 Appendix in S4 File). The form documents each cohort’s: (1) format and frequency of collecting and geocoding addresses, as well as the types of addresses collected (e.g., residential, work, school) and level of geocoding (e.g., latitude-longitude or Zip + 4); (2) types of neighborhoods assessed, such as administrative data appended to geocodes (e.g., city, state, county) and the availability of buffers (e.g., Euclidian/straight line or network/street buffers); and (3) measures of the environment collected, created, or attached to participant geocodes. Specifically, the form assessed available neighborhood-level measures (i.e., already linked to participant locations) of the natural environment (e.g., greenspace, air quality), built environment (e.g., food stores, social and walking destinations, housing density), and social or economic environment (e.g., racial and ethnic composition, neighborhood residents’ socioeconomic status). Within each of these three broad categories, the form asks about specific neighborhood metrics, the method of collection or derivation (e.g., survey, GIS, audit), the data source and year, the level of measurement (e.g., census tract or Euclidian buffer), and the time period(s) in which the measure was linked to locations (e.g., baseline, visit year). Lastly, the form collects COVID-specific neighborhood measures: number of cases, number of deaths, vaccination rates, policies, and disparity indices (e.g., race- or income-based differences). For each COVID-specific neighborhood metric, the form asks for the specific metrics available, the source of the data (e.g., survey, government agency), the level of measurement (e.g., city, zip), and the timing of measurement (e.g., by years, months). The draft inventory form was reviewed for clarity and completeness by three investigators and was piloted to assess usability in one of the C4R cohorts before finalization for use with the remaining 13 C4R cohorts.
Cohort study design
The C4R Study design, as previously described [20], followed a cohort ancillary studies model. Researchers in each cohort study were directly responsible for accomplishing data collection in accordance with the standard protocols and under the supervision of their own observational studies monitoring board, steering committee, institutional review board (IRB), and any other applicable regulatory authorities. Columbia University served as the Data Coordination and Harmonization Center (DCHC) for C4R (Columbia University Institutional Review Board, IRB-AAAT3035). A full list of cohort IRBs supervising implementation of the C4R protocols is provided in S2 Table. For the purposes of the work presented in this manuscript, data were accessed at Drexel University from February 28, 2022, to April 26, 2024. Drexel IRB determined that work within this manuscript was not human subjects research. Cohort data inventories were conducted beginning February 28, 2022, and finalized January 9, 2023. Study data were previously recorded as part of the C4R Study and were handled in accordance with all applicable confidentiality and data protection regulations.
Cohort geospatial data acquisition process and ensuring confidentiality
We (JAH, LMB, MPJ, STD, TC) contacted representatives from each of the 14 C4R cohorts requesting data manuals for each cohort’s geospatial data from which information was extracted using the geospatial inventory form. We also looked for published manuscripts using measures of interest. This was done to minimize effort for each individual cohort. When each inventory was complete, we confirmed with the cohort that our assessment was accurate.
To provide data on the geographic coverage of C4R, we requested a list of geospatial units for participants’ residential locations. We requested cohorts to provide the geographic administrative units to which participant addresses were geocoded (e.g., census tract, ZIP code, county, state) which were the smallest spatial resolution shareable without potential participant identification. While select cohorts have workplace locations, we did not include these. Geocoding data were provided to us in an anonymized format (i.e., as a list of administrative geographies) and no individual participant data were linked to these locations.
Public Data Sources Linked to Geographic Coverage of C4R
Utilizing the American Community Survey (ACS) 2015−2019 5-Year estimates and the geospatial data acquired for cohorts, we derived relevant population, social, and economic context information for a comparative analysis between areas where C4R participants have lived and those in which C4R participants have not lived. We used census tracts that match the 2015−2019 ACS boundaries, derived from the National Historical Geographic Information System (NHGIS) 2019 Census Tract Boundary File: United States and Puerto Rico [21]. These measures included total population density (pop/km2), median household income, as well as percentages of older adults (≥65 years, ≥ 85 years), non-Hispanic Whites, foreign-born residents, home ownership (% rent, % own), residents below poverty, residents with formal educations (% with high school (HS) degree, % with college degree), residents employed, commuting by car, and crowding. Index of Concentration at the Extremes (ICE) was calculated from ACS data using standard methods [22] for both the high-income White households versus low-income Black households and the high-income White non-Hispanic households versus low-income Hispanic households. ICE is scaled from −1 to +1 where −1 indicates that 100% of the population in the given area is concentrated among the most deprived group, and +1 means that 100% of the population is concentrated in the most privileged group.
Data regarding urbanicity were derived from the McAlexander et al. community type classification [23]. We operationalized urbanicity as the percent of total tracts that were classified as “high density urban”, “low density urban”, and “small town/suburban”, which make up all categories other than “rural” (and any undesignated tracts) as classified by McAlexander et al. To measure hospital access, we used staffed hospital bed counts per county from the Area Health Resources Files 2018–2019 to create a number of beds per 100,000 population density measure [24]. Data on greenness was derived from the 2019 National Land Cover Database (NLCD) [25] by including six categories of land cover previously used to generate a greenness measure (proportion of each category) at the census tract level [14]. We reclassified the categories into one single “green” category and used zonal statistics to get the percent areal coverage per census tract.
We captured metrics of neighborhood COVID burden from the Johns Hopkins COVID-19 dashboard (cases per quarter) and the Centers for Disease Control (CDC) COVID Data Tracker (vaccination rates) [26,27]. We aggregated daily COVID-19 case counts per county from the Johns Hopkins COVID-19 Dashboard to quarterly counts then used population to create a quarterly cases per 100,000 population density measure. The CDC publishes data on COVID-19 vaccination rates per county, which we used to indicate completed primary series vaccination rates at the end of each quarter (Jan-Dec). We examined case rates from Q1 of 2020 until Q4 of 2021, and vaccination rates from Q1 of 2021 (when a vaccine was available to a majority of the population) to Q4 of 2021. Despite access to data for 2022, we did not examine these rates due to a lack of reporting that likely makes these data not representative of overall COVID levels.
Analyses
Inventory data on neighborhood measures within each cohort were grouped and summarized based on overall availability and construct. For example, since many cohorts did not obtain neighborhood housing type measures (i.e., age of housing, mix of multi-family vs. single family), but more collected home ownership (i.e., % within neighborhood who own or rent), we combined these into a single category (i.e., “Home ownership/housing conditions”). Similarly, many cohorts measured greenspace through park data or through satellite images, making it harder to differentiate between greenness measures and park access (which are related), so these were combined into “Greenspace/Parks.” Within each category, measures were counted and summarized based as 0 measures, 1 measure, 2–5 measures, 6 or more measures. For example, for “Access to Food Stores/Food Environment,” a cohort might have data on densities of fast-food establishments, restaurants, supermarkets, convenience stores, wholesale markets, and fruit and vegetable stands, as well as ratios of types of stores (e.g., healthy to unhealthy sources). This would count as 6 + measures. Similarly, data were summarized as being derived from GIS (e.g., the densities in the prior example) or derived from surveys to participants (e.g., self-reported access to fruit and vegetables within their neighborhood). We did not receive information from the cohorts indicating that they have audited any environmental characteristics using in-person audit tools. Finally, cohorts either had measures from only one time point (e.g., baseline or during an ancillary study) or multiple time points longitudinally (e.g., continuously, at each exam); therefore, we summarized data within each category using that dichotomy. All neighborhood measures were visualized in a heatmap that could illustrate data timing (color family), number of measures (intensity of color), and type of measures (pattern). Cleaning and summarization of inventory data was done in Excel, and the heatmap was created in R (ggplot2 package, RStudio 2023.06.1 with Base R version 4.3.0).
We received participant locations from each cohort in different aggregated geographic administrative levels and vintages. For example, one cohort provided counties in which participants had lived in 2015, while another provided tracts from four separate decennial censuses, ranging from 1980 to 2010. To harmonize these disparate boundaries, we merged and dissolved all files into one dataset and then overlaid it with 2019 census tracts from NHGIS. Tracts with at least 5% of their area overlapped by C4R participant geographies were included as “C4R areas” in the final tract-level datasets for comparison. All counties that these tracts fell in were used in final county-level datasets. Remaining census tracts and boundaries in the contiguous US were considered “non-C4R areas.” Only the contiguous US was considered for any aspect of this study, meaning that “national” measures refer to data for that subset only. All spatial processing was performed in ArcGIS Pro 2.9 (ESRI, Redlands, CA).
All public data sources on social, economic, and other contexts were joined to either the county- or tract-level geographies on matching Federal Information Processing Standards (FIPS) codes. These were then aggregated nationally, for C4R coverage area and for non-C4R coverage area. To calculate representativeness and coverage, total population and land area measures were summed in the aggregation. All other ACS 2015–2019 measures were aggregated by means. Therefore, final aggregated measures are means of tract-level measures, not percentages or indices of that level. For example, the HS Degree (%) measure is the mean of all tract-level percentages of persons with at least a high school diploma, rather than the total percentage of persons with a high school diploma. We used this method of aggregation for all additional measures. We described means and standard deviations (SD) for population, sociodemographic, economic, and environmental context data for C4R locations, non-C4R locations, and nationally. We used t-tests to calculate p-values comparing the means of C4R locations to means of non-C4R locations. All aggregation and analyses were done in Python 3.9 (Wilmington, DE).
Results
Address and geolocation data
COPDGene was the only cohort that did not centralize recorded addresses on inception of the cohort (i.e., address information kept within the 21 sites), instead centralizing addresses during a subsequent exam year once informed consent language was updated. Most cohorts updated their participants’ address information each time an exam was performed, with some cohorts updating information more frequently. For example, REGARDS recorded an initial address at the beginning of the study, then asked participants to update their addresses annually and confirmed addresses every six months during a follow-up call.
While collection and maintenance of address data was (and still is) done by all 14 cohorts, fewer cohorts (n = 12) have processed these participants’ addresses to identify their geographic location (i.e., geolocation or geocoding to a latitude-longitude for linkage to other geographic layers). Twelve of the cohorts (ARIC, CARDIA, COPDGene, FHS, HCHS/SOL, JHS, MASALA, MESA, NOMAS, REGARDS, SHS, SPIROMICS) have derived geospatial measures for administrative boundaries (e.g., tracts, zip codes, or counties), and six cohorts (ARIC, CARDIA, HCHS/SOL, JHS, MESA, REGARDS) have derived measures for Euclidian or network buffers (e.g., 1 km around residence). Cohorts varied in use of Euclidean or Network buffers and size of buffers around residential addresses; some cohorts calculated buffer sizes using metric (1 km, 3 km, 5 km) and others used imperial (0.5 mi, 1 mi, 3 mi, 5 mi). Buffers used were also specific to certain GIS measures (e.g., some cohorts measured distance to nearest roadway, which is not a buffer, while others calculated densities of road network within a buffer around a participant’s residence).
C4R geospatial coverage
C4R cohort participants covered 28% of the US contiguous land area, which includes 52% of the US total population (Table 1). Generally, compared to the US, C4R locations are more urban, with higher population densities, lower proportion non-Hispanic White residents, and higher foreign-born populations. The higher urbanicity of C4R locations can also be seen in higher proportions renting, a lower proportion commuting by car, and slightly lower greenness. Several locations within the US have participants from multiple C4R cohorts (“cohort overlap”) (Fig 1). These areas include San Francisco, CA, Minneapolis, MN, Chicago, IL, Birmingham, AL, Winston-Salem, NC, Baltimore, MD, and New York, NY. These cities represent potential opportunities to harmonize and attach raw environmental GIS data available in one cohort to another cohort that may not have these data.
Geographic coverage obtained from cohorts anonymized and aggregated by spatial unit (e.g., county, census tract) and time (e.g., baseline, most recent exam, all exams). Since participant locations were provided at varying time periods and geographies, the map does not necessarily represent locations during COVID-19 pandemic. Geographic boundaries shown in Fig 1 come from U.S. Census Bureau's TIGER/Line shapefiles.
Neighborhood measures
In our inventory of the C4R cohorts’ geospatial data, twelve cohorts (ARIC, CARDIA, COPDGene, FHS, HCHS/SOL, JHS, MASALA, MESA, NOMAS, REGARDS, SHS, SPIROMICS) indicated that they have measured neighborhood geospatial data for their study participants (Fig 2). Two cohorts (FIP/PrePF, SARP) indicated they have not collected, processed, or linked any neighborhood geospatial data for their participants (blank columns, Fig 2).
Abbreviations: ARIC, Atherosclerosis Risk in Communities; BCL, Biorepository and Central Laboratory; C4R, Collaborative Cohort of Cohorts for COVID-19 Research; CARDIA, Coronary Artery Risk Development in Young Adults; CONNECTS, Collaborating Network of Networks for Evaluating COVID-19 and Therapeutic Strategies; COPDGene, Genetic Epidemiology of COPD; FHS, Framingham Heart Study; HCHS/SOL, Hispanic Community Health Study/Study of Latinos; JHS, Jackson Heart Study; MASALA, Mediators of Atherosclerosis in South Asians Living in America; MESA, Multi-Ethnic Study of Atherosclerosis; NOMAS, Northern Manhattan Study; PASC, Post-Acute Sequelae of SARS-CoV-2 Infection; PrePF, Prevent Pulmonary Fibrosis; REGARDS, REasons for Geographic and Racial Differences in Stroke; SARP, Severe Asthma Research Program; SPIROMICS, Subpopulations and Intermediate Outcome Measures in COPD Study; SHS, Strong Heart Study.
Built environment data
Built environment data were collected by nine cohorts (ARIC, CARDIA, FHS, HCHS/SOL, JHS, MASALA, MESA, NOMAS, REGARDS). The most common data collected by cohorts was access to food stores or food environment. These cohorts measured neighborhood amenities for retail food stores, such as wholesale clubs, grocery and convenience stores, chain stores, co-ops, fast-food, specialty, meat and fish, and supermarkets. Three cohorts (JHS, MESA, REGARDS) utilized data from the National Establishment Time Series (NETS) dataset to calculate changes in the business environments of their cohorts over time [28]. This dataset allows for cohorts to track the locations and types of businesses dating back to 1990, enabling longitudinal neighborhood change analysis. CARDIA, HCHS/SOL, JHS, MESA, and REGARDS all retained data on commercial buildings, land use, walkability, recreation/physical activity facilities, and alcohol/tobacco/marijuana outlets (counts per buffer or density per land area). These data were often derived from the NETS dataset, InfoUSA/ReferenceUSA, or similar dataset offerings from ESRI (i.e., ArcGIS Business Analyst). Traffic and road infrastructure, often primarily calculated using distance to arterial roads, were measured by eight cohorts (ARIC, CARDIA, HSCS/SOL, FHS, JHS, MESA, NOMAS, REGARDS); this was disparate from street connectivity, measured using both density of intersections, ratios of Network to Euclidean buffers, and other measures of connectivity.
Social environment data
Social environment data were the most plentiful data type connected to the C4R cohorts. Most cohorts have collected aggregated data on neighborhood socioeconomic status and data on urbanicity and population density, with all cohorts in possession of neighborhood-level data reporting some level of socioeconomic status. Most commonly, cohorts linked participant data to neighborhood employment data, median household income, education levels of residents, and poverty derived from the US decennial census and ACS. Cohorts that measured urbanicity of their participants’ communities often relied on rural-urban commuting area (RUCA) codes, a tract-based classification that uses standard census measures of population density, levels of urbanization, and journey-to-work commuting to characterize all US census tracts with respect to their rural/urban status [29].
Five cohorts reported measuring levels of neighborhood social engagement and cohesion (CARDIA, JHS, HCHS/SOL, MASALA, MESA). As expected for these types of constructs, these measures were assessed using survey tools administered to cohort participants. Uniquely, MESA measured these constructs simultaneously at the community level through administration of the survey to nearby, non-MESA participants and then created neighborhood estimates within a set buffer of MESA participant residences using Bayesian estimation techniques [30]. Five cohorts incorporated levels of violence and crime (HCSC/SOL, JHS, MASALA, MESA, REGARDS). Sometimes perceptions of violence or crime were measured (e.g., in MESA and their community survey), while other times crime data were GIS-based from local municipalities or the CrimeRisk Index (provided through ESRI from Applied Geographic Solutions) [31].
Natural environment data
Natural environment data were the least available construct of geospatial data available among cohorts. Nine measured air quality (ARIC, CARDIA, HCHS/SOL, FHS, JHS, MESA, NOMAS, REGARDS, SPIROMICS), five measured greenspace/parks (CARDIA, HCHS/SOL, JHS, MASALA, MESA, REGARDS), three measured bluespace (CARDIA, HCHS/SOL, REGARDS) and three measured temperature/weather (CARDIA, FHS, MESA), and only two measured hilliness/elevation (CARDIA, MESA). Cohorts that measured air quality included data on multiple pollutants, especially nitrogen dioxide, ozone, particulate matter less than 10 um or 2.5um (PM10, PM2.5), and nitrous oxide, measuring them in many different ways (e.g., modeling, averages during specific time periods). Some cohorts collected air quality data from EPA air pollution sensors [32], while many used more complex modeling techniques to estimate fine-scale air pollution [33]. Cohorts that measured greenspace/parks with park access generally quantified park density, size, and characteristics frequently sourced from local parks departments, city governments, or ParkServe [34]. When measuring natural vegetation, cohorts often used the National Land Cover Database from the USGS [25] or the Normalized Difference Vegetation Index (NDVI), a measure of vegetation calculated using ArcGIS from Landsat or Moderate Resolution Imaging Spectroradiometer (MODIS) satellite imagery [35]. Three cohorts measured some type of bluespace data (CARDIA, HCSC/SOL, REGARDS), measuring distances to the nearest shoreline and surface land use that is coded as water or ice. The three cohorts (CARDIA, FHS, MESA) that measured temperature (e.g., average variation, extreme temperature events) and weather (e.g., precipitation levels) tended to source data from the National Oceanic and Atmospheric Administration (NOAA) local climatological data archive [36] or local monitors (e.g., Logan International Airport). Almost no cohorts measured light pollution or noise, so these were not included in Fig 2.
COVID Pandemic contextual data
To date, none of the surveyed cohorts have linked neighborhood-level COVID data to participants’ addresses. While the C4R project gathered self-reported COVID within participants’ friend and family networks, a large gap remains in connecting participants to the COVID levels or pandemic responses (e.g., policies) they experienced around their residences.
Forthcoming geospatial or neighborhood data
On or before January 2024, several cohorts reported geospatial measures being developed by investigators (through funded projects, ancillary studies, or approved manuscript proposals). These measures do not appear in this manuscript as they are not yet readily available for outside researchers through the coordinating centers. Examples include new segregation measures in ARIC, new disaster and climate measures in MESA, and new segregation, redlining, and neighborhood disinvestment measures in JHS (S3 Table).
3.8.1. Sociodemographic, economic, environmental, and COVID representativeness of C4R geographics.
Areas where C4R participants have lived or presently reside are more dense, urban, affluent, formally educated with a higher proportion foreign born residents and less car-dependent, with lower proportions employed and green than the non-C4R locations and the US overall (Table 1). Variability exists across cohorts (not shown). However, unsurprisingly, since many cohorts are based at institutions located in major cities, the geographic coverage of C4R is more dense (2613.08 population/km2) and urban (82.69) than the entire US (2087.58 and 75.61, respectively) and non-C4R locations (1575.32 and 68.72, respectively, p < 0.001). Relatedly, C4R areas have higher percentages of residents with a college degree and higher median household income, but also have a higher proportion of residents renting, below the poverty line, and unemployed. Tracking with their higher urbanicity, C4R locations have lower percentage of residents commuting by car and more racial and ethnic diversity as represented by both higher proportion foreign-born residents and lower non-Hispanic White residents. Interestingly, using ICE comparing only high-income White households to low-income Black households showed C4R locations have less concentrated privilege, while using ICE comparing high-income non-Hispanic White households to low-income Hispanic households showed C4R locations to have more concentrated privilege than non-C4R locations or the national average.
Estimates for the community COVID context of C4R participant locations showed that they generally experienced higher rates of infection and slightly higher rates of vaccination (Table 2). For example, in Q2 and Q3 of 2020, when the pandemic first peaked, the case rates for C4R locations were 536.74 cases per 100,000 and 1498.62 cases per 100,000 people, respectively, while the national rates were 489.68 cases per 100,000 people and 1435.37 cases per 100,000 people, respectively. We observed lower case rates in C4R locations in Q4 of 2020 and Q4 of 2021 than the case rates of the nation and non-C4R areas. Vaccination rates were always higher in the C4R locations than for the nation and non-C4R areas.
Discussion
In the current study, we cataloged the available geospatial and neighborhood data and geospatial coverage of the 14 cohorts included in C4R. Twelve of the 14 cohorts had at least one neighborhood measure available at the time of this inventory. Social environmental data were the most common data element, with most cohorts having data on at least one neighborhood measure of socioeconomic status. The most common built environment measures were related to food access, followed by other destination-based measures such as walkability. Natural environment data were available in the fewest cohorts, with emphasis on air quality or greenspace. The C4R sample is generally spatially and socially representative of the overall nation, covering 28% of US land area and 52% of the US population. Areas where C4R participants have lived or currently reside differed from non-C4R locations and the nation overall, although these differences were modest.
Application of findings to research questions
Our inventory provides a comprehensive overview of environmental data available across C4R cohorts and highlights opportunities for future harmonization relevant to COVID research. The comprehensive characterization of pre- and mid-pandemic neighborhood conditions with sufficient geospatial variability and participant diversity can enable interdisciplinary research. Specifically, these pre- and mid-pandemic data allow the examination of environmental factors’ contribution to COVID risk and disparities, including on infection, morbidity, mortality, persistent symptoms, and vaccination uptake. Place-based differences in COVID impact by urbanization or city characteristics have been demonstrated across countries [37,38] and within the US [39]. Examinations of how neighborhoods impacted COVID outcomes informed previous prevention efforts, such as opening up park space or closing streets to create protected pedestrian spaces [40,41]. This work can also inform the creation of environments that are resilient to future public health disasters [39,42,43]. The geographic representation and sociodemographic diversity of the C4R cohorts, along with their combined sample size, permits research into how policies or neighborhoods contribute to COVID and related health disparities. Emergency measures in other domains implemented by some locales, such as eviction moratoriums, could be helpful to public health and used as a basis for future policies to protect populations [39,44–47].
The geographic and individual diversity of C4R could also be leveraged to understand how neighborhoods and policies differentially impact COVID and other health outcomes by participant characteristics such as race and ethnicity, sex/gender, age, and socioeconomic status. This may help identify the conditions and populations for which policies are most effective. Ultimately, while the COVID pandemic intensity wanes, lessons learned about neighborhoods during the COVID pandemic and recovery period can be applied to enhance future disaster preparedness [48]. However, this can only happen if efforts are made to integrate individual-level risk information to avoid both ecological bias [48] and confounding by prior conditions. Longitudinal data from C4R helps to overcome some of these limitations, allowing for better control of changing ecological conditions over time.
C4R cohorts’ neighborhood environment and contextual data could also be leveraged to understand critical shifts that occurred in built, social, and natural environments post-pandemic and their subsequent health impacts. These may be primary shifts that occurred early during the pandemic (e.g., policies around eviction, stay-home orders, pedestrianized streets) or enduring changes to retail environments, physical spaces, economic conditions, or population distributions. Evidence suggests that early pandemic prevention measures lowered pollution levels, limited access to public spaces, closed retail establishments, and produced lost income [49–51]. A number of cities made intentional health-supportive changes to their environments (e.g., pedestrianized streets, parklets), with varying success retaining these as the pandemic has progressed [52–54]. Similar evidence has begun to emerge regarding shifts in health-related contexts and social environments, including crime [55–57], neighborhood trust or social cohesion [58,59], and movement patterns [60–62]. As the country and world ease into a “new normal,” it is essential to understand long-term impacts of the pandemic period on our lived experience, including neighborhood conditions. Further, the addition of novel forthcoming geospatial measures (e.g., segregation, resilience to climate disasters) could be integrated into existing datasets to encourage additional inquiries as conditions evolve. Connecting C4R neighborhood data to the extensive pre-pandemic (1975–2020) and planned (2020-present) deep phenotyping in the cohorts on physical function, cardiac measures, brain measures, sleep, biomarkers, and genotyping, among others, overcomes confounding by lack of knowledge regarding previous health behaviors and trajectories. Linking neighborhood data with these cohorts would provide an opportunity to understand how the COVID pandemic and related neighborhood or community changes affected other health conditions commonly assessed across the cohorts, such as cardiovascular risk, aging, and related diseases.
Future harmonization
Despite the myriad opportunities afforded by C4R neighborhood data, challenges remain for harmonization and future work. First, the pandemic disrupted data collection and infrastructure systems across the country and world. Specific to the US, there are known concerns about the quality and viability of Census 2020 and American Community Survey data that include the year 2020 (i.e., all five-year estimates from 2016–2025). This is similar for any other measure traditionally collected and processed in-person (e.g., building inspection measures), but is less likely to impact remote sensing measures (e.g., greenness). Furthermore, the validity and meaning of these measures may have changed due to pandemic restrictions (e.g., land uses designated for commercial purposes were closed and their public functions suspended). Second, although many cohorts have measures within a domain, these measures use different metrics, were collected at different levels of resolution (e.g., counties vs. census tracts), or were based on different underlying administrative and feature data. For example, some cohorts may source land use measures from parcel data collected in specific jurisdictions while others use national satellite imagery. Similarly, some may calculate intersection density for street networks while others use block length or a ratio of Network buffers to Euclidean buffers. Existing survey data are rich and comprehensive within each cohort but very sparse across cohorts, and without it, we may lack the ability to understand individuals’ perceptions of neighborhood amenities or culture pre-, mid-, and “post”-pandemic. Even within cohorts that collect survey data on perceived safety, social cohesion, and belonging, there are no signs of convergence toward a single standard set of tools. These inconsistencies make harmonization of existing data challenging. In addition, in several cohorts, neighborhood and spatial data were created, stored, and managed by ancillary studies rather than centrally within the coordinating center that currently interfaces with C4R. Harmonization is further hindered by the sensitive nature of spatial data; there are human subjects concerns when sharing geospatial data such as addresses and the various cohorts and university IRBs have different restrictions about the identifiability of participant latitudes-longitudes.
Despite privacy concerns, options for harmonization exist that circumvent some of these obstacles. One option is to create national-level metrics at a small administrative scale (i.e., block groups or tracts) that could be linked by each cohort to their participant locations. Alternatively, the creation of a national set of gridded metrics (i.e., as GIS rasters) would facilitate linkage to locations across a smooth, interpolated surface. These options create a balance between consistency across cohorts and ensuring data are backwards compatible, especially for measures where there is little to no ability to create historic metrics. While these workarounds may help with privacy concerns, some cohorts have limited internal resources, experience, and trained staff for managing geographic or spatial data to do the final linkage. Although harmonization is important to study neighborhood impact on COVID, pandemic impacts on everyday contexts, and their subsequent effects on health, these are only one arm of the many pressing scientific lines of inquiry possible for these cohorts. Ultimately, all cohorts are balancing different competing priorities and must remain accountable to their primary missions (e.g., cardiovascular, pulmonary), existing ancillary studies, and participants. Harmonization will require considerable financial resources across an extended timeline, and execution by a committed and experienced team of diverse researchers.
Limitations and strengths
This paper provides a comprehensive description of the neighborhood and spatial data across cohorts but is not without limitations. First, many cohorts were missing some of the details on software for geocoding and, as in many transdisciplinary studies, we needed to establish common language around geography terminology and concepts (e.g., administrative boundary, uncertain geographic context problem, modifiable areal unit problem). Missing technical details may have occurred when neighborhood measures were compiled for ancillary studies but not distributed through the broader network of cohort researchers. When inventorying neighborhood data, the use of a survey and self-report by staff within cohort coordinating centers may introduce recall bias. We minimized these issues by collecting data documentation to confirm reported survey data and looking for published manuscripts using measures of interest. To facilitate harmonization, future work should prioritize documentation of software details and centralization of ancillary study datasets, code, and documentation at coordinating centers. In addition, as stated in our results, the inventory is a snapshot of what cohorts had in February-November 2022 and does not represent projects underway. For example, some cohorts had recently funded grants to incorporate new measures that were not yet calculated at the time of this inventory. We have tried our best to capture these forthcoming measures in the Supplemental files but may have missed upcoming research.
Our calculation of spatial extent has limitations which may introduce ecological bias and related issues, including modifiable area unit problem. Due to differences in human subject restrictions or levels of geocoding, cohorts sent residential location data at different spatial scales (e.g., cities, counties, census tracts) and across different spatial vintages (e.g., census 2000 tracts, census 2010 tracts). For anonymity, we only present data aggregated across all cohorts (i.e., counts of cohorts in a location, not the named cohort) and some cohorts additionally aggregated their data across time, providing all locations that a participant has lived across any exam. Additionally, cohorts with older adults (e.g., ARIC, MESA) have some participants who live in the southern US during the wintertime, and those wintertime locations are not necessarily captured in our presented data. Therefore, the data we presented on spatial representation combines time, spatial unit, and cohort. As such, the map of the geographic extent of C4R participants does not necessarily represent participants’ locations during the COVID-19 pandemic and the comparison of this extent to national demographics may include some spatial and/or temporal misclassification that could introduce bias. For research on specific exposure-outcome pairs, future work should try to gain approval and access to participant data including their COVID-19 specific outcomes and data collected at a spatial and temporal resolution best fitting the study’s specific aims, which would help to minimize ecological bias (e.g., modifiable area unit problem, ecological fallacy). The map also represents residential areas of C4R without weighting to account for differential counts of participants for given geographies (e.g., some census tracts may have one participant, some may have many). Similarly, our calculations of sociodemographic and COVID measures for the geographic extent required reaggregation of the spatial scale at which those data were provided (e.g., census tract for census, county for COVID data). This assumes an even distribution of each characteristic across the areal unit. Furthermore, we were limited to counties for national COVID data, and different reporting issues occurred at different points in the pandemic (lack of access to testing early in the pandemic, use of at-home tests later in the pandemic that may have not been reported to the city, etc.). Finally, to ease interpretation, all data presented are summaries and do not represent harmonized measures across the cohorts. More specific details on the availability of geospatial data for individual cohorts should be requested from the authors or directly from each cohort; access to these data should be requested through the approved processes of each cohort.
Despite these limitations, several strengths enhance the utility of this paper to guide future research on how environmental context shapes health outcomes and disparities. First, we used a centralized, systematic method for cataloging available geospatial data through a uniform data collection tool and process across all cohorts. Our graphical display of these data allows researchers to see potential synergies that would otherwise be difficult (or impossible) to recognize. In addition to sparking novel research questions and enabling well-powered analyses across diverse samples and regions, our clear summarization could be used as a model for other projects. Our method could be applied to those seeking to harmonize data across cohorts or as a guide for the creation of new measures (e.g., noise) that could be developed for one or more cohorts. Second, capturing key elements across these metrics, including differences in how data were collected (e.g., survey data vs. GIS data), allows researchers interested in utilizing these data to be informed about the inconsistencies across cohorts. This points researchers toward potential limitations of their analyses and provides information to interpret disparate study results across cohorts. Third, our systematic approach provides a roadmap for researchers who would like to collect their own geospatial data. For example, our collection process has already led to discussions across cohorts regarding best measures for specific environmental features. We believe that this summarization will encourage selection and operationalization of measures and methods that increase the coherence of the evidence created. This would enable even more scientific rigor and lead to the creation of larger and more diverse datasets that can answer highly complex questions about the intersection of environment and human health. Finally, our diverse and interdisciplinary team of national leaders in neighborhood health demonstrates best practices for team science. Our process ensured that multiple viewpoints were represented, ultimately encouraging shared methods and promoting collaborations with a wide range of disciplines with great potential to improve public health.
Conclusion
In this study, we cataloged the geographic and spatial data available across a large geographically, socioeconomically, racially and ethnically diverse study of 14 pooled cohorts in the US. Our analyses described the spatial extent of the C4R population and examined representativeness compared to areas without C4R participants and compared to the nation overall, finding that C4R had relatively representative coverage of urban areas nationwide. We found that social environmental data were the most common types of contextual data collected, followed by built environment features. Our inventory also indicated that all cohorts were missing neighborhood COVID contextual data, several cohorts were lacking any geocoding or environmental data, and that even cohorts with contextual data had gaps in natural environment metrics, such as weather or heat, which may be relevant in an age of increasing climate-related events and vector-borne diseases [63–66]. Despite numerous advances in the assessment of neighborhood effects on health, our inventory of neighborhood measures highlights the need to develop a systematic methodology to facilitate harmonization. Agreement on common metrics, repositories of national environmental measures, and guidance documents for cohorts would encourage and enable valid and rigorous testing of the health effects of neighborhoods or their subsequent impacts on health disparities. Given the large quantity of data within the C4R project, methods will also need to be developed to integrate data across multiple spatial and temporal scales. Once harmonized, the associated data and methods could be shared widely to facilitate international collaborations and pooling of data for large-scale research projects (e.g., with the International Human Exposome Network (IHEN)). Work remains to understand the implications of the COVID pandemic on our neighborhoods, the way we interact with them, and their subsequent impact on our health and well-being. During a time when the environmental context is rapidly changing due to human activities and substantial climate change shifts, lessons learned from the pandemic period may help us create more resilient, health-supportive, equitable communities. This project opens possibilities for research collaborations to understand the health effects of neighborhood changes experienced during the pandemic and work toward sustainable, thoughtful neighborhood interventions to support health and well-being as future disasters arise.
Supporting information
S1 Table. Basic Characteristics of the 14 C4R cohorts March 1, 2020.
https://doi.org/10.1371/journal.pone.0352170.s001
(DOCX)
S2 Table. Cohort Institutional Review Boards (IRBs) supervising implementation of the C4R protocols.
Columbia University serves as the data coordinating center for C4R (IRB-AAAT3035). The study received initial IRB approval on December 4, 2020, with approval continuing through the current date (August 8, 2025). To date, four approval letters have been issued by the Columbia University IRB; these are provided in Supporting Document S5. This file is a list of site-specific IRB approval numbers for each study for the duration of the study period.
https://doi.org/10.1371/journal.pone.0352170.s002
(DOCX)
S3 Table. Forthcoming data not inventoried in this paper.
We surveyed the cohorts’ coordinating centers on or before November 2022. Several cohorts reported geospatial measures being developed by investigators (through funded projects, ancillary studies, or approved manuscript proposals). These measures do not appear in this manuscript as they are not yet readily available for outside researchers through the coordinating centers.
https://doi.org/10.1371/journal.pone.0352170.s003
(DOCX)
S4 File. Appendix of the Survey to Inventory Geospatial and Neighborhood Data in C4R.
This form was used by our team during data collection from cohort coordinating centers and administrators.
https://doi.org/10.1371/journal.pone.0352170.s004
(DOCX)
Acknowledgments
The authors would like to thank Martha Daviglus, Priya Palta, Neil Schneiderman, and Ann Chang for their contributions to these efforts. This manuscript has been reviewed by CARDIA for scientific content.
SPIROMICS co-authors would like to acknowledge the following current and former investigators of the SPIROMICS sites and reading centers: Neil E Alexis, MD; Wayne H Anderson, PhD; Mehrdad Arjomandi, MD; Igor Barjaktarevic, MD, PhD; R Graham Barr, MD, DrPH; Patricia Basta, PhD; Lori A Bateman, MS; Christina Bellinger, MD; Surya P Bhatt, MD; Eugene R Bleecker, MD; Richard C Boucher, MD; Russell P Bowler, MD, PhD; Russell G Buhr, MD, PhD; Stephanie A Christenson, MD; Alejandro P Comellas, MD; Christopher B Cooper, MD, PhD; David J Couper, PhD; Gerard J Criner, MD; Ronald G Crystal, MD; Jeffrey L Curtis, MD; Claire M Doerschuk, MD; Mark T Dransfield, MD; M Bradley Drummond, MD; Christine M Freeman, PhD; Craig Galban, PhD; Katherine Gershner, DO; MeiLan K Han, MD, MS; Nadia N Hansel, MD, MPH; Annette T Hastie, PhD; Eric A Hoffman, PhD; Yvonne J Huang, MD; Robert J Kaner, MD; Richard E Kanner, MD; Mehmet Kesimer, PhD; Eric C Kleerup, MD; Jerry A Krishnan, MD, PhD; Wassim W Labaki, MD; Lisa M LaVange, PhD; Stephen C Lazarus, MD; Fernando J Martinez, MD, MS; Merry-Lynn McDonald, PhD; Deborah A Meyers, PhD; Wendy C Moore, MD; John D Newell Jr, MD; Elizabeth C Oelsner, MD, MPH; Jill Ohar, MD; Wanda K O'Neal, PhD; Victor E Ortega, MD, PhD; Robert Paine, III, MD; Laura Paulin, MD, MHS; Stephen P Peters, MD, PhD; Cheryl Pirozzi, MD; Nirupama Putcha, MD, MHS; Sanjeev Raman, MBBS, MD; Stephen I Rennard, MD; Donald P Tashkin, MD; J Michael Wells, MD; Robert A Wise, MD; and Prescott G Woodruff, MD, MPH. The project officers from the Lung Division of the National Heart, Lung, and Blood Institute were Lisa Postow, PhD, and Lisa Viviano, BSN. The authors also acknowledge the University of North Carolina at Chapel Hill BioSpecimen Processing Facility and Alexis Lab for sample processing, storage, and sample disbursements.
In addition to funding directly awarded to co-authors for this manuscript, many grants, contracts, cooperative agreements, institutional resources, foundation contributions, and industry contributions have supported these C4R cohorts over time. ARIC: HHSN268201700001I, HHSN268201700002I, HHSN268201700003I, HHSN268201700005I, HHSN268201700004I; 2U01HL096812, 2U01HL096814, 2U01HL096899, 2U01HL096902, 2U01HL096917; R01NS102715. CARDIA parent study contracts: HHSN268201800005I, HHSN268201800007I, HHSN268201800003I, HHSN268201800006I, HHSN268201800004I. COPDGene: U01HL089897, U01HL089856. COPDGene is also supported by the COPD Foundation through contributions made to an Industry Advisory Board comprising AstraZeneca, Boehringer-Ingelheim, Genentech, GlaxoSmithKline, Novartis, Pfizer, Siemens, and Sunovion. FHS: NO1-HC-25195, HHSN268201500001I, 75N92019D00031; R01HL109263; P01ES009825; P01AG031093. HCHS/SOL parent study contracts: HHSN268201300001I, N01-HC-65233, HHSN268201300004I, N01-HC-65234, HHSN268201300002I, N01-HC-65235, HHSN268201300003I, N01-HC-65236, HHSN268201300005I, N01-HC-65237; R01HL148463. The following Institutes/Centers/Offices have contributed to HCHS/SOL through a transfer of funds to the NHLBI: National Institute on Minority Health and Health Disparities, National Institute on Deafness and Other Communication Disorders, National Institute of Dental and Craniofacial Research, National Institute of Diabetes and Digestive and Kidney Diseases, National Institute of Neurological Disorders and Stroke, and NIH Office of Dietary Supplements. JHS parent study contracts: HHSN268201800013I, HHSN268201800014I, HHSN268201800015I, HHSN268201800010I, HHSN268201800011I, HHSN268201800012I. MASALA institutional/CTSI support: UL1RR024131. MESA parent study / SHARe / TOPMed / CTSA / pooled cohort infrastructure: 75N92020D00001, HHSN268201500003I, N01-HC-95159, 75N92020D00005, N01-HC-95160, 75N92020D00002, N01-HC-95161, 75N92020D00003, N01-HC-95162, 75N92020D00006, N01-HC-95163, 75N92020D00004, N01-HC-95164, 75N92020D00007, N01-HC-95165, N01-HC-95166, N01-HC-95167, N01-HC-95168, N01-HC-95169, R01HL077612, R01HL093081, R01HL130506, R01HL127028, R01HL127659, R01HL098433, R01HL101250, R01HL135009, R01AG058969, UL1TR000040, UL1TR001079, UL1TR001420, N02-HL-64278, 3U54HG003067-13S1, 3R01HL117626-02S1, HHSN268201800002I, HHSN268201600034I, HHSN268201600032I, HHSN268201600038I, 3R01HL120393, U01HL120393, HHSN268180001I, UL1TR001881, DK063491, R21HL153700, K23HL130627, R21HL129924, R21HL121457, R01HL131610. NOMAS: R01AG066162. PrePF: X01ES101947, R01HL095393, RC2HL1011715, R21/33HL120770, R01HL097163, X01HL134585, UH2/3HL123442, P01HL092870, UG3/UH3HL151865, W81XWH-17-1-0597. REGARDS parent study / non-PI-author / institutional support: R01AA028552; Urban Health Collaborative at Drexel University; Built Environment and Health Research Group at Columbia University. SARP: U10HL109164, U10HL109257, U10HL109146, U10HL109172, U10HL109250, U10HL109168, U10HL109152. Industry partnerships also provided additional support from AstraZeneca, Boehringer-Ingelheim, Genentech, GlaxoSmithKline, MedImmune, Novartis, Regeneron, Sanofi, and TEVA. Spirometers used in SARP III were provided by nSpire Health. SPIROMICS parent / non-PI-author / industry support: HHSN268200900013C, HHSN268200900014C, HHSN268200900015C, HHSN268200900016C, HHSN268200900017C, HHSN268200900018C, HHSN268200900019C, HHSN268200900020C, U01HL137880, U24HL141762, R01HL182622, R01HL144718. SPIROMICS was supplemented by contributions made through the Foundation for the NIH and the COPD Foundation from Amgen; AstraZeneca/MedImmune; Bayer; Bellerophon Therapeutics; Boehringer-Ingelheim Pharmaceuticals, Inc.; Chiesi Farmaceutici S.p.A.; Forest Research Institute, Inc.; Genentech; GlaxoSmithKline; Grifols Therapeutics, Inc.; Ikaria, Inc.; MGC Diagnostics; Novartis Pharmaceuticals Corporation; Nycomed GmbH; Polarean; ProterixBio; Regeneron Pharmaceuticals, Inc.; Sanofi; Sunovion; Takeda Pharmaceutical Company; and Theravance Biopharma and Mylan/Viatris. SHS: 75N92019D00027, 75N92019D00028, 75N92019D00029, 75N92019D00030, R01HL109315, R01HL109301, R01HL109284, R01HL109282, R01HL109319, U01HL41642, U01HL41652, U01HL41654, U01HL65520, U01HL65521.
We also acknowledge the dedication of other investigators, staff, and participants of all C4R cohorts, without whom this research would not be possible
References
- 1. Bryan MS, Sun J, Jagai J, Horton DE, Montgomery A, Sargis R, et al. Coronavirus disease 2019 (COVID-19) mortality and neighborhood characteristics in Chicago. Ann Epidemiol. 2021;56:47–54.e5.
- 2. McGowan VJ, Bambra C. COVID-19 mortality and deprivation: pandemic, syndemic, and endemic health inequalities. Lancet Public Health. 2022;7(11):e966–75. pmid:36334610
- 3. Mude W, Oguoma VM, Nyanhanda T, Mwanri L, Njue C. Racial disparities in COVID-19 pandemic cases, hospitalisations, and deaths: A systematic review and meta-analysis. J Glob Health. 2021;11:05015.
- 4. Alidadi M, Sharifi A. Effects of the built environment and human factors on the spread of COVID-19: A systematic literature review. Sci Total Environ. 2022;850:158056. pmid:35985590
- 5. Chen JT, Krieger N. Revealing the unequal burden of COVID-19 by income, race/ethnicity, and household crowding: US county versus zip code analyses. Journal of Public Health Management and Practice. 2021;27(1):S43–56.
- 6. Frumkin H. COVID-19, the Built Environment, and Health. Environ Health Perspect. 2021;129(7):75001. pmid:34288733
- 7. Weaver AK, Head JR, Gould CF, Carlton EJ, Remais JV. Environmental Factors Influencing COVID-19 Incidence and Severity. Annu Rev Public Health. 2022;43:271–91. pmid:34982587
- 8. Hamidi S, Sabouri S, Ewing R. Does density aggravate the COVID-19 pandemic? Early findings and lessons for planners. Journal of the American Planning Association. 2020;86(4):495–509.
- 9. Bourdrel T, Annesi-Maesano I, Alahmad B, Maesano CN, Bind M-A. The impact of outdoor air pollution on COVID-19: a review of evidence from in vitro, animal, and human studies. Eur Respir Rev. 2021;30(159):200242. pmid:33568525
- 10. Hernandez Carballo I, Bakola M, Stuckler D. The impact of air pollution on COVID-19 incidence, severity, and mortality: A systematic review of studies in Europe and North America. Environ Res. 2022;215(Pt 1):114155. pmid:36030916
- 11. Marquès M, Domingo JL. Positive association between outdoor air pollution and the incidence and severity of COVID-19. A review of the recent scientific evidences. Environ Res. 2022;203:111930. pmid:34425111
- 12. Klompmaker JO, Hart JE, Holland I, Sabath MB, Wu X, Laden F, et al. County-level exposures to greenness and associations with COVID-19 incidence and mortality in the United States. Environ Res. 2021;199:111331. pmid:34004166
- 13. Russette H, Graham J, Holden Z, Semmens EO, Williams E, Landguth EL. Greenspace exposure and COVID-19 mortality in the United States: January-July 2020. Environ Res. 2021;198:111195. pmid:33932476
- 14. Yang Y, Lu Y, Jiang B. Population-weighted exposure to green spaces tied to lower COVID-19 mortality rates: A nationwide dose-response study in the USA. Sci Total Environ. 2022;851(Pt 2):158333. pmid:36041607
- 15. Oelsner EC, Krishnaswamy A, Balte PP, Allen NB, Ali T, Anugu P, et al. Collaborative Cohort of Cohorts for COVID-19 Research (C4R) Study: Study Design. Am J Epidemiol. 2022;191(7):1153–73. pmid:35279711
- 16. Arcaya MC, Tucker-Seeley RD, Kim R, Schnake-Mahl A, So M, Subramanian SV. Research on neighborhood effects on health in the United States: A systematic review of study characteristics. Soc Sci Med. 2016;168:16–29. pmid:27637089
- 17. Brownson RC, Hoehner CM, Day K, Forsyth A, Sallis JF. Measuring the built environment for physical activity: state of the science. Am J Prev Med. 2009;36(4 Suppl):S99-123.e12. pmid:19285216
- 18. Lovasi GS, Grady S, Rundle A. Steps Forward: Review and Recommendations for Research on Walkability, Physical Activity and Cardiovascular Health. Public Health Rev. 2012;33(4):484–506. pmid:25237210
- 19. Schaefer-McDaniel N, Caughy MO, O’Campo P, Gearey W. Examining methodological details of neighbourhood observations and the relationship to health: a literature review. Soc Sci Med. 2010;70(2):277–92. pmid:19883966
- 20. Oelsner EC, Krishnaswamy A, Rustamov R, Balte PP, Ali T, Allen NB, et al. Classifying COVID-19 hospitalizations in epidemiology cohort studies: The C4R study. PLoS One. 2025;20(2):e0316198. pmid:39928595
- 21.
Manson S, et al. IPUMS National Historical Geographic Information System: Version 17.0. Minneapolis, MN: IPUMS University of Minnesota. 2022.
- 22. Krieger N, Waterman PD, Spasojevic J, Li W, Maduro G, Van Wye G. Public Health Monitoring of Privilege and Deprivation With the Index of Concentration at the Extremes. Am J Public Health. 2016;106(2):256–63. pmid:26691119
- 23. McAlexander TP, Algur Y, Schwartz BS, Rummo PE, Lee DC, Siegel KR, et al. Categorizing community type for epidemiologic evaluation of community factors and chronic disease across the United States. Soc Sci Humanit Open. 2022;5(1):100250. pmid:35369036
- 24.
Health Resources & Services Administration. Area Health Resources Files. Rockville, MD. 2022.
- 25.
Dewitz J. National Land Cover Database (NLCD) 2019 Products (ver 2.0, June 2021). Sioux Falls, SD: U.S. Geological Survey. 2021.
- 26. Dong E, Du H, Gardner L. An interactive web-based dashboard to track COVID-19 in real time. Lancet Infect Dis. 2020;20(5):533–4. pmid:32087114
- 27.
Centers for Disease Control and Prevention. COVID Data Tracker. Atlanta, GA: Department of Health and Human Services. 2022.
- 28. Hirsch JA, Moore KA, Cahill J, Quinn J, Zhao Y, Bayer FJ, et al. Business Data Categorization and Refinement for Application in Longitudinal Neighborhood Health Research: a Methodology. J Urban Health. 2021;98(2):271–84. pmid:33005987
- 29.
Economic Research Service. Rural-Urban Commuting Area Codes. Washington, DC: U.S. Department of Agriculture. 2022.
- 30. Mujahid MS, Diez Roux AV, Morenoff JD, Raghunathan T. Assessing the measurement properties of neighborhood scales: from psychometrics to ecometrics. Am J Epidemiol. 2007;165(8):858–67. pmid:17329713
- 31.
Applied Geographic Solutions. Thousand Oaks, CA: Applied Geographic Solutions. 2023.
- 32.
Environmental Protection Agency. Air data: Air quality data collected at outdoor monitors across the US. Research Triangle Park, NC: Environmental Protection Agency. 2023.
- 33. Kirwa K, Szpiro AA, Sheppard L, Sampson PD, Wang M, Keller JP, et al. Fine-Scale Air Pollution Models for Epidemiologic Research: Insights From Approaches Developed in the Multi-ethnic Study of Atherosclerosis and Air Pollution (MESA Air). Curr Environ Health Rep. 2021;8(2):113–26. pmid:34086258
- 34.
Trust for Public Land. ParkServe. Los Angeles, CA. 2023.
- 35. Casey JA, James P, Cushing L, Jesdale BM, Morello-Frosch R. Race, Ethnicity, Income Concentration and 10-Year Change in Urban Greenness in the United States. Int J Environ Res Public Health. 2017;14(12):1546. pmid:29232867
- 36.
National Centers for Environmental Information. Local Climatological Data (LCD). Silver Spring, MD: National Oceanic and Atmospheric Administration. 2023.
- 37. Naudé W, Nagler P. COVID-19 and the city: Did urbanized countries suffer more fatalities?. Cities. 2022;131:103909. pmid:35966968
- 38. González‐Val R, Sanz‐Gracia F. Urbanization and COVID‐19 incidence: A cross‐country investigation. Papers in Regional Science. 2022;101(2):399–416.
- 39. Bilal U, McCulley E, Li R, Rollins H, Schnake-Mahl A, Mullachery PH, et al. Tracking COVID-19 Inequities Across Jurisdictions Represented in the Big Cities Health Coalition (BCHC): The COVID-19 Health Inequities in BCHC Cities Dashboard. Am J Public Health. 2022;112(6):904–12. pmid:35420892
- 40. Firth CL, et al. Not quite a block party: COVID-19 street reallocation programs in Seattle, WA and Vancouver, BC. SSM-population health. 2021;14:100769.
- 41.
Douglas G, Moore D. Analyzing the use and impacts of Oakland slow streets and potential scalability beyond covid-19. 2022.
- 42. Liu C, Liu Z, Guan C. The impacts of the built environment on the incidence rate of COVID-19: A case study of King County, Washington. Sustain Cities Soc. 2021;74:103144. pmid:34306992
- 43. Kashem SB, Baker DM, González SR, Lee CA. Exploring the nexus between social vulnerability, built environment, and the prevalence of COVID-19: A case study of Chicago. Sustain Cities Soc. 2021;75:103261. pmid:34580620
- 44. Leifheit KM, Linton SL, Raifman J, Schwartz GL, Benfer EA, Zimmerman FJ, et al. Expiring Eviction Moratoriums and COVID-19 Incidence and Mortality. Am J Epidemiol. 2021;190(12):2503–10. pmid:34309643
- 45. Sandoval-Olascoaga S, Venkataramani AS, Arcaya MC. Eviction Moratoria Expiration and COVID-19 Infection Risk Across Strata of Health and Socioeconomic Status in the United States. JAMA Netw Open. 2021;4(8):e2129041. pmid:34459904
- 46. Versey HS. The impending eviction cliff: housing insecurity during COVID-19. American Public Health Association. 2021;:1423–7.
- 47. Ali AK, Wehby GL. State Eviction Moratoriums During The COVID-19 Pandemic Were Associated With Improved Mental Health Among People Who Rent: Study examines associations between the mental health of renters and state eviction moratoriums during COVID-19. Health Affairs, 2022;41(11):1583–9.
- 48.
Bakerjian D, Nguyen A. Lessons learned from coronavirus disease 2019 recovery: policy implications for the health and well-being of older adults. Oxford University Press US. 2022.
- 49. Patrício Silva AL, Prata JC, Walker TR, Duarte AC, Ouyang W, Barcelò D, et al. Increased plastic pollution due to COVID-19 pandemic: Challenges and recommendations. Chem Eng J. 2021;405:126683. pmid:32834764
- 50. Berman JD, Ebisu K. Changes in U.S. air pollution during the COVID-19 pandemic. Sci Total Environ. 2020;739:139864. pmid:32512381
- 51. Dang HAH, Nguyen CV. Gender inequality during the COVID-19 pandemic: Income, expenditure, savings, and job loss. World Development. 2021;140:105296.
- 52. Rhoads D, et al. A sustainable strategy for open streets in (post) pandemic cities. Communications Physics. 2021;4(1):183.
- 53.
Rhoads D, et al. Planning for sustainable open streets in pandemic cities. In: 2020. https://doi.org/arXiv:2009.12548
- 54. Rojas-Rueda D, Morales-Zamora E. Built environment, transport, and COVID-19: a review. Current Environmental Health Reports. 2021;8(2):138–45.
- 55. Abrams DS. COVID and crime: An early empirical look. J Public Econ. 2021;194:104344. pmid:33518828
- 56. Ashby MPJ. Initial evidence on the relationship between the coronavirus pandemic and crime in the United States. Crime Sci. 2020;9(1):6. pmid:32455094
- 57. Boman JH, Gallupe O. Has COVID-19 changed crime? Crime rates in the United States during the pandemic. American Journal of Criminal Justice. 2020;45:537–45.
- 58. Jiang N, Wu AM, Cheng EW. Social trust and stress symptoms among older adults during the COVID-19 pandemic: evidence from Asia. BMC Geriatr. 2022;22(1):330. pmid:35428191
- 59.
Zangger C. Help thy neighbor. Neighborhood relations, subjective well-being, and trust during the COVID-19 pandemic. SocArXiv. 2021. 19.
- 60.
Haslag PH, Weagley D. From LA to Boise: How migration has changed during the COVID-19 pandemic. In: 2022. https://ssrn.com/abstract=3808326
- 61. González-Leonardo M, Rowe F, Fresolone-Caparrós A. Rural revival? The rise in internal migration to rural areas during the COVID-19 pandemic. Who moved and where?. Journal of Rural Studies. 2022;96:332–42.
- 62.
Coven JA. Gupta IY. JUE Insight: Urban flight seeded the COVID-19 pandemic across the United States. Journal of urban economics, 2023. 133: p. 103489.
- 63. Caminade C, McIntyre KM, Jones AE. Impact of recent and future climate change on vector‐borne diseases. Annals of the New York Academy of Sciences. 2019;1436(1):157–73.
- 64. Schiermeier Q. Europe’s mega-heatwave boosted by climate change. Nat. 2019;571(7764):155.
- 65. Strauss BH, Orton PM, Bittermann K, Buchanan MK, Gilford DM, Kopp RE, et al. Economic damages from Hurricane Sandy attributable to sea level rise caused by anthropogenic climate change. Nat Commun. 2021;12(1):2720. pmid:34006886
- 66. Williams AP, et al. Observed impacts of anthropogenic climate change on wildfire in California. Earth’s Future. 2019;7(8):892–910.