Synthetic U.S. Spatiotemporal Covid-19 Datasets Generated Through USPopulationSampler Package
Authors/Creators
Description
Dataset Summaries
This repository serves as the official data companion to the USPopulationSampler package. It provides standardized, pre-processed datasets designed to facilitate high-resolution geographic sampling, demographic analysis, machine learning, and epidemiological modeling. When the package is loaded into R, users have access to the following datasets described below:
1. Synthetic Spatiotemporal Covid-19 Datasets
Description: A collection of 28 simulated epidemiological datasets designed to mimic the statistical properties, temporal trends, and transmission dynamics of the actual pandemic that were generated by the USPopulationSampler package stored as parquet files. Each dataset comprises of randomly generated and independent spatiotemporal data with the locations being generated uniformly in census blockgroups which themselves are randomly selected according to population counts within a given county. Datasets were generated from 5/27/2026 to 5/28/2026.
The dimension of each synthetic data is 100,501,034 rows and 5 columns with each row representing a case of Covid-19. The data columns include longitude, latitude, date of each case from 2020 to 2023, the fips identifier code, and the geo_id.
Number of variables: 5
Number of cases/rows: 100,501,034
Variable List:
| Variable Name | Description | Unit/Format |
| lon | Longitude coordinate of sampled location point | Decimal Degrees (WGS84) |
| lat | Latitude coordinate of sampled location point | Decimal Degrees (WGS84) |
| fips | Five-digit Federal Information Processing Standards (FIPS) county code | Character String |
| GEO_ID | Census Block Group Geographic Identifier (GEOID) corresponding to the sampled location | Character String |
| date | Assigned observed date of the sampled location point | YYYY-MM-DD |
2. Processed Covid-19 Dataset
Description: A curated, cleaned historical record of the Covid-19 pandemic in the United States, spanning from its onset in 2020 through 2023. The original raw data was taken from the https://github.com/nytimes/covid-19-data.
Number of variables: 4
Number of cases/rows: 100,501,034
Variable List:
| Variable Name | Description | Unit/Format |
| fips | Five-digit Federal Information Processing Standards (FIPS) county code | Character String |
| date | Recorded date of the Covid-19 case | YYYY-MM-DD |
| county | Name of U.S county in which Covid-19 case was recorded | Character String |
| state | Name of the U.S state/territory in which Covid-19 case was recorded | Character String |
3. USPopulationSampler Reference Data
Description: Embedded directly within or complemented by the package, this dataset provides a highly granular geographic and demographic baseline of the United States based on the official published 2020 Decennial Census - Census Block Maps. The data goes down to the block group level which is the smallest geographic unit for which the Census Bureau publishes population data).
Number of variables: 9
Number of cases/rows: 242,335
| Variable Name | Description | Unit/Format |
| GEO_ID | Census Block Group Geographic Identifier (GEOID) (21-digit identifier) | Character String |
| NAMELSAD | Recorded date of the Covid-19 case | YYYY-MM-DD |
| STATEFP | Name of U.S county in which Covid-19 case was recorded (2-digit identifier) | Character String |
| COUNTYFP | Name of the U.S state/territory in which Covid-19 case was recorded (3-digit identifier) | Character String |
| TRACTCE | Census Tract Identifier within the county (6-digit identifier) | Character String |
| BLKGRPCE | Census blockgroup identifier within the tract (1-digit identifier) | Character String |
| geometry | Polygon geometry representing the spatial boundary of the Census Block Group | Multipolygon geometry (WGS84 / EPSG:4326) |
| pop | Population of county blockgroups recorded as number of individuals in that blockgroup | Numeric |
| FIPS | Five-digit Federal Information Processing Standards (FIPS) county code | Character String |
Specialized formats or other abbreviations used:
FIPS = Federal Information Processing Standards geographic code identifying U.S states and territories and counties.
GEOID = Census Geographic Identifier. Values are Census Block Group identifiers in the format 1500000USSSCCCTTTTTTB, where:
- ss = state FIPS code (2-digit)
- ccc = county FIPS code (3-digit)
- TTTTTT = census tract code (6-digit)
- B = blockgroup number
lat = Latitude in decimal degrees (WGS84)
lon = Longitude in decimal degrees (WGS84)
Files
Files
(49.2 GB)
| Name | Size | |
|---|---|---|
|
md5:5b01c78b5c8e96c3687a867a2e094bd6
|
15.2 MB | Download |
|
md5:57402a17dfd78a8aceec775ebab6fd98
|
1.8 GB | Download |
|
md5:206672946863f072e599369ef971877c
|
1.8 GB | Download |
|
md5:4a4493ba0d82f0d8f2b30cd899bb36ec
|
1.8 GB | Download |
|
md5:9213e00d8aed7f6705a386022d2fdd12
|
1.8 GB | Download |
|
md5:707b574164790e37b3349405d6110c66
|
1.8 GB | Download |
|
md5:1471056ae1d86890bcdf28cdcee91532
|
1.8 GB | Download |
|
md5:b94cf0d283a16f99b4f89f19e0b7dbb3
|
1.8 GB | Download |
|
md5:f45c4bbe5c96e8104bfd5bd1d26d4848
|
1.8 GB | Download |
|
md5:4b1507c8173a326e458850017c6a2ec9
|
1.8 GB | Download |
|
md5:9e4f470dfcceea7fdbe3054901135e18
|
1.8 GB | Download |
|
md5:2f75e6ec5da26fdbd2a199e5183ab274
|
1.8 GB | Download |
|
md5:0afb3631712a461403bc85054217d28d
|
1.8 GB | Download |
|
md5:a83e44f22c252c012a02457703dac9fd
|
1.8 GB | Download |
|
md5:d1f6527050428872541159b2673533fd
|
1.8 GB | Download |
|
md5:f90759bb709ed922b63f6072b3cb59fa
|
1.8 GB | Download |
|
md5:1ca57366125f6fb729a9066771397ba0
|
1.8 GB | Download |
|
md5:28c6c3ed9596503bbe1deb4cad705d46
|
1.8 GB | Download |
|
md5:97dbf774c61583743d5462043c6bfe83
|
1.8 GB | Download |
|
md5:b6022f68ca7268a8432389488883c311
|
1.8 GB | Download |
|
md5:94544f8885e916a739ed3a12710ec451
|
1.8 GB | Download |
|
md5:622c685b82076d1fef0fbc09fc6e9783
|
1.8 GB | Download |
|
md5:36992e3c50ad78557ce4bbac0eebd179
|
1.8 GB | Download |
|
md5:6b5f270affe0622d91a09fbeb1326ed2
|
1.8 GB | Download |
|
md5:63a7a2d0a262f5c3c87aca1ce1e4ddde
|
1.8 GB | Download |
|
md5:9659eee718f297499f306e965bb53e83
|
1.8 GB | Download |
|
md5:0f2017f8885d43fbffbd114ab0bb47e9
|
1.8 GB | Download |
|
md5:f7b98284e667738533cfdd4f64af7597
|
1.8 GB | Download |
|
md5:8cd62b4350d25345a4b00c5c9379d1bf
|
1.8 GB | Download |
|
md5:371bdac5cbba21acc32d6c6449df9f05
|
78.4 MB | Download |
Additional details
Funding
- National Institutes of Health
- Deep Bayesian Biostatistics: AI-Accelerated Inference for Big "N" and Big "P" R35 GM159431
- U.S. National Science Foundation
- CAREER: Data-Centric Evolutionary Contagion Models with Parallel and Quantum Parallel Computing DMS 2236854
Software
- Repository URL
- https://github.com/Techavoan/USPopulationSampler
- Programming language
- R
- Development Status
- Active