Published July 19, 2026 | Version v8

Synthetic U.S. Spatiotemporal Covid-19 Datasets Generated Through USPopulationSampler Package

  • 1. ROR icon University of California, Los Angeles

Description

Dataset Summaries

This repository serves as the official data companion to the USPopulationSampler package. It provides standardized, pre-processed datasets designed to facilitate high-resolution geographic sampling, demographic analysis, machine learning, and epidemiological modeling. When the package is loaded into R, users have access to the following datasets described below: 

 

1. Synthetic Spatiotemporal Covid-19 Datasets 

Description: A collection of 28 simulated epidemiological datasets designed to mimic the statistical properties, temporal trends, and transmission dynamics of the actual pandemic that were generated by the USPopulationSampler package stored as parquet files. Each dataset comprises of randomly generated and independent spatiotemporal data with the locations being generated uniformly in census blockgroups which themselves are randomly selected according to population counts within a given county.  Datasets were generated from 5/27/2026 to 5/28/2026. 

The dimension of each synthetic data is 100,501,034 rows and 5 columns with each row representing a case of Covid-19. The data columns include longitude, latitude, date of each case from 2020 to 2023, the fips identifier code, and the geo_id. 

Number of variables: 

Number of cases/rows: 100,501,034

Variable List: 

Variable Name Description Unit/Format
lon Longitude coordinate of sampled location point Decimal Degrees (WGS84)
lat Latitude coordinate of sampled location point  Decimal Degrees (WGS84)
fips Five-digit Federal Information Processing Standards (FIPS) county code  Character String 
GEO_ID Census Block Group Geographic Identifier (GEOID) corresponding to the sampled location Character String 
date Assigned observed date of the sampled location point YYYY-MM-DD

2. Processed Covid-19 Dataset 

Description: A curated, cleaned historical record of the Covid-19 pandemic in the United States, spanning from its onset in 2020 through 2023. The original raw data was taken from the https://github.com/nytimes/covid-19-data. 

Number of variables: 4

Number of cases/rows: 100,501,034

Variable List:

Variable Name Description Unit/Format
fips Five-digit Federal Information Processing Standards (FIPS) county code  Character String 
date Recorded date of the Covid-19 case  YYYY-MM-DD
county Name of U.S county in which Covid-19 case was recorded  Character String
state Name of the U.S state/territory in which Covid-19 case was recorded  Character String

3. USPopulationSampler Reference Data 

Description: Embedded directly within or complemented by the package, this dataset provides a highly granular geographic and demographic baseline of the United States based on the official published 2020 Decennial Census - Census Block Maps. The data goes down to the block group level which is the smallest geographic unit for which the Census Bureau publishes population data).

Number of variables: 9

Number of cases/rows: 242,335

Variable Name Description Unit/Format
GEO_ID Census Block Group Geographic Identifier (GEOID) (21-digit identifier) Character String 
NAMELSAD Recorded date of the Covid-19 case  YYYY-MM-DD
STATEFP Name of U.S county in which Covid-19 case was recorded (2-digit identifier) Character String
COUNTYFP Name of the U.S state/territory in which Covid-19 case was recorded (3-digit identifier) Character String
TRACTCE Census Tract Identifier within the county (6-digit identifier) Character String
BLKGRPCE Census blockgroup identifier within the tract (1-digit identifier) Character String
geometry Polygon geometry representing the spatial boundary of the Census Block Group Multipolygon geometry (WGS84 / EPSG:4326)
pop Population of county blockgroups recorded as number of individuals in that blockgroup Numeric 
FIPS Five-digit Federal Information Processing Standards (FIPS) county code  Character String 

 

Specialized formats or other abbreviations used:

FIPS = Federal Information Processing Standards geographic code identifying U.S states and territories and counties.

GEOID = Census Geographic Identifier. Values are Census Block Group identifiers in the format 1500000USSSCCCTTTTTTB, where:

  • ss = state FIPS code (2-digit)
  • ccc = county FIPS code (3-digit)
  • TTTTTT = census tract code (6-digit)
  • B = blockgroup number

lat = Latitude in decimal degrees (WGS84)

lon = Longitude in decimal degrees (WGS84) 

Files

Files (49.2 GB)

Name Size
md5:5b01c78b5c8e96c3687a867a2e094bd6
15.2 MB Download
md5:57402a17dfd78a8aceec775ebab6fd98
1.8 GB Download
md5:206672946863f072e599369ef971877c
1.8 GB Download
md5:4a4493ba0d82f0d8f2b30cd899bb36ec
1.8 GB Download
md5:9213e00d8aed7f6705a386022d2fdd12
1.8 GB Download
md5:707b574164790e37b3349405d6110c66
1.8 GB Download
md5:1471056ae1d86890bcdf28cdcee91532
1.8 GB Download
md5:b94cf0d283a16f99b4f89f19e0b7dbb3
1.8 GB Download
md5:f45c4bbe5c96e8104bfd5bd1d26d4848
1.8 GB Download
md5:4b1507c8173a326e458850017c6a2ec9
1.8 GB Download
md5:9e4f470dfcceea7fdbe3054901135e18
1.8 GB Download
md5:2f75e6ec5da26fdbd2a199e5183ab274
1.8 GB Download
md5:0afb3631712a461403bc85054217d28d
1.8 GB Download
md5:a83e44f22c252c012a02457703dac9fd
1.8 GB Download
md5:d1f6527050428872541159b2673533fd
1.8 GB Download
md5:f90759bb709ed922b63f6072b3cb59fa
1.8 GB Download
md5:1ca57366125f6fb729a9066771397ba0
1.8 GB Download
md5:28c6c3ed9596503bbe1deb4cad705d46
1.8 GB Download
md5:97dbf774c61583743d5462043c6bfe83
1.8 GB Download
md5:b6022f68ca7268a8432389488883c311
1.8 GB Download
md5:94544f8885e916a739ed3a12710ec451
1.8 GB Download
md5:622c685b82076d1fef0fbc09fc6e9783
1.8 GB Download
md5:36992e3c50ad78557ce4bbac0eebd179
1.8 GB Download
md5:6b5f270affe0622d91a09fbeb1326ed2
1.8 GB Download
md5:63a7a2d0a262f5c3c87aca1ce1e4ddde
1.8 GB Download
md5:9659eee718f297499f306e965bb53e83
1.8 GB Download
md5:0f2017f8885d43fbffbd114ab0bb47e9
1.8 GB Download
md5:f7b98284e667738533cfdd4f64af7597
1.8 GB Download
md5:8cd62b4350d25345a4b00c5c9379d1bf
1.8 GB Download
md5:371bdac5cbba21acc32d6c6449df9f05
78.4 MB Download

Additional details

Funding

National Institutes of Health
Deep Bayesian Biostatistics: AI-Accelerated Inference for Big "N" and Big "P" R35 GM159431
U.S. National Science Foundation
CAREER: Data-Centric Evolutionary Contagion Models with Parallel and Quantum Parallel Computing DMS 2236854

Software

Repository URL
https://github.com/Techavoan/USPopulationSampler
Programming language
R
Development Status
Active