Published September 22, 2021 | Version v1

Re-identification of Individuals in Genomic Datasets Using Public Face Images

  • 1. Washington University in St. Louis
  • 2. Vanderbilt University

Description

Image-genome pairs in these synthetic datasets were created by combining a subset of the publicly available face image dataset, CelebA, and genotypes from OpenSNP. The genome in a given pair does not correspond to the individual in the image (taken from CelebA), but comes instead from an individual with the same set of phenotypes (taken from OpenSNP). Artificial genotypes were created for each image (genotype refers only to the small subset of SNPs we are interested in) using all available data from OpenSNP where self-reported phenotypes are present.

In the Synthetic-Ideal dataset, to each image, we assigned a genotype from OpenSNP that corresponds to an individual with the same phenotypes, such that the probability of the selected phenotypes is maximized, given the genotype. In other words, we picked the genotype from the OpenSNP data that is most representative of an individual with a given set of phenotypes.

In the Synthetic-Realistic dataset, to each image, we assigned a genotype from OpenSNP that corresponds to an individual with the same phenotypes, but at random according to the empirical distribution of phenotypes for particular SNPs in our data.

Since CelebA does not have labels for all considered phenotypes, 1000 images from this dataset were manually labeled by one of the authors. After cleaning and removing ambiguous cases, the resulting datasets consist of 456 records.

Notes

National Institutes of Health (Grant RM1HG009034) National Science Foundation (Grant IIS-1905558)

Files

GenomicReID.zip

Files (22.6 GB)

Name Size
md5:1c3044a30e0b20497e20513b6229bcd5
22.6 GB Preview Download