Published February 6, 2020 | Version v1
Dataset Open

Dataset of Grouped Commit Author IDs after Identity Resolution

  • 1. University of Tennessee

Description

This Dataset contains the SHA1 values of IDs for 5,427,024 commit authors who have created commits in git version control system, and have more than 1 ID in git. It is a compressed CSV file (separated by ; ) with 14,861,538 author IDs, where the first column is the group ID, which is same as the first (randomly selected) author ID of the group, and the second column is the author ID that is part of the group. If an author was found to have 2 different IDs: I1, I2, then it is recorded in the file in 2 separate lines, with the lines being I1;I1 and I1;I2, i.e. the first column is the group identifier, which is one of the IDs in a group, and the second column contains the different author IDs in separate lines. Author IDs consist of the Author's name and email address in the format: Name <Email>.

Files

Files (3.4 GB)

Name Size Download all
md5:55179ecb58b7d38122682bc44753bb4f
3.4 GB Download