Published August 19, 2024 | Version v1

System Fingerprint Recognition for Deepfake Audio (SFR) - Compressed Set

Description

The rapid progress of deep speech synthesis models has posed significant threats to society such as malicious manip
ulation of content. This has led to an increase in studies aimed at detecting so-called “deepfake audio”. However, existing works
focus on the binary detection of real audio and fake audio. In real-world scenarios such as model copyright protection and
digital evidence forensics, it is needed to know what tool or model generated the deepfake audio to explain the decision. This
motivates us to ask: ‘Can we recognize the system fingerprints of deepfake audio?’ In this paper, we present the first deepfake
audio dataset for System Fingerprint Recognition (SFR) and conduct an initial investigation. We collected the dataset from
the speech synthesis systems of seven Chinese vendors that use the latest state-of-the-art deep learning technologies, including
both clean and compressed sets. In addition, we provide extensive benchmarks and research findings to facilitate the further development of system fingerprint recognition methods. The dataset is publicly available. 
 
The subsets 01, 02, and 03 represent the training set, development set, and test set, respectively.
 
This data set is licensed with a CC BY-NC-ND 4.0 license.

Files

compressed_01.zip

Files (36.9 GB)

Name Size
md5:1bfa77c6070adf5cc2136e64007bed53
20.5 GB Preview Download
md5:8024bc2cbffcb8e4bbac0a5f8ec18314
6.2 GB Preview Download
md5:a6bb9044658ac5bc41c00774803ee668
10.1 GB Preview Download

Additional details

References

  • Yan X, Yi J, Wang C, et al. System fingerprint recognition for deepfake audio: an initial dataset and investigation[J]. arXiv preprint arXiv:2208.10489, 2022.