Published December 14, 2025 | Version v1

HQ-MPSD: A Multilingual Benchmark for Partial Deepfake Speech Detection

Description

HQ-MPSD Dataset v1

 

The structure of dataset folder is below:

HQ-MPSD/
├── English/
│ ├── Bonafide/
│ │ ├──<speakerID>_<audiobookId>_<segmentID>.flac
│ │ └── ...
│ ├── Fully_Fake/
│ │ ├── <speakerID>_<audiobookId>_<segmentID>_f.flac
│ │ └── ...
│ ├── Partial_Fake_Clean/
│ │ ├── <speakerID>_<audiobookId>_<segmentID>_p.flac
│ │ └── ...
│ └── ...
│ ├── Partial_Fake_Noisy/
│ │ ├── <speakerID>_<audiobookId>_<segmentID>_pn.flac
│ │ └── ...
│ └── Frame_level_label.txt
│ └── Noise_augmentation_info.txt
├── French/
│ └── (same structure as English)
├── German/
│ └── (same structure as English)
└── ...

 

Frame labels are defined as:

  • 0 → Bonafide frame

  • 1 → Deepfake frame

  • 2 → Transition frame

Each row in the Frame_level_label.txt is formatted as follows: 
<id> <label_for_each_30ms_frame>

 

This file provides detailed provenance for noise-augmented audio samples, including:

  • the augmentation type or noise category

  • the specific noise file from OpenSLR26

  • the specific noise file from MUSAN

Each row in the Noise_augmentation_info.txt is formatted as follows:
<id> <augmentation_label> <used_file_path_in_OpneSLR26> <<used_file_path_in_MUSAN>

Files

Dutch.zip

Files (41.9 GB)

Name Size
md5:2c1d4b34a76e641587a887dad54c598b
1.1 GB Preview Download
md5:c89346355d9afb0ba8dca4247c35dbe6
3.2 GB Preview Download
md5:a643a8e8e8bcdc9ab01e54c676c9f28a
7.2 GB Preview Download
md5:57e6824dff2d5aad08b44b68b9722286
6.9 GB Preview Download
md5:7adc2aaa212d4c7e0fad9ddbae1058b4
3.4 GB Preview Download
md5:2049d3c5220083e53537a621fd1c7835
8.8 GB Preview Download
md5:66e84dde89d3b58a078c6306b40f36f5
3.4 GB Preview Download
md5:a89f54ff2348962fb4074eec0662b099
7.9 GB Preview Download