Published September 11, 2025 | Version 1.0

ML-DF: A Proof-of-Concept EU AI Act Compliance Supporting Deepfake Speech Dataset

  • 1. ROR icon Brno University of Technology

Description

ML-DF - A Multilingual Deepfake Speech Dataset for Fairness and Compliance Analyses

ML-DF is a proof-of-concept dataset for evaluating deepfake (synthetic) speech detectors with a focus on fairness, auditability, and EU AI Act–aligned data governance. It contains 448,000 utterances (balanced 224,000 bona fide + 224,000 deepfake) designed to support subgroup analysis and compliance-oriented audits (e.g., by gender and language). Most samples are 2–6 seconds long. The dataset includes structured metadata (language, speaker info, synthesis tool) to enable fine-grained analyses. The dataset is balanced and targets realistic attack coverage by combining modern TTS and VC systems.

Contents

  • Total size: 448,000 utterances (224k genuine, 224k deepfake)
  • Languages: English, German, French, Spanish, Italian
  • Speakers: 106 total (53 male + 53 female)
  • Synthesizers
    • TTS: VITS [1], ZMM-TTS [2]
    • VC: LVC-VC [3], DDDM-VC [4]
  • Typical duration: 2–6 s per utterance
  • Balance: Balanced across real/fake, gender, and languages; convenient for statistical evaluation

All audio files are provided in 16-bit PCM WAV format at 16 kHz, consisting of synthetic (deepfake) and genuine speech samples.

Source & Metadata

Source speech is derived from Multilingual Librispeech (MLS) [5]. Each speaker contributed 3 hours of speech for model training, ensuring sufficient data for high-quality synthesis. For the development set, 70 minutes of speech per speaker were used. For the test set, another 70 minutes of speech per speaker were utilized. All subsets were speaker-disjoint to prevent data leakage and allow fair model evaluation - the data used for training the synthesizers is different from the ones used for synthesis, following the train-dev-test split of the MLS corpus. Therefore, ML-DF provides speaker-disjoint genuine vs. spoofed content and structured metadata (language, synthesis method details, reference speaker information). The language subset archives contain protocols with per-recording metadata and demographics.

Details about the synthesis procedure (including the training pipeline) are available separately in the file synthesis.pdf. All code for segmenting and aligning data, training synthesizers, generating deepfake recordings, as well as training and evaluating detectors, is available in the code.zip archived repository.

To prevent shortcut learning and inflating detector performance [6], we measured several auxiliary properties (artefacts) of recordings in ML-DF. We used the measured values as simple detector scores to compute Equal Error Rates (EER) to assess their discriminative power. We consider an artefact sufficiently controlled when its distribution between bona fide and deepfake samples largely overlaps, i.e., EER ≥ 30%. The results of this analysis are in the file artefacts.pdf.

License

CC BY 4.0 (commercial use permitted). The dataset builds upon MLS (also CC BY 4.0) and is released under the same license.

References

[1] J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in Proceedings of the 38th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 5530–5540. [Online]. Available: https://proceedings.mlr.press/v139/kim21f.html

[2] C. Gong, X. Wang, E. Cooper, D. Wells, L. Wang, J. Dang, K. Richmond, and J. Yamagishi, “Zmm-tts: Zero-shot multilingual and multispeaker speech synthesis conditioned on self-supervised discrete speech representations,” vol. 32, p. 4036–4051, Sep. 2024. [Online]. Available: https://doi.org/10.1109/TASLP.2024.3451951

[3] W. Kang, M. Hasegawa-Johnson, and D. Roy, “End-to-end zero-shot voice conversion with location-variable convolutions,” in Interspeech 2023, 2023, pp. 2303–2307

[4] H.-Y. Choi, S.-H. Lee, and S.-W. Lee, “Dddm-vc: Decoupled denoising diffusion models with disentangled representation and prior mixup for verified robust voice conversion,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, 2024, pp. 17 862–17 870.

[5] V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A large-scale multilingual dataset for speech research,” in Interspeech 2020, 2020, pp. 2757–2761

[6] N. Müller, F. Dieckmann, P. Czempin, R. Canals, K. Böttinger, and J. Williams, “Speech is silver, silence is golden: What do asvspoof-trained models really learn?” in 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, 2021, pp. 55–60.

Files

README.md

Files (43.9 GB)

Name Size
md5:d44a913ae3bf5e2666829325c19e5ffa
106.7 kB Preview Download
md5:8266381936b3b8324aba44d5ab2a2bf6
121.0 kB Preview Download
md5:b9a6867f592538cfbe78ee75831a7fa0
160.1 MB Preview Download
md5:f89e3db74171413c0b2b75f6c6a7ae40
10.5 GB Download
md5:c0e252f0d56d1d3b716fc5f970b6d7aa
10.8 GB Download
md5:d2271cd1a73721192712c94ec4a6e142
10.4 GB Download
md5:3f98271034b477088c6c3a4300b89a6e
10.7 GB Download
md5:c3ce93f9566605e0a5ad2e3cda099d7d
1.5 GB Download
md5:527dc6cad772ccb187d5bfe5af738204
18.7 kB Download
md5:25cc69e8d9234a22c1f38222e0bfdebf
2.1 MB Preview Download
md5:ce92531b0f374aad533e5e95ef98e243
130.8 kB Preview Download
md5:5938ef0e49cdcea3ae3146499d5d92b1
4.7 kB Preview Download
md5:fb51bc9abb966697b859ecd23e4c3791
505.8 kB Preview Download

Additional details

Funding

Brno University of Technology
Reliable, Secure, and Intelligent Computer Systems FIT-S-23-8151
e-INFRA
e-INFRA CZ LM2018140 ID:90254