Fairness and underspecification in acoustic scene classification: The case for disaggregated evaluations

Andreas Triantafyllopoulos; Manuel Milling; Konstantinos Drossos; Björn W. Schuller

doi:10.5281/zenodo.5552489

Published October 6, 2021 | Version v1

Preprint Open

Fairness and underspecification in acoustic scene classification: The case for disaggregated evaluations

1. audEERING GmbH, Gilching, Germany & EIHW – Chair of Embedded Intelligence for Health Care and Wellbeing, University of Augsburg, Germany
2. EIHW – Chair of Embedded Intelligence for Health Care and Wellbeing, University of Augsburg, Germany
3. Audio Research Group, Tampere University, Tampere, Finland
4. audEERING GmbH, Gilching, Germany & EIHW – Chair of Embedded Intelligence for Health Care and Wellbeing, University of Augsburg, Germany & GLAM – Group for Audio, Language, & Music, Imperial College, London, UK

Underspecification and fairness in machine learning (ML) applications have recently become two prominent issues in the ML community. Acoustic scene classification (ASC) applications have so far remained unaffected by this discussion, but are now becoming increasingly used in real-world systems where fairness and reliability are critical aspects. In this work, we argue for the need of a more holistic evaluation process for ASC models through disaggregated evaluations. This entails taking into account performance differences across several factors, such as city, location, and recording device. Although these factors play a well-understood role in the performance of ASC models, most works report single evaluation metrics taking into account all different strata of a particular dataset. We argue that metrics computed on specific sub-populations of the underlying data contain valuable information about the expected real-world behaviour of proposed systems, and their reporting could improve the transparency and trustability of such systems. We demonstrate the effectiveness of the proposed evaluation process in uncovering underspecification and fairness problems exhibited by several standard ML architectures when trained on two widely-used ASC datasets. Our evaluation shows that all examined architectures exhibit large biases across all factors taken into consideration, and in particular with respect to the recording location. Additionally, different architectures exhibit different biases even though they are trained with the same experimental configurations.

Files

DCASE2021___Holistic_Evaluation.pdf

Files (166.9 kB)

Name	Size	Download all
DCASE2021___Holistic_Evaluation.pdf md5:b9e9e81f466c9163a1f986428d2f3a8a	166.9 kB	Preview Download

Additional details

Is published in: Conference paper: 10.5281/zenodo.5770113 (DOI)
Is supplemented by: Dataset: 10.5281/zenodo.1228142 (DOI); Dataset: 10.5281/zenodo.1228235 (DOI)

European Commission
MARVEL – Multimodal Extreme Scale Data Analytics for Smart Cities Environments 957337

	All versions	This version
Views	257	255
Downloads	136	135
Data volume	23.7 MB	23.5 MB

Fairness and underspecification in acoustic scene classification: The case for disaggregated evaluations

Files

DCASE2021___Holistic_Evaluation.pdf

Files (166.9 kB)

Additional details

Related works

Funding

Fairness and underspecification in acoustic scene classification: The case for disaggregated evaluations

Creators

Description

Files

DCASE2021___Holistic_Evaluation.pdf

Files (166.9 kB)

Additional details

Related works

Funding