Published August 23, 2023 | Version v1

Abusive Speech Detection in Indic Languages Using Acoustic Features

  • 1. EDMO icon Technical University of Munich
  • 1. EDMO icon Technical University of Munich
  • 2. ROR icon Universitas Islam Negeri Antasari Banjarmasin
  • 3. Imperial College London

Description

Abusive content in online social networks is a well-known problem that can cause serious psychological harm and incite hatred. The ability to upload audio data increases the importance of developing methods to detect abusive content in speech recordings. However, simply transferring the mechanisms from written abuse detection would ignore relevant information such as emotion and tone. In addition, many current algorithms require training in the specific language for which they are being used. This paper proposes to use acoustic and prosodic features to classify abusive content. We used the ADIMA data set, which contains recordings from ten Indic languages, and trained different models in multilingual and cross-lingual settings. Our results show that it is possible to classify abusive and non-abusive content using only acoustic and prosodic features. The most important and influential features are discussed.

Files

spiesberger23_interspeech.pdf

Files (2.4 MB)

Name Size Download all
md5:472bb12964c94db394dd20cc46500ea2
2.4 MB Preview Download

Additional details

Funding

European Commission
SHIFT - MetamorphoSis of cultural Heritage Into augmented hypermedia assets For enhanced accessibiliTy and inclusion 101060660

Dates

Available
2023-08-23