Continual Learning for Automated Audio Captioning Using The Learning Without Forgetting Approach

Jan Berg; Konstantinos Drossos

doi:10.5281/zenodo.5723053

Published November 15, 2021 | Version v1

Conference paper Open

Continual Learning for Automated Audio Captioning Using The Learning Without Forgetting Approach

1. Audio Research Group, Tampere University, Finland

Automated audio captioning (AAC) is the task of automatically creating textual descriptions (i.e. captions) for the contents of a general audio signal. Most AAC methods are using existing datasets to optimize and/or evaluate upon. Given the limited information held by the AAC datasets, it is very likely that AAC methods learn only the information contained in the utilized datasets. In this paper we present a first approach for continuously adapting an AAC method to new information, using a continual learning method. In our scenario, a pre-optimized AAC method is used for some unseen general audio signals and can update its parameters in order to adapt to the new information, given a new reference caption. We evaluate our method using a freely available, pre-optimized AAC method and two freely available AAC datasets. We compare our proposed method with three scenarios, two of training on one of the datasets and evaluating on the other and a third of training on one dataset and fine-tuning on the other. Obtained results show that our method achieves a good balance between distilling new knowledge and not forgetting the previous one.

Notes

The authors wish to acknowledge CSC-IT Center for Science, Finland, for computational resources. K. Drossos has received funding from the European Union's Horizon 2020 research and innovation programme under grant agreement No 957337, project MARVEL.

Files

DCASE2021_Berg_et_al_continual_learning.pdf

Files (304.6 kB)

Name	Size	Download all
DCASE2021_Berg_et_al_continual_learning.pdf md5:459252e10945e22955f4c30c240023c2	304.6 kB	Preview Download

Additional details

Is published in: Conference paper: 10.5281/zenodo.5770113 (DOI)
Is supplemented by: Software: https://github.com/JanBerg1/AAC-LwF (URL); Dataset: 10.5281/zenodo.4783391 (DOI)

European Commission
MARVEL – Multimodal Extreme Scale Data Analytics for Smart Cities Environments 957337

	All versions	This version
Views	267	267
Downloads	143	143
Data volume	47.8 MB	47.8 MB

Continual Learning for Automated Audio Captioning Using The Learning Without Forgetting Approach

Notes

Files

DCASE2021_Berg_et_al_continual_learning.pdf

Files (304.6 kB)

Additional details

Related works

Funding

Continual Learning for Automated Audio Captioning Using The Learning Without Forgetting Approach

Creators

Description

Notes

Files

DCASE2021_Berg_et_al_continual_learning.pdf

Files (304.6 kB)

Additional details

Related works

Funding