Published September 2, 2018 | Version v1

Wide Learning for Auditory Comprehension

  • 1. University of Tübingen

Description

Classical linguistic, cognitive and engineering models for speech recognition and human auditory comprehension posit representations for sounds and words that mediate between the acoustic signal and interpretation. Recent advances in automatic speech recognition have shown, using deep learning, that state-of-the-art performance is obtained without such units. We present a cognitive model of auditory comprehension based on wide rather than deep learning that was trained on 20 to 80 hours of TV news broadcasts. Just as deep network models, our model is an end-to-end system that does not make use of phonemes and phonological wordform representations. Nevertheless, it performs well on the difficult task of single word identification (model accuracy 11.37%, Mozilla DeepSpeech: 4.45%). The architecture of the model is a simple two-layered wide neural network with weighted connections between the acoustic frequency band features as inputs and lexical outcomes (pointers to semantic vectors) as outputs. Model performance shows hardly any degredation when trained on speech in noise rather than on clean speech. Performance was further enhanced by adding a second network to a standard wide network. The present word recognition module is designed to become part of a larger system modeling the comprehension of running speech.

Files

ShafaeiBaayenInterspeech2018.pdf

Files (277.5 kB)

Name Size Download all
md5:6015ec29d95fd18f0922d51fbc4d3b05
277.5 kB Preview Download

Additional details

Funding

European Commission
WIDE - Wide Incremental learning with Discrimination nEtworks 742545