Learning spectro-temporal representations of complex sounds with parameterized neural networks

Rachid Riad; Julien Karadayi; Anne-Catherine Bachoud-Lévi; Emmanuel Dupoux

doi:10.1121/10.0005482

Learning spectro-temporal representations of complex sounds with parameterized neural networks

J Acoust Soc Am. 2021 Jul;150(1):353. doi: 10.1121/10.0005482.

Authors

Rachid Riad¹, Julien Karadayi¹, Anne-Catherine Bachoud-Lévi², Emmanuel Dupoux¹

Affiliations

¹ Ecole des Hautes Etudes en Sciences Sociales, CNRS, Institut National de Recherche informatique et Automatique, Département d'Études Cognitives, Ecole Normale Supérieure-Paris Sciences et Lettres University, 29 Rue d'Ulm, 75005 Paris, France.
² NeuroPsychologie Interventionnelle, Département d'Études Cognitives, Ecole Normale Supérieure, Institut National de la Santé et de la Recherche Médicale, Institut Mondor de Recherche Biomédicale, Neuratris, Université Paris-Est Créteil, Paris Sciences et Lettres University, 29 Rue d'Ulm, 75005 Paris, France.

PMID: 34340514
DOI: 10.1121/10.0005482

Abstract

Deep learning models have become potential candidates for auditory neuroscience research, thanks to their recent successes in a variety of auditory tasks, yet these models often lack interpretability to fully understand the exact computations that have been performed. Here, we proposed a parametrized neural network layer, which computes specific spectro-temporal modulations based on Gabor filters [learnable spectro-temporal filters (STRFs)] and is fully interpretable. We evaluated this layer on speech activity detection, speaker verification, urban sound classification, and zebra finch call type classification. We found that models based on learnable STRFs are on par for all tasks with state-of-the-art and obtain the best performance for speech activity detection. As this layer remains a Gabor filter, it is fully interpretable. Thus, we used quantitative measures to describe distribution of the learned spectro-temporal modulations. Filters adapted to each task and focused mostly on low temporal and spectral modulations. The analyses show that the filters learned on human speech have similar spectro-temporal parameters as the ones measured directly in the human auditory cortex. Finally, we observed that the tasks organized in a meaningful way: the human vocalization tasks closer to each other and bird vocalizations far away from human vocalizations and urban sounds tasks.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Acoustic Stimulation
Auditory Cortex*
Auditory Perception
Neural Networks, Computer
Speech Perception*