Shennong: A Python toolbox for audio speech features extraction

Mathieu Bernard; Maxime Poli; Julien Karadayi; Emmanuel Dupoux

doi:10.3758/s13428-022-02029-6

Shennong: A Python toolbox for audio speech features extraction

Behav Res Methods. 2023 Dec;55(8):4489-4501. doi: 10.3758/s13428-022-02029-6. Epub 2023 Feb 7.

Authors

Mathieu Bernard^#^{1

2}, Maxime Poli^#³, Julien Karadayi³, Emmanuel Dupoux^{3

4}

Affiliations

¹ Cognitive Machine Learning, PSL Research University, CNRS, EHESS, ENS, Inria, Paris, France. mathieu.bernard.2@cnrs.fr.
² EconomiX (UMR 7235), Université Paris Nanterre, CNRS, Nanterre, France. mathieu.bernard.2@cnrs.fr.
³ Cognitive Machine Learning, PSL Research University, CNRS, EHESS, ENS, Inria, Paris, France.
⁴ Meta AI Research, Paris, France.

^# Contributed equally.

PMID: 36750521
DOI: 10.3758/s13428-022-02029-6

Abstract

We introduce Shennong, a Python toolbox and command-line utility for audio speech features extraction. It implements a wide range of well-established state-of-the-art algorithms: spectro-temporal filters such as Mel-Frequency Cepstral Filterbank or Predictive Linear Filters, pre-trained neural networks, pitch estimators, speaker normalization methods, and post-processing algorithms. Shennong is an open source, reliable and extensible framework built on top of the popular Kaldi speech processing library. The Python implementation makes it easy to use by non-technical users and integrates with third-party speech modeling and machine learning tools from the Python ecosystem. This paper describes the Shennong software architecture, its core components, and implemented algorithms. Then, three applications illustrate its use. We first present a benchmark of speech features extraction algorithms available in Shennong on a phone discrimination task. We then analyze the performances of a speaker normalization model as a function of the speech duration used for training. We finally compare pitch estimation algorithms on speech under various noise conditions.

Keywords: Features extraction; Pitch estimation; Python; Software; Speech processing.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Algorithms
Ecosystem*
Humans
Neural Networks, Computer
Software
Speech*

Abstract

Publication types

MeSH terms

Grants and funding