Biosignal Sensors and Deep Learning-Based Speech Recognition: A Review

Wookey Lee; Jessica Jiwon Seong; Busra Ozlu; Bong Sup Shim; Azizbek Marakhimov; Suan Lee

doi:10.3390/s21041399

Biosignal Sensors and Deep Learning-Based Speech Recognition: A Review

Sensors (Basel). 2021 Feb 17;21(4):1399. doi: 10.3390/s21041399.

Authors

Wookey Lee¹, Jessica Jiwon Seong², Busra Ozlu³, Bong Sup Shim³, Azizbek Marakhimov⁴, Suan Lee⁵

Affiliations

¹ Biomedical Science and Engineering & Dept. of Industrial Security Governance & IE, Inha University, 100 Inharo, Incheon 22212, Korea.
² Department of Industrial Security Governance, Inha University, 100 Inharo, Incheon 22212, Korea.
³ Biomedical Science and Engineering & Department of Chemical Engineering, Inha University, 100 Inharo, Incheon 22212, Korea.
⁴ Frontier College, Inha University, 100 Inharo, Incheon 22212, Korea.
⁵ School of Computer Science, Semyung University, Jecheon 27136, Korea.

Abstract

Voice is one of the essential mechanisms for communicating and expressing one's intentions as a human being. There are several causes of voice inability, including disease, accident, vocal abuse, medical surgery, ageing, and environmental pollution, and the risk of voice loss continues to increase. Novel approaches should have been developed for speech recognition and production because that would seriously undermine the quality of life and sometimes leads to isolation from society. In this review, we survey mouth interface technologies which are mouth-mounted devices for speech recognition, production, and volitional control, and the corresponding research to develop artificial mouth technologies based on various sensors, including electromyography (EMG), electroencephalography (EEG), electropalatography (EPG), electromagnetic articulography (EMA), permanent magnet articulography (PMA), gyros, images and 3-axial magnetic sensors, especially with deep learning techniques. We especially research various deep learning technologies related to voice recognition, including visual speech recognition, silent speech interface, and analyze its flow, and systematize them into a taxonomy. Finally, we discuss methods to solve the communication problems of people with disabilities in speaking and future research with respect to deep learning components.

Keywords: EMG; artificial larynx; biosignal; deep learning; mouth interface; voice production.

Publication types

Review

MeSH terms

Deep Learning*
Humans
Speech
Speech Perception*
Voice*

Abstract

Publication types

MeSH terms

Grants and funding