Using Acoustic Speech Patterns From Smartphones to Investigate Mood Disorders: Scoping Review

Olivia Flanagan; Amy Chan; Partha Roop; Frederick Sundram

doi:10.2196/24352

Using Acoustic Speech Patterns From Smartphones to Investigate Mood Disorders: Scoping Review

JMIR Mhealth Uhealth. 2021 Sep 17;9(9):e24352. doi: 10.2196/24352.

Authors

Olivia Flanagan¹, Amy Chan², Partha Roop³, Frederick Sundram¹

Affiliations

¹ Department of Psychological Medicine, Faculty of Medical and Health Sciences, University of Auckland, Auckland, New Zealand.
² School of Pharmacy, Faculty of Medical and Health Sciences, University of Auckland, Auckland, New Zealand.
³ Faculty of Engineering, University of Auckland, Auckland, New Zealand.

PMID: 34533465
PMCID: PMC8486998
DOI: 10.2196/24352

Abstract

Background: Mood disorders are commonly underrecognized and undertreated, as diagnosis is reliant on self-reporting and clinical assessments that are often not timely. Speech characteristics of those with mood disorders differs from healthy individuals. With the wide use of smartphones, and the emergence of machine learning approaches, smartphones can be used to monitor speech patterns to help the diagnosis and monitoring of mood disorders.

Objective: The aim of this review is to synthesize research on using speech patterns from smartphones to diagnose and monitor mood disorders.

Methods: Literature searches of major databases, Medline, PsycInfo, EMBASE, and CINAHL, initially identified 832 relevant articles using the search terms "mood disorders", "smartphone", "voice analysis", and their variants. Only 13 studies met inclusion criteria: use of a smartphone for capturing voice data, focus on diagnosing or monitoring a mood disorder(s), clinical populations recruited prospectively, and in the English language only. Articles were assessed by 2 reviewers, and data extracted included data type, classifiers used, methods of capture, and study results. Studies were analyzed using a narrative synthesis approach.

Results: Studies showed that voice data alone had reasonable accuracy in predicting mood states and mood fluctuations based on objectively monitored speech patterns. While a fusion of different sensor modalities revealed the highest accuracy (97.4%), nearly 80% of included studies were pilot trials or feasibility studies without control groups and had small sample sizes ranging from 1 to 73 participants. Studies were also carried out over short or varying timeframes and had significant heterogeneity of methods in terms of the types of audio data captured, environmental contexts, classifiers, and measures to control for privacy and ambient noise.

Conclusions: Approaches that allow smartphone-based monitoring of speech patterns in mood disorders are rapidly growing. The current body of evidence supports the value of speech patterns to monitor, classify, and predict mood states in real time. However, many challenges remain around the robustness, cost-effectiveness, and acceptability of such an approach and further work is required to build on current research and reduce heterogeneity of methodologies as well as clinical evaluation of the benefits and risks of such approaches.

Keywords: data science; diagnosis; monitoring; mood disorders; smartphone; speech patterns.

©Olivia Flanagan, Amy Chan, Partha Roop, Frederick Sundram. Originally published in JMIR mHealth and uHealth (https://mhealth.jmir.org), 17.09.2021.

Publication types

Research Support, Non-U.S. Gov't
Review

MeSH terms

Acoustics
Humans
Monitoring, Physiologic
Mood Disorders / diagnosis
Smartphone*
Speech*