Multi-Task Joint Learning Model for Chinese Word Segmentation and Syndrome Differentiation in Traditional Chinese Medicine

Chenyuan Hu; Shuoyan Zhang; Tianyu Gu; Zhuangzhi Yan; Jiehui Jiang

doi:10.3390/ijerph19095601

Multi-Task Joint Learning Model for Chinese Word Segmentation and Syndrome Differentiation in Traditional Chinese Medicine

Int J Environ Res Public Health. 2022 May 5;19(9):5601. doi: 10.3390/ijerph19095601.

Authors

Chenyuan Hu¹, Shuoyan Zhang¹, Tianyu Gu¹, Zhuangzhi Yan², Jiehui Jiang²

Affiliations

¹ School of Communication and Information Engineering, Shanghai University, Shanghai 200444, China.
² Institute of Biomedical Engineering, School of Life Science, Shanghai University, Shanghai 200444, China.

Abstract

Evidence-based treatment is the basis of traditional Chinese medicine (TCM), and the accurate differentiation of syndromes is important for treatment in this context. The automatic differentiation of syndromes of unstructured medical records requires two important steps: Chinese word segmentation and text classification. Due to the ambiguity of the Chinese language and the peculiarities of syndrome differentiation, these tasks pose a daunting challenge. We use text classification to model syndrome differentiation for TCM, and use multi-task learning (MTL) and deep learning to accomplish the two challenging tasks of Chinese word segmentation and syndrome differentiation. Two classic deep neural networks—bidirectional long short-term memory (Bi-LSTM) and text-based convolutional neural networks (TextCNN)—are fused into MTL to simultaneously carry out these two tasks. We used our proposed method to conduct a large number of comparative experiments. The experimental comparisons showed that it was superior to other methods on both tasks. Our model yielded values of accuracy, specificity, and sensitivity of 0.93, 0.94, and 0.90, and 0.80, 0.82, and 0.78 on the Chinese word segmentation task and the syndrome differentiation task, respectively. Moreover, statistical analyses showed that the accuracies of the non-joint and joint models were both within the 95% confidence interval, with pvalue < 0.05. The experimental comparison showed that our method is superior to prevalent methods on both tasks. The work here can help modernize TCM through intelligent differentiation.

Keywords: deep learning; joint learning; multi-task learning; syndrome differentiation.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

China
Humans
Language*
Medicine, Chinese Traditional* / methods
Neural Networks, Computer
Syndrome

Grants and funding

This work was supported by the National Key Research and Development Program of China under Grant No 2018YFC1707704.