Extracting causal relations from the literature with word vector mapping

Ning An; Yongbo Xiao; Jing Yuan; Jiaoyun Yang; Gil Alterovitz

doi:10.1016/j.compbiomed.2019.103524

Extracting causal relations from the literature with word vector mapping

Comput Biol Med. 2019 Dec:115:103524. doi: 10.1016/j.compbiomed.2019.103524. Epub 2019 Nov 27.

Authors

Ning An¹, Yongbo Xiao², Jing Yuan³, Jiaoyun Yang⁴, Gil Alterovitz⁵

Affiliations

¹ Key Laboratory of Knowledge Engineering with Big Data of Ministry of Education, Hefei University of Technology, Hefei, China; School of Computer Science and Information Engineering, Hefei University of Technology, Hefei, China. Electronic address: ning.g.an@acm.org.
² Key Laboratory of Knowledge Engineering with Big Data of Ministry of Education, Hefei University of Technology, Hefei, China; School of Computer Science and Information Engineering, Hefei University of Technology, Hefei, China. Electronic address: xyb1996@mail.hfut.edu.cn.
³ Department of Neurology, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences, Beijing, China. Electronic address: yuanjing@pumch.cn.
⁴ Key Laboratory of Knowledge Engineering with Big Data of Ministry of Education, Hefei University of Technology, Hefei, China; School of Computer Science and Information Engineering, Hefei University of Technology, Hefei, China. Electronic address: jiaoyun@hfut.edu.cn.
⁵ Boston Children's Hospital, Harvard Medical School, Boston, USA. Electronic address: gil_alterovitz@hms.harvard.edu.

PMID: 31698234
DOI: 10.1016/j.compbiomed.2019.103524

Abstract

Causal graphs play an essential role in the determination of causalities and have been applied in many domains including biology and medicine. Traditional causal graph construction methods are usually data-driven and may not deliver the desired accuracy of a graph. Considering the vast number of publications with causality knowledge, extracting causal relations from the literature to help to establish causal graphs becomes possible. Current supervised-learning-based causality extraction methods requires sufficient labeled data to train a model, and rule-based causality extraction methods are limited by the predefined patterns. This paper proposes a causality extraction framework by integrating rule-based methods and unsupervised learning models to overcome these limitations. The proposed method consists of three modules, including data preprocessing, syntactic pattern matching, and causality determination. In data preprocessing, abstracts are crawled based on attribute names before sentences are extracted and simplified. In syntactic pattern matching, these simplified sentences are parsed to obtain the part-of-speech tags, and triples are achieved based on these tags by matching the two designed syntactic patterns. In causality determination, four verb seed sets are initialized, and word vectors are constructed for the verbs in both the seed sets and the triples by applying an unsupervised machine learning model. Causal relations are identified by comparing the similarity between the verbs in each triple and that in each seed set to overcome the limitation of the seed sets. Causality extraction results on the attributes from the risk factors for Alzheimer's disease show that our method outperforms Bui's method and Alashri's method in terms of precision, recall, specificity, accuracy and F-score, with increases in the F-score of 8.29% and 5.37%, respectively.

Keywords: Causal extraction; Causal graph; Causality; Literature analysis; Word vector.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Data Mining*
Humans
Semantics*
Support Vector Machine*