CMedTEX: A Rule-based Temporal Expression Extraction and Normalization System for Chinese Clinical Notes

AMIA Annu Symp Proc. 2017 Feb 10:2016:818-826. eCollection 2016.

Abstract

Time is an important aspect of information and is very useful for information utilization. The goal of this study was to analyze the challenges of temporal expression (TE) extraction and normalization in Chinese clinical notes by assessing the performance of a rule-based system developed by us on a manually annotated corpus (including 1,778 clinical notes of 281 hospitalized patients). In order to develop system conveniently, we divided TEs into three categories: direct, indirect and uncertain TEs, and designed different rules for each category of them. Evaluation on the independent test set shows that our system achieves an F-score of93.40% on TE extraction, and an accuracy of 92.58% on TE normalization under "exact-match" criterion. Compared with HeidelTime for Chinese newswire text, our system is much better, indicating that it is necessary to develop a specific TE extraction and normalization system for Chinese clinical notes because of domain difference.

MeSH terms

  • China
  • Electronic Health Records*
  • Humans
  • Information Storage and Retrieval
  • Natural Language Processing*
  • Time