Chinese unknown word recognition for PCFG-LA parsing

Qiuping Huang; Liangye He; Derek F Wong; Lidia S Chao

doi:10.1155/2014/959328

Chinese unknown word recognition for PCFG-LA parsing

ScientificWorldJournal. 2014:2014:959328. doi: 10.1155/2014/959328. Epub 2014 Apr 9.

Authors

Qiuping Huang¹, Liangye He¹, Derek F Wong¹, Lidia S Chao¹

Affiliation

¹ NLP CT Laboratory, Department of Computer and Information Science, University of Macau, Macau.

Abstract

This paper investigates the recognition of unknown words in Chinese parsing. Two methods are proposed to handle this problem. One is the modification of a character-based model. We model the emission probability of an unknown word using the first and last characters in the word. It aims to reduce the POS tag ambiguities of unknown words to improve the parsing performance. In addition, a novel method, using graph-based semisupervised learning (SSL), is proposed to improve the syntax parsing of unknown words. Its goal is to discover additional lexical knowledge from a large amount of unlabeled data to help the syntax parsing. The method is mainly to propagate lexical emission probabilities to unknown words by building the similarity graphs over the words of labeled and unlabeled data. The derived distributions are incorporated into the parsing process. The proposed methods are effective in dealing with the unknown words to improve the parsing. Empirical results for Penn Chinese Treebank and TCT Treebank revealed its effectiveness.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Asian People
Humans
Language
Recognition, Psychology
Vocabulary*