A novel feature selection strategy for enhanced biomedical event extraction using the Turku system

Biomed Res Int. 2014:2014:205239. doi: 10.1155/2014/205239. Epub 2014 Apr 6.

Abstract

Feature selection is of paramount importance for text-mining classifiers with high-dimensional features. The Turku Event Extraction System (TEES) is the best performing tool in the GENIA BioNLP 2009/2011 shared tasks, which relies heavily on high-dimensional features. This paper describes research which, based on an implementation of an accumulated effect evaluation (AEE) algorithm applying the greedy search strategy, analyses the contribution of every single feature class in TEES with a view to identify important features and modify the feature set accordingly. With an updated feature set, a new system is acquired with enhanced performance which achieves an increased F-score of 53.27% up from 51.21% for Task 1 under strict evaluation criteria and 57.24% according to the approximate span and recursive criterion.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • Abstracting and Indexing / methods*
  • Algorithms*
  • Artificial Intelligence*
  • Data Mining / methods*
  • Natural Language Processing*
  • Pattern Recognition, Automated / methods*
  • Semantics*
  • Vocabulary, Controlled*