Discretization of gene expression data revised

Brief Bioinform. 2016 Sep;17(5):758-70. doi: 10.1093/bib/bbv074. Epub 2015 Sep 22.

Abstract

Gene expression measurements represent the most important source of biological data used to unveil the interaction and functionality of genes. In this regard, several data mining and machine learning algorithms have been proposed that require, in a number of cases, some kind of data discretization to perform the inference. Selection of an appropriate discretization process has a major impact on the design and outcome of the inference algorithms, as there are a number of relevant issues that need to be considered. This study presents a revision of the current state-of-the-art discretization techniques, together with the key subjects that need to be considered when designing or selecting a discretization approach for gene expression data.

Keywords: data mining; data preprocessing; discretization; gene expression analysis; gene expression data; machine learning.

MeSH terms

  • Algorithms
  • Data Mining
  • Gene Expression Profiling
  • Gene Expression*