DIRECTION: a machine learning framework for predicting and characterizing DNA methylation and hydroxymethylation in mammalian genomes

Bioinformatics. 2017 Oct 1;33(19):2986-2994. doi: 10.1093/bioinformatics/btx316.

Abstract

Motivation: 5-Methylcytosine and 5-Hydroxymethylcytosine in DNA are major epigenetic modifications known to significantly alter mammalian gene expression. High-throughput assays to detect these modifications are expensive, labor-intensive, unfeasible in some contexts and leave a portion of the genome unqueried. Hence, we devised a novel, supervised, integrative learning framework to perform whole-genome methylation and hydroxymethylation predictions in CpG dinucleotides. Our framework can also perform imputation of missing or low quality data in existing sequencing datasets. Additionally, we developed infrastructure to perform in silico, high-throughput hypotheses testing on such predicted methylation or hydroxymethylation maps.

Results: We test our approach on H1 human embryonic stem cells and H1-derived neural progenitor cells. Our predictive model is comparable in accuracy to other state-of-the-art DNA methylation prediction algorithms. We are the first to predict hydroxymethylation in silico with high whole-genome accuracy, paving the way for large-scale reconstruction of hydroxymethylation maps in mammalian model systems. We designed a novel, beam-search driven feature selection algorithm to identify the most discriminative predictor variables, and developed a platform for performing integrative analysis and reconstruction of the epigenome. Our toolkit DIRECTION provides predictions at single nucleotide resolution and identifies relevant features based on resource availability. This offers enhanced biological interpretability of results potentially leading to a better understanding of epigenetic gene regulation.

Availability and implementation: http://www.pradiptaray.com/direction, under CC-by-SA license.

Contacts: pradiptaray@gmail.com or mchen@utdallas.edu or michael.zhang@utdallas.edu.

Supplementary information: Supplementary data are available at Bioinformatics online.

MeSH terms

  • 5-Methylcytosine / analogs & derivatives*
  • 5-Methylcytosine / metabolism*
  • Algorithms
  • Animals
  • CpG Islands
  • DNA / chemistry
  • DNA / metabolism
  • DNA Methylation*
  • Epigenesis, Genetic
  • Humans
  • Machine Learning*
  • Mammals / genetics
  • Software

Substances

  • 5-hydroxymethylcytosine
  • 5-Methylcytosine
  • DNA