iDNA-ABF: multi-scale deep biological language learning model for the interpretable prediction of DNA methylations

Genome Biol. 2022 Oct 17;23(1):219. doi: 10.1186/s13059-022-02780-1.

Abstract

In this study, we propose iDNA-ABF, a multi-scale deep biological language learning model that enables the interpretable prediction of DNA methylations based on genomic sequences only. Benchmarking comparisons show that our iDNA-ABF outperforms state-of-the-art methods for different methylation predictions. Importantly, we show the power of deep language learning in capturing both sequential and functional semantics information from background genomes. Moreover, by integrating the interpretable analysis mechanism, we well explain what the model learns, helping us build the mapping from the discovery of important sequential determinants to the in-depth analysis of their biological functions.

Keywords: DNA methylation; Deep learning; Interpretable analysis; Multi-scale information processing.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • DNA Methylation*
  • Genomics
  • Language*
  • Models, Biological