Biological sequence modeling with convolutional kernel networks

Dexiong Chen; Laurent Jacob; Julien Mairal

doi:10.1093/bioinformatics/btz094

Biological sequence modeling with convolutional kernel networks

Bioinformatics. 2019 Sep 15;35(18):3294-3302. doi: 10.1093/bioinformatics/btz094.

Authors

Dexiong Chen¹, Laurent Jacob², Julien Mairal¹

Affiliations

¹ Université Grenoble Alpes, INRIA, CNRS, Grenoble INP, LJK, Grenoble, Isère France.
² University of Lyon, Université Lyon 1, CNRS, Laboratoire de Biométrie et Biologie Évolutive UMR 5558, Lyon, Rhône France.

PMID: 30753280
DOI: 10.1093/bioinformatics/btz094

Abstract

Motivation: The growing number of annotated biological sequences available makes it possible to learn genotype-phenotype relationships from data with increasingly high accuracy. When large quantities of labeled samples are available for training a model, convolutional neural networks can be used to predict the phenotype of unannotated sequences with good accuracy. Unfortunately, their performance with medium- or small-scale datasets is mitigated, which requires inventing new data-efficient approaches.

Results: We introduce a hybrid approach between convolutional neural networks and kernel methods to model biological sequences. Our method enjoys the ability of convolutional neural networks to learn data representations that are adapted to a specific task, while the kernel point of view yields algorithms that perform significantly better when the amount of training data is small. We illustrate these advantages for transcription factor binding prediction and protein homology detection, and we demonstrate that our model is also simple to interpret, which is crucial for discovering predictive motifs in sequences.

Availability and implementation: Source code is freely available at https://gitlab.inria.fr/dchen/CKN-seq.

Supplementary information: Supplementary data are available at Bioinformatics online.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Algorithms*
Neural Networks, Computer*
Protein Binding
Software