Interpreting Neural Networks for Biological Sequences by Learning Stochastic Masks

Johannes Linder; Alyssa La Fleur; Zibo Chen; Ajasja Ljubeti; David Baker; Sreeram Kannan; Georg Seelig

doi:10.1038/s42256-021-00428-6

Interpreting Neural Networks for Biological Sequences by Learning Stochastic Masks

Nat Mach Intell. 2022 Jan;4(1):41-54. doi: 10.1038/s42256-021-00428-6. Epub 2022 Jan 25.

Authors

Johannes Linder¹, Alyssa La Fleur¹, Zibo Chen², Ajasja Ljubeti², David Baker², Sreeram Kannan³, Georg Seelig^{1

3}

Affiliations

¹ Paul G. Allen School of Computer Science and Engineering, University of Washington.
² Institute for Protein Design, University of Washington.
³ Department of Electrical and Computer Engineering, University of Washington.

Abstract

Sequence-based neural networks can learn to make accurate predictions from large biological datasets, but model interpretation remains challenging. Many existing feature attribution methods are optimized for continuous rather than discrete input patterns and assess individual feature importance in isolation, making them ill-suited for interpreting non-linear interactions in molecular sequences. Building on work in computer vision and natural language processing, we developed an approach based on deep learning - Scrambler networks - wherein the most salient sequence positions are identified with learned input masks. Scramblers learn to predict Position-Specific Scoring Matrices (PSSMs) where unimportant nucleotides or residues are scrambled by raising their entropy. We apply Scramblers to interpret the effects of genetic variants, uncover non-linear interactions between cis-regulatory elements, explain binding specificity for protein-protein interactions, and identify structural determinants of de novo designed proteins. We show that Scramblers enable efficient attribution across large datasets and result in high-quality explanations, often outperforming state-of-the-art methods.

Grants and funding

R21 HG010945/HG/NHGRI NIH HHS/United States