'Bingo'-a large language model- and graph neural network-based workflow for the prediction of essential genes from protein data

Brief Bioinform. 2023 Nov 22;25(1):bbad472. doi: 10.1093/bib/bbad472.

Abstract

The identification and characterization of essential genes are central to our understanding of the core biological functions in eukaryotic organisms, and has important implications for the treatment of diseases caused by, for example, cancers and pathogens. Given the major constraints in testing the functions of genes of many organisms in the laboratory, due to the absence of in vitro cultures and/or gene perturbation assays for most metazoan species, there has been a need to develop in silico tools for the accurate prediction or inference of essential genes to underpin systems biological investigations. Major advances in machine learning approaches provide unprecedented opportunities to overcome these limitations and accelerate the discovery of essential genes on a genome-wide scale. Here, we developed and evaluated a large language model- and graph neural network (LLM-GNN)-based approach, called 'Bingo', to predict essential protein-coding genes in the metazoan model organisms Caenorhabditis elegans and Drosophila melanogaster as well as in Mus musculus and Homo sapiens (a HepG2 cell line) by integrating LLM and GNNs with adversarial training. Bingo predicts essential genes under two 'zero-shot' scenarios with transfer learning, showing promise to compensate for a lack of high-quality genomic and proteomic data for non-model organisms. In addition, the attention mechanisms and GNNExplainer were employed to manifest the functional sites and structural domain with most contribution to essentiality. In conclusion, Bingo provides the prospect of being able to accurately infer the essential genes of little- or under-studied organisms of interest, and provides a biological explanation for gene essentiality.

Keywords: adversarial training; biological interpretation; essential gene prediction; graph neural network; large language model.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • Animals
  • Drosophila Proteins* / genetics
  • Drosophila melanogaster / genetics
  • Genes, Essential*
  • Mice
  • Microfilament Proteins / genetics
  • Neural Networks, Computer
  • Proteins / genetics
  • Proteomics
  • Workflow

Substances

  • Proteins
  • shot protein, Drosophila
  • Microfilament Proteins
  • Drosophila Proteins