Comprehensive assessment of protein loop modeling programs on large-scale datasets: prediction accuracy and efficiency

Tianyue Wang; Langcheng Wang; Xujun Zhang; Chao Shen; Odin Zhang; Jike Wang; Jialu Wu; Ruofan Jin; Donghao Zhou; Shicheng Chen; Liwei Liu; Xiaorui Wang; Chang-Yu Hsieh; Guangyong Chen; Peichen Pan; Yu Kang; Tingjun Hou

doi:10.1093/bib/bbad486

Comprehensive assessment of protein loop modeling programs on large-scale datasets: prediction accuracy and efficiency

Brief Bioinform. 2023 Nov 22;25(1):bbad486. doi: 10.1093/bib/bbad486.

Authors

Tianyue Wang¹, Langcheng Wang², Xujun Zhang¹, Chao Shen¹, Odin Zhang¹, Jike Wang¹, Jialu Wu¹, Ruofan Jin³, Donghao Zhou⁴, Shicheng Chen¹, Liwei Liu⁵, Xiaorui Wang⁶, Chang-Yu Hsieh¹, Guangyong Chen⁷, Peichen Pan¹, Yu Kang¹, Tingjun Hou¹

Affiliations

¹ College of Pharmaceutical Sciences, Zhejiang University, Hangzhou 310058, Zhejiang, China.
² Department of Pathology, New York University Medical Center, 550 First Avenue, New York, NY 10016, USA.
³ College of Life Sciences, Zhejiang University, Hangzhou 310058, Zhejiang, China.
⁴ Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, Guangdong, China.
⁵ Advanced Computing and Storage Laboratory, Central Research Institute, 2012 Laboratories, Huawei Technologies Co., Ltd., Shenzhen 518129, Guangdong, China.
⁶ State Key Laboratory of Quality Research in Chinese Medicines, Macau University of Science and Technology, Macao, China.
⁷ Zhejiang Lab, Zhejiang University, Hangzhou 311121, Zhejiang, China.

Abstract

Protein loops play a critical role in the dynamics of proteins and are essential for numerous biological functions, and various computational approaches to loop modeling have been proposed over the past decades. However, a comprehensive understanding of the strengths and weaknesses of each method is lacking. In this work, we constructed two high-quality datasets (i.e. the General dataset and the CASP dataset) and systematically evaluated the accuracy and efficiency of 13 commonly used loop modeling approaches from the perspective of loop lengths, protein classes and residue types. The results indicate that the knowledge-based method FREAD generally outperforms the other tested programs in most cases, but encountered challenges when predicting loops longer than 15 and 30 residues on the CASP and General datasets, respectively. The ab initio method Rosetta NGK demonstrated exceptional modeling accuracy for short loops with four to eight residues and achieved the highest success rate on the CASP dataset. The well-known AlphaFold2 and RoseTTAFold require more resources for better performance, but they exhibit promise for predicting loops longer than 16 and 30 residues in the CASP and General datasets. These observations can provide valuable insights for selecting suitable methods for specific loop modeling tasks and contribute to future advancements in the field.

Keywords: AlphaFold2; artificial intelligence; deep learning; loop modeling; protein loop.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Protein Conformation
Proteins* / chemistry

Substances

Proteins

Abstract

Publication types

MeSH terms

Substances

Grants and funding