Predictive modeling of Pseudomonas syringae virulence on bean using gradient boosted decision trees

PLoS Pathog. 2022 Jul 25;18(7):e1010716. doi: 10.1371/journal.ppat.1010716. eCollection 2022 Jul.

Abstract

Pseudomonas syringae is a genetically diverse bacterial species complex responsible for numerous agronomically important crop diseases. Individual P. syringae isolates are assigned pathovar designations based on their host of isolation and the associated disease symptoms, and these pathovar designations are often assumed to reflect host specificity although this assumption has rarely been rigorously tested. Here we developed a rapid seed infection assay to measure the virulence of 121 diverse P. syringae isolates on common bean (Phaseolus vulgaris). This collection includes P. syringae phylogroup 2 (PG2) bean isolates (pathovar syringae) that cause bacterial spot disease and P. syringae phylogroup 3 (PG3) bean isolates (pathovar phaseolicola) that cause the more serious halo blight disease. We found that bean isolates in general were significantly more virulent on bean than non-bean isolates and observed no significant virulence difference between the PG2 and PG3 bean isolates. However, when we compared virulence within PGs we found that PG3 bean isolates were significantly more virulent than PG3 non-bean isolates, while there was no significant difference in virulence between PG2 bean and non-bean isolates. These results indicate that PG3 strains have a higher level of host specificity than PG2 strains. We then used gradient boosting machine learning to predict each strain's virulence on bean based on whole genome k-mers, type III secreted effector k-mers, and the presence/absence of type III effectors and phytotoxins. Our model performed best using whole genome data and was able to predict virulence with high accuracy (mean absolute error = 0.05). Finally, we functionally validated the model by predicting virulence for 16 strains and found that 15 (94%) had virulence levels within the bounds of estimated predictions. This study strengthens the hypothesis that P. syringae PG2 strains have evolved a different lifestyle than other P. syringae strains as reflected in their lower level of host specificity. It also acts as a proof-of-principle to demonstrate the power of machine learning for predicting host specific adaptation.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • Decision Trees
  • Host Specificity
  • Phaseolus* / microbiology
  • Plant Diseases / microbiology
  • Pseudomonas syringae*
  • Virulence

Grants and funding

This study was funded by a grant from the Natural Sciences and Engineering Research Council of Canada (NSERC) to DSG. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.