Clustering millions of tandem mass spectra

Ari M Frank; Nuno Bandeira; Zhouxin Shen; Stephen Tanner; Steven P Briggs; Richard D Smith; Pavel A Pevzner

doi:10.1021/pr070361e

Clustering millions of tandem mass spectra

J Proteome Res. 2008 Jan;7(1):113-22. doi: 10.1021/pr070361e. Epub 2007 Dec 8.

Authors

Ari M Frank¹, Nuno Bandeira, Zhouxin Shen, Stephen Tanner, Steven P Briggs, Richard D Smith, Pavel A Pevzner

Affiliation

¹ Department of Computer Science and Engineering, University of California, San Diego, La Jolla, California 92093-0404, USA. arf@cs.ucsd.edu

Abstract

Tandem mass spectrometry (MS/MS) experiments often generate redundant data sets containing multiple spectra of the same peptides. Clustering of MS/MS spectra takes advantage of this redundancy by identifying multiple spectra of the same peptide and replacing them with a single representative spectrum. Analyzing only representative spectra results in significant speed-up of MS/MS database searches. We present an efficient clustering approach for analyzing large MS/MS data sets (over 10 million spectra) with a capability to reduce the number of spectra submitted to further analysis by an order of magnitude. The MS/MS database search of clustered spectra results in fewer spurious hits to the database and increases number of peptide identifications as compared to regular nonclustered searches. Our open source software MS-Clustering is available for download at http://peptide.ucsd.edu or can be run online at http://proteomics.bioprojects.org/MassSpec.

Publication types

Research Support, N.I.H., Extramural
Research Support, Non-U.S. Gov't
Research Support, U.S. Gov't, Non-P.H.S.

MeSH terms

Amino Acid Sequence
Cluster Analysis*
Computational Biology
Molecular Sequence Data
Peptides / analysis*
Proteomics / methods*
Tandem Mass Spectrometry*

Substances

Peptides

Abstract

Publication types

MeSH terms

Substances

Grants and funding