A Statistical Model for Identifying Proteins by Tandem Mass Spectrometry

Статистическая модель идентификации белков методом тандемной масс-спектрометрии
Ruedi Aebersold, Alexey I. Nesvizhskii, Eugene Kolker, Andrew Keller
2003-07-15

expectation-maximizationfalse positive identification ratepeptide-to-spectrum assignmentprotein identificationtandem mass spectrometry
A statistical model is presented for computing probabilities that proteins are present in a sample on the basis of peptides assigned to tandem mass (MS/MS) spectra acquired from a proteolytic digest of the sample. Peptides that correspond to more than a single protein in the sequence database are apportioned among all corresponding proteins, and a minimal protein list sufficient to account for the observed peptide assignments is derived using the expectation-maximization algorithm. Using peptide assignments to spectra generated from a sample of 18 purified proteins, as well as complex H. influenzae and Halobacterium samples, the model is shown to produce probabilities that are accurate and have high power to discriminate correct from incorrect protein identifications. This method allows filtering of large-scale proteomics data sets with predictable sensitivity and false positive identification error rates. Fast, consistent, and transparent, it provides a standard for publishing large-scale protein identification data sets in the literature and for comparing the results obtained from different experiments.
1
A statistical model computes probabilities that proteins are present based on peptides assigned to tandem MS/MS spectra from proteolytic digests.
2
Peptides matching multiple proteins are apportioned among corresponding proteins using an expectation-maximization (EM) algorithm to derive a minimal protein list explaining observed peptides.
3
The approach is fast, consistent, transparent, and suitable as a standard for publishing and comparing large-scale protein identification results.
4
The method enables filtering of large-scale proteomics datasets with predictable sensitivity and false positive identification error rates.
5
When applied to 18 purified proteins and complex H. influenzae and Halobacterium samples, the model produces accurate protein-presence probabilities with high power to discriminate correct from incorrect identifications.

Proteins present in a biological sample identified from tandem mass spectrometry (MS/MS) peptide assignments

Statistical computation of probabilities for protein presence and derivation of a minimal protein list from peptide-to-spectrum assignments (including apportionment of shared peptides) to control sensitivity and false positive identification rates

Publication Details
Publication Date
2003-07-15
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Ruedi Aebersold
Alexey I. Nesvizhskii
Eugene Kolker
Andrew Keller
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%