Written and maintained by CASRAI Editorial Board
Last updated
A protein identification is a statistical claim, not an observation. Nothing about a tandem mass spectrum says “this is peptide X” the way a barcode scanner reads a barcode — a search engine compares the spectrum against every theoretical fragmentation pattern a sequence database can produce, scores the best match, and reports it alongside an estimate of how often a match that good happens by chance. Whether that identification actually stands or quietly belongs in the noise is decided by a short list of parameter choices made before the search ever runs: which enzyme setting, how many missed cleavages, which modifications, what mass tolerance, and what false discovery rate (FDR) threshold. Get those wrong and the identification list is not just less complete — it is wrong in a way that is invisible unless you know to look.
This guide walks the bottom-up (shotgun) workflow end to end, then spends most of its length on the database-search decisions that determine whether a peptide-spectrum match survives scrutiny. The wet-lab side — protease choice for digestion, cleanup, DDA vs DIA acquisition and run QC — is covered in depth in choosing DDA or DIA, sample prep, and run QC; the instrument coupling itself is in how liquid chromatography and mass spectrometry are coupled. This page picks up where those leave off: the spectra exist, and now something has to decide what they mean.
The bottom-up workflow, start to finish
Bottom-up (shotgun) proteomics identifies proteins indirectly, by identifying the peptides their digestion produces and inferring the parent proteins back from that peptide evidence. The stages, in order:
- Extraction and denaturation. Proteins are solubilized and unfolded so the digestion enzyme can reach internal cleavage sites, typically with a chaotropic agent (urea) or detergent, plus reduction and alkylation of cysteines to prevent disulfide reformation and non-specific side reactions.
- Enzymatic digestion. A protease cleaves proteins into peptides at defined sequence positions. This is a wet-lab step, but the enzyme and its cleavage rule have to be declared again later, identically, as a database-search parameter — the search engine has to know what kind of peptide to expect.
- Cleanup and separation. Digests are desalted, then peptides are separated by reversed-phase liquid chromatography so they enter the mass spectrometer sequentially rather than all at once.
- Tandem MS acquisition. Precursor ions are measured (MS1), selected or systematically windowed, then fragmented and their product ions measured (MS2). This is where DDA and DIA diverge, covered in the sample-prep guide linked above.
- Database search. Every acquired MS2 spectrum is compared against theoretical spectra generated by in silico digesting a protein sequence database under the parameters below, and scored for best match. This is the stage the rest of this page covers.
- Statistical validation (FDR control). Matches are filtered using a target-decoy estimate of the false discovery rate, not accepted on score alone.
- Protein inference. Peptide identifications that survive FDR filtering are assembled back into a protein-level parts list, which is where shared peptides and ambiguous protein groups become a real problem, not just a theoretical one.
Everything after step 4 is deterministic given steps 1–4’s outputs and the parameters chosen — which is exactly why those parameters are worth understanding rather than leaving on default.
Building the search space: the sequence database
The search engine can only identify a peptide whose sequence is actually present in the database it searches. A database that is too narrow (a single protein, when the sample is a lysate) misses real hits entirely; one that is unnecessarily broad (every organism’s proteome, when the sample is human) inflates the search space with sequences that can never legitimately match, which does not create false positives directly but does dilute the statistics used to control them. The practical rule is to search the smallest database that could plausibly contain the sample’s real content — typically a single well-annotated proteome such as UniProt’s reviewed Swiss-Prot set for the study organism, with a defined common-contaminant list (keratins, trypsin itself, serum albumin) appended so contaminant peptides get identified as contaminants instead of forced into a spurious match against a real protein.
A target-decoy search also requires a decoy database appended to (or generated alongside) the target: the same sequences reversed or shuffled, so they preserve amino acid composition and peptide length distribution without corresponding to any real protein. Decoy hits are not noise to be discarded — they are the entire mechanism the FDR estimate below depends on, and most modern search engines and search-engine wrappers (Comet, MS-GF+, MaxQuant/Andromeda, Mascot with a decoy option enabled, X!Tandem) generate or accept a decoy database as standard practice, not an optional extra step.
Enzyme specificity and missed cleavages
The digestion enzyme has to be declared to the search engine, and it has to match what was actually used at the bench — a mismatch here is a common, easily overlooked cause of a search that returns far fewer identifications than the data supports. Trypsin, cleaving C-terminal to lysine and arginine (and blocked by a following proline), is the default for a reason: it is efficient, reproducible, and produces peptides in a length and charge-state range electrospray and collision-induced fragmentation handle well.
Two settings determine how strictly the enzyme rule is enforced:
- Enzyme specificity. “Fully specific” (also called fully tryptic) requires both peptide termini to match a valid cleavage site; “semi-specific” requires only one terminus to match, allowing the other to fall anywhere. Semi-specific searches are necessary for applications with a genuinely non-tryptic terminus — N-terminomics, some post-translational-processing studies, samples with real in-vivo proteolysis — but they multiply the search space enormously, since almost any position can now be a valid terminus. Run semi-specific only when there is a real reason to expect non-tryptic termini; running it by default trades identification confidence for a marginal recall gain that mostly is not real.
- Missed cleavages. Digestion is never 100% efficient; some peptides retain one or more internal cleavage sites the enzyme did not cut. Allowing 1–2 missed cleavages in the search accounts for this and is standard practice; allowing more recovers a few additional real peptides but expands the theoretical peptide population faster than it adds true positives, again diluting the score distribution the FDR calculation relies on. A search set to zero missed cleavages will under-identify a normally-digested sample; a search set to 4–5 is rarely buying anything beyond noise unless the digestion itself was known to be incomplete.
Modifications: fixed and variable
Post-translational and chemical modifications change a peptide’s mass, and the search engine only considers a modified peptide if that modification was declared in advance — it does not discover unanticipated modifications on its own (open or unrestricted searches exist for that purpose and are a separate, much more computationally expensive workflow).
Modifications are declared as one of two kinds:
- Fixed modifications are applied to every instance of a residue, because the chemistry guarantees it. Carbamidomethylation of cysteine (from the iodoacetamide alkylation step in sample prep) is the standard example — essentially every cysteine in the sample carries it, so declaring it fixed rather than variable is both more accurate and computationally cheaper.
- Variable modifications may or may not be present on any given instance of a residue — oxidation of methionine (a common artifact of sample handling) is the standard example. The search engine has to consider both the modified and unmodified state at every eligible residue, which multiplies the number of theoretical peptide variants it checks against each spectrum.
This multiplication is the reason variable-modification lists should stay short and deliberate. Each additional variable modification roughly multiplies the effective search space, which increases the chance a decoy peptide scores well enough to pass the same threshold a real peptide would need — in effect quietly eroding the FDR control described below, even though the reported FDR number does not visibly change. A search declaring cysteine carbamidomethylation as fixed and methionine oxidation as the only variable modification is a defensible, tight default; a search with eight or ten variable modifications added “just in case” is a common way to generate a longer identification list that is not actually more trustworthy.
Precursor and fragment mass tolerance
Tolerance windows define how close an observed mass has to be to a theoretical mass to count as a match, and they should be set by what the instrument that acquired the data can actually measure, not left on a generic default carried over from a different instrument. A high-resolution accurate-mass analyser (Orbitrap, high-resolution time-of-flight) supports a narrow precursor tolerance, typically single-digit parts-per-million — see how much resolving power and mass accuracy actually buy you for the arithmetic behind that number. Fragment tolerance depends on which analyser records the MS2 scan: an Orbitrap or QTOF fragment scan again supports a tight, ppm-scale tolerance, while a fragment scan recorded on a lower-resolution ion trap needs a wider, Dalton-scale tolerance to match its actual measurement precision. Setting tolerances tighter than the instrument supports causes real matches to fall outside the window and be missed; setting them wider than necessary admits more candidate matches per spectrum and, like an over-long modification list, dilutes the score distribution the FDR estimate depends on.
Scoring: how a search engine picks a “best” match
Every candidate peptide within tolerance of a given precursor mass gets a theoretical fragmentation spectrum generated and compared against the observed spectrum; the comparison is reduced to a single score, and the highest-scoring candidate is reported as the identification for that spectrum. The scoring model differs by search engine — Mascot’s probability-based Mowse score, SEQUEST’s and Comet’s cross-correlation (Xcorr), X!Tandem’s hyperscore, MaxQuant/Andromeda’s probability-based score, MS-GF+’s spectral-probability score — but the underlying question is the same in every case: how much better does the top candidate explain this spectrum than chance would predict. None of these raw scores is comparable across search engines, and none of them is, by itself, evidence of a correct identification without the statistical validation step that follows — a high score on a bad database or an over-permissive parameter set is still a high score.
Target-decoy FDR: the statistical gate that decides what “identified” means
Score alone cannot tell you how many of the peptides above a given threshold are wrong, because there is no independent ground truth for a real sample. Target-decoy searching solves this empirically: since decoy sequences cannot correspond to any real peptide in the sample, any decoy match that scores above a threshold is, by construction, a false positive, and the rate of decoy matches above a threshold is used as an estimate of the rate of target false positives at that same threshold. The conventional cutoff across the field is 1% FDR, applied by finding the score threshold at which decoy hits make up no more than 1% of hits above it.
The detail that gets lost most often: FDR is calculated separately at different levels, and 1% at one level is not 1% at another.
- PSM-level FDR (peptide-spectrum match) is the rate among individual spectrum-to-peptide matches, including repeated identifications of the same peptide across multiple spectra.
- Peptide-level FDR collapses to unique peptide sequences before calculating the rate.
- Protein-level FDR is calculated after protein inference (below), and is characteristically looser than PSM-level FDR at the same nominal 1% setting — a protein can be “identified” at 1% protein FDR while individual supporting PSMs sit at a less favourable real rate, because the errors do not propagate one-to-one through inference.
A reported identification count is only meaningful when the FDR level it was filtered at is stated. Comparing a protein count from one analysis against a peptide or PSM count from another — or against a count from a different lab’s parameter set entirely — is comparing incompatible numbers dressed up as the same statistic.
Protein inference: from peptides back to proteins
A confidently identified peptide does not automatically produce a confidently identified protein, because many peptides — especially short ones, and anything from a conserved domain — match more than one protein sequence in the database, particularly across isoforms and close homologues. Standard practice applies the parsimony principle: report the smallest set of proteins that explains all observed peptides, rather than every protein any peptide could theoretically belong to. Where a peptide is shared between two candidate proteins and one of them also has unique peptide evidence, that peptide is assigned to the protein with independent support — the razor peptide convention — and proteins that cannot be distinguished on the available peptide evidence are reported together as a protein group rather than arbitrarily picking one.
A protein identified on a single peptide, seen in a single spectrum (a “one-hit wonder”), deserves more scrutiny than one supported by multiple peptides across multiple spectra, even when both clear the same FDR threshold — the FDR estimate is a population-level statistic about the whole result set, not a per-identification guarantee, and single-peptide identifications are disproportionately where a population-level 1% rate concentrates in practice.
When an identification does not actually stand
A peptide-spectrum match can pass the FDR filter and still be worth a second look before it goes into a paper or a report. Common warning signs:
- Single PSM support with no complementary evidence — no other peptides from the same protein, no replicate spectrum, nothing corroborating beyond the one match.
- A score close to the decoy-crossing threshold rather than well clear of it — the FDR filter accepts it, but it is exactly the kind of match the decoy rate is describing as a coin-flip proposition at the margin.
- A match that only works because of an unusual modification or missed-cleavage count that was allowed broadly rather than because there was specific reason to expect it on this peptide.
- A known contaminant sequence — keratin, trypsin autolysis products, common reagent proteins — matched instead against a real target because the contaminant was not included in the search database.
- Disagreement between search engines on the same spectrum, when more than one is run — not proof either is wrong, but a reasonable trigger for manual spectral inspection on anything the result depends on.
None of these turn a passed FDR filter into a rejected identification by themselves. They are the difference between an identification that is statistically licensed and one that has actually been looked at.
Frequently asked questions
What does a 1% FDR actually mean in a proteomics results list?
It means the score threshold used was chosen so that decoy (necessarily false) matches make up no more than 1% of everything above that threshold at the level the FDR was calculated — PSM, peptide, or protein. It is a statement about the whole filtered list’s expected error rate, not a certainty about any single identification in it.
How many missed cleavages should a search allow?
One to two is standard for a normal tryptic digest. Allowing more recovers a small number of additional real peptides at the cost of a much larger search space, which dilutes the statistics the FDR estimate depends on; only raise it if there is a specific, known reason the digestion was incomplete.
Should I run a semi-specific search by default?
No. Semi-specific (or non-specific) enzyme settings are appropriate when there is a real reason to expect non-tryptic peptide termini — N-terminomics, in-vivo proteolysis, some post-translational processing studies. Run fully specific by default; semi-specific searches enormously expand the search space for most standard tryptic-digest samples without a corresponding gain in real identifications.
Why does adding more variable modifications reduce confidence rather than add it?
Each additional variable modification multiplies the number of theoretical peptide forms the search engine checks against every spectrum. A larger search space makes it statistically easier for a decoy (false) peptide to score as well as a real one would, which erodes FDR control even though the reported FDR number does not visibly flag it.
Is a one-hit-wonder identification reliable?
It can pass the FDR filter and still deserve more scrutiny than a multi-peptide identification, because FDR is a population-level estimate across the whole result set, not a per-identification guarantee. Corroborating evidence — a second peptide, a replicate spectrum, manual inspection — matters most exactly for single-peptide calls.
Do different search engines give the same identification list from the same data?
Not exactly. Scoring models differ (Mascot’s Mowse score, SEQUEST/Comet’s Xcorr, X!Tandem’s hyperscore, MaxQuant/Andromeda’s and MS-GF+’s probability-based scores are not on the same scale), so the same spectra searched against the same database and parameters in two engines typically produce overlapping but non-identical identification lists. Running more than one engine on the same data is a reasonable cross-check, not a requirement.
Related CASRAI guides
For the acquisition and sample-prep decisions upstream of the search covered here, see choosing DDA or DIA, sample prep, and run QC. For the instrument fundamentals, see LC-MS explained and resolving power and mass accuracy. For where identified proteins and their supporting spectra get archived and reused, see UniProt, the Protein Data Bank, and ProteomeXchange and PXD accession numbers.








