Skip to main content
v2026.11,610 entries · CC-BY 4.0

NIST Mass Spectral Library Search: Match Factor, Reverse Match and the Retention-Index Cross-Check

How to interpret match factor, reverse match factor and probability in NIST mass spectral library search, why a high-scoring hit can still be the wrong compound for isomers and homologous series, and the retention-index cross-check that catches it.

Ask about NIST Mass Spectral Library Search: Match Factor, Reverse Match and the Retention-Index Cross-Check

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

A NIST library search does not identify a compound. It ranks candidates. The distinction matters because the three numbers NIST MS Search hands back for the top hit — match factor, reverse match factor, and probability — describe how well a spectrum resembles a library entry, not whether the unknown actually is that entry. A match factor in the 900s feels conclusive. On a chromatogram with a co-eluting shoulder, a homologous series, or a pair of positional isomers with near-identical electron-ionization fragmentation, it can also be confidently wrong. Reading the score correctly, and knowing what independent check catches the cases where it lies, is the difference between a defensible identification and a false one that survives because nobody looked past the top hit.

What the search algorithm is actually comparing

The NIST/EPA/NIH Mass Spectral Library, maintained and periodically updated by NIST’s Mass Spectrometry Data Center (recent major release lines have carried the NIST17, NIST20 and NIST23 designations), pairs library spectra with NIST MS Search software. For a standard electron-ionization (EI) full-scan spectrum, the search compares the unknown’s peak list — each fragment’s mass-to-charge ratio and relative intensity — against every library entry using a weighted dot-product (cosine similarity) algorithm. Each peak is weighted by both its intensity and its m/z, on the reasoning that a higher-mass fragment is more structurally diagnostic than a low-mass one and shouldn’t be swamped by an intense but generic low-mass ion. The output is scaled to a 0–999 (sometimes reported 0–1000) index rather than a raw similarity coefficient, which is why “match factor” reads like a percentage but isn’t one.

Three related numbers come out of that comparison, and they answer different questions:

Match factor (forward match): does everything in the unknown belong?

The forward match factor (MF) scores the unknown spectrum against the library spectrum in both directions — it penalizes library peaks missing from the unknown and unknown peaks not present in the library entry. That second penalty is the important one operationally: if the unknown spectrum carries extra ions from co-elution, column bleed, or background subtraction that wasn’t quite clean, the forward match factor drops even when every peak that is shared matches well. A depressed MF on an otherwise textbook compound is often a chromatography or background problem, not a wrong-library-entry problem — check peak purity and background subtraction before concluding the identification itself is bad.

Reverse match factor: does the library entry’s pattern show up in the unknown?

The reverse match factor (RMF) only checks the library peaks against the unknown — it ignores any unknown peaks that the library entry doesn’t have. That makes RMF the more forgiving, and more useful, metric for a spectrum you already suspect is contaminated or co-eluting: if RMF is high while MF is noticeably lower, the compound’s diagnostic fragments are genuinely present, but something extraneous is riding along with them. The gap between MF and RMF is itself informative — a wide gap is a purity flag, not noise to average away.

Probability: a statement about the candidate list, not about the top hit

Probability (Prob%), the third figure NIST MS Search reports, is not a calibrated “percent chance this identification is correct.” It’s derived from how the top several hits in the returned candidate list score relative to each other — specifically, how much better the top hit scores than the runners-up. A high probability means the search found one candidate that stands out clearly from the rest of the library; a low probability means several structurally different candidates scored similarly, which is common for compound classes with many close chemical relatives (positional isomers, homologous series, stereoisomers that EI-MS can’t distinguish at all, since EI spectra are generally insensitive to stereochemistry). A low probability is a legitimate reason to distrust an otherwise high match factor — it means the library itself can’t cleanly separate your compound from a near-identical relative, no matter how good the top score looks in isolation.

Why a high score can still be a false identification

Electron ionization fragments molecules in a way that’s often insensitive to exactly the structural differences that matter for identification. Positional isomers (ortho/meta/para-substituted aromatics, many terpene and steroid isomers), and compounds within a homologous series that differ by a repeating unit, frequently produce EI spectra similar enough that the library search alone cannot reliably tell them apart — the match factor can sit in the high 800s or 900s for more than one real candidate. The library search is doing its job correctly in that situation; the job it does is limited to spectral similarity, and spectral similarity does not uniquely determine molecular structure for every compound class. This is precisely the gap a retention-index check closes.

The retention-index cross-check

The linear (Kovats) retention index expresses a compound’s GC retention time relative to a bracketing pair of n-alkanes run on the same column under the same temperature program, producing a number that’s far more portable across instruments and labs than a raw retention time. Two facts make it the right independent check against a library match:

  • It measures a different physical property. Retention index reflects a compound’s interaction with the stationary phase (volatility, polarity, and shape) — nothing to do with how it fragments under electron ionization. A false identification driven by fragmentation similarity between two isomers has no reason to also produce a matching retention index by coincidence.
  • Reference values exist independently of your search result. The NIST database itself carries retention-index data (contributed for many library entries, most extensively for standard non-polar and semi-standard non-polar phases) alongside the spectra, and a large published literature reports RI values by column phase for common analyte classes. You compare your compound’s measured RI, calculated from your own n-alkane run on your own column, against the literature or database RI logged for the library hit — not against a value from a different phase type.

The workflow in practice: run (or maintain a current run of) an n-alkane calibration series on the same column and temperature program used for the sample, calculate the unknown’s linear retention index from its retention time relative to the bracketing alkanes, then compare that measured RI against the RI on record for the library’s top hit. A compound reporting a match factor above 900 whose measured RI is 40–50 index units or more away from the literature RI for that same phase type is a strong signal you have the right spectral neighborhood but the wrong specific compound — exactly the isomer/homolog failure mode above. RI agreement, by contrast, doesn’t retroactively prove the identification on its own, but it removes the single most common source of a false-positive high-match-factor result, and it’s the standard corroborating check recommended alongside library search in forensic and environmental GC-MS identification protocols specifically because it’s cheap to generate from data you’re already collecting.

One caveat worth stating plainly: retention index is phase-dependent. A literature RI measured on a different stationary-phase chemistry (a polar wax-type phase versus a non-polar dimethylpolysiloxane phase, for instance) is not directly comparable to your value even for the correct compound — match phase type to phase type, not just “an RI value” to “an RI value,” or the cross-check will produce false alarms of its own.

A practical identification workflow

  1. Run the library search and record MF, RMF, and Prob% for the top several hits, not just the single best one.
  2. If MF is markedly lower than RMF, suspect co-elution or background contamination before doubting the compound identity itself — check peak purity.
  3. If Prob% is low despite a high top-hit MF, treat the result as “this chemical family, specific member unconfirmed” rather than a settled identification — the library is telling you it found close relatives it can’t separate.
  4. Calculate the measured linear retention index from your own n-alkane series on the same column/method, and compare it against the literature or database RI for the top candidate on a matching phase type.
  5. Where the identification is load-bearing (regulatory, forensic, publication-grade), corroborate with an authentic reference standard run under identical conditions, or with an orthogonal technique (accurate mass, MS/MS fragmentation, or NMR) rather than resting the call on library search alone.

Frequently asked questions

What match factor counts as a confirmed identification?

There’s no single universal cutoff, and treating one number as a pass/fail line is exactly the mistake this guide is arguing against. A high match factor (commonly cited informal ranges put “good” matches above roughly 900 and “tentative” matches in the 800s) narrows the candidate list; it doesn’t close the case on its own, particularly for compound classes prone to isomer confusion. Corroborate with retention index and, where the identification matters, an authentic standard.

Why would reverse match factor be much higher than the forward match factor?

The unknown spectrum contains ions that the library entry doesn’t — typically co-eluting background, a column-bleed artifact, or incomplete background subtraction riding on top of a genuinely good match. RMF ignores those extra peaks; MF penalizes them. A wide MF/RMF gap is a data-quality flag on the chromatography, not evidence the identification itself is wrong.

Can two different compounds have essentially the same EI mass spectrum?

Yes, routinely. Electron ionization is a destructive, high-energy process that often produces very similar fragmentation for positional isomers, some stereoisomers (which EI-MS generally cannot distinguish at all), and adjacent members of a homologous series. This is the specific failure mode the retention-index cross-check exists to catch, since RI depends on physical interaction with the column rather than on fragmentation chemistry.

Does a high probability score mean the identification is statistically likely to be correct?

Not in the sense of a calibrated per-identification confidence. Probability reflects how distinctly the top hit separates from the next-best candidates in the search results — a measure of how ambiguous the candidate list is, not an estimate of ground truth. A high probability with a high match factor is reassuring; a low probability is a legitimate reason to look harder even when the top match factor alone looks fine.

Do I need to run my own retention index, or can I trust a database value?

You need your own measured value for the unknown — that’s the whole point of the cross-check, since it’s an independent measurement against which the database or literature RI for the candidate compound is compared. Run an n-alkane series on the same column and temperature program as the sample, calculate the linear retention index from the unknown’s position relative to the bracketing alkanes, and compare that measured number to the reference RI logged for your top library candidate on a matching phase type.

Related CASRAI guides

For the column, carrier gas and detector choices that determine what spectrum you get in the first place, see gas chromatography columns, carrier gases and detectors explained. For how GC-MS relates to the liquid-phase alternative, see GC-MS vs LC-MS and LC-MS explained. For the analogous database-search-parameter problem in proteomics identification, see protein identification by mass spectrometry.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.