Written and maintained by CASRAI Editorial Board
Last updated
Most guidance on MeSH stops at “MeSH is PubMed’s controlled vocabulary, so use MeSH terms and your search will be more precise.” That is true and almost useless. It tells you nothing about the three mechanics that actually change what a MeSH search returns — explosion, major-topic restriction, and subheadings — and it says nothing at all about the failure mode that matters most to anyone running a review: a MeSH-only strategy silently under-retrieves the most recent literature, because the newest citations in PubMed have not been indexed with MeSH yet, and a large minority of them never will be.
None of those failures announce themselves. A MeSH search that misses half the current year’s relevant papers returns a clean, plausible, well-ranked result set. This page sets out what NLM actually documents about each mechanic, and pairs it with counts measured directly against PubMed so you can see the size of each effect rather than take it on faith.
How to read this page
Three tiers, kept separate throughout, in the same way as CASRAI’s companion page on which Google Scholar operators actually work:
- Documented — stated by NLM in the PubMed User Guide, which is the authoritative specification for PubMed’s query syntax.
- Measured — a retrieval count obtained by running the query against PubMed’s E-utilities
esearchAPI on 26 August 2026. PubMed grows daily and MEDLINE indexing is continuous, so every count here is a snapshot. The ratios are the durable finding; the absolute numbers will have moved by the time you read this, and re-running them is a two-minute check. - Observed only — reproducible behaviour that NLM has not specified in writing. Flagged as such rather than presented as a rule.
What a MeSH term is, and where it comes from
MeSH (Medical Subject Headings) is NLM’s controlled vocabulary for describing the subject of every article indexed for MEDLINE. Per the PubMed User Guide, MeSH is updated annually, and its terms are “arranged hierarchically by subject categories with more specific terms arranged beneath broader terms.” Applying it “ensures that articles are uniformly indexed by subject, whatever the author’s words” — which is the whole point: an indexer assigns Myocardial Infarction regardless of whether the author wrote “heart attack”, “MI”, or “coronary thrombosis”.
Two consequences follow immediately, and they pull in opposite directions:
- MeSH gives you recall that free text cannot. One heading collects every synonym an author might have used, plus every non-English-language record whose title and abstract you would never match on.
- MeSH only exists on records a human or an algorithm has already indexed. Free text is present the moment the publisher deposits the citation. MeSH is not. That asymmetry is the whole of §6 below.
The indexing itself is no longer a purely manual process. The User Guide states plainly that “MEDLINE articles are automatically indexed with MeSH terms using a well-refined algorithm,” and PubMed exposes the method as a searchable value. Measured 26 August 2026: indexingmethod_automated returns 6,888,662 citations and indexingmethod_curated (algorithmic suggestions reviewed, and possibly modified, by a human) returns 1,794,516. If you are writing a methods section that characterises MeSH as human-assigned indexing, that description is now out of date for a large and growing share of the database.
Finding the right term: the MeSH Database
Do not guess headings. The MeSH Database is the lookup tool, and the User Guide notes that MeSH terms “can be selected for searching in the MeSH database and from the advanced search builder index.” For each descriptor it gives you the four things you need before you commit it to a strategy:
- The preferred descriptor — the string PubMed actually indexes on.
- The entry terms — the synonyms that map to it. This is how you discover that Odontalgia is an entry term for Toothache, and not a heading in its own right.
- The tree position, which tells you what explosion will pull in.
- The allowable subheadings for that descriptor, and the year the descriptor was introduced — worth checking, because a heading cannot have been applied to records indexed before it existed.
The field tags, in one table
All of these are documented in the PubMed User Guide’s search field tag list. Case and spacing inside the brackets do not matter: crabs [mh] and Crabs[mh] are the same query.
| Tag | Example | What it does |
|---|---|---|
[mh] (or [mesh]) |
hypertension[mh] |
Searches the term as a MeSH heading. Explodes by default — includes every narrower term beneath it in the tree. |
[mh:noexp] |
neoplasms[mh:noexp] |
Turns explosion off. Retrieves only records indexed to that exact heading. |
[majr] |
sepsis[majr] |
MeSH Major Topic — restricts to records where the heading is one of the main topics of the article, marked with an asterisk on the indexed record. |
[majr:noexp] |
hypertension[majr:noexp] |
Major topic, explosion off. Both restrictions at once. |
[sh] |
toxicity[sh] |
Searches a subheading on its own (“free-floating”), unattached to any heading. Also explodes by default. |
[sh:noexp] |
therapy[sh:noexp] |
Free-floating subheading, explosion off. |
Heading/Subheading |
neoplasms/diet therapy |
Directly attached subheading. The [mh] tag is optional here; [majr] may be used instead. |
[mhda] |
2026/03[mhda] |
MeSH Date — the date MeSH terms were added to the citation. Not searched by All Fields; the tag is required. |
[sb] |
medline[sb] |
Citation status subset. The diagnostic tool for §6. |
Mechanic 1: explosion
Documented. “MeSH terms in PubMed automatically include the more specific MeSH terms in a search. To turn off this automatic feature, use the search syntax [mh:noexp].”
This is the single most consequential default in PubMed’s query language, and it is on unless you switch it off. It is also the opposite of the default on the Ovid platform, where explosion is opt-in via exp — a difference that matters if you are moving a strategy between platforms, as CASRAI’s guide to translating a PubMed strategy into Embase sets out in more detail.
Measured (26 August 2026):
| Query | Records | Effect of explosion |
|---|---|---|
"sepsis"[mh] |
158,361 | Explosion nearly doubles retrieval: the narrower headings beneath Sepsis contribute 76,430 records, 48% of the total. |
"sepsis"[mh:noexp] |
81,931 | |
"hypertension"[mh] |
340,893 | A flatter branch: explosion adds 65,207 records, 19% of the total. |
"hypertension"[mh:noexp] |
275,686 |
The practical reading: the size of the explosion effect is a property of the branch, not a constant. You cannot estimate it. Run both forms and look at the difference before deciding whether :noexp is buying you precision worth the recall it costs. In an evidence synthesis, that decision has to be justified in the protocol, not made silently at the keyboard.
Mechanic 2: major-topic restriction
Documented. MeSH Major Topic [majr] is “a MeSH term that is one of the main topics discussed in the article denoted by an asterisk on the MeSH term or MeSH/Subheading combination, e.g., Cytokines/physiology*”.
Every record carries a mix of headings: the ones the article is about, and the ones that merely describe its context — the population, the setting, the incidental comparator. [majr] keeps only the first kind. It is the cleanest precision lever PubMed has, and it is also the one most likely to quietly lose you eligible studies, because whether a heading was starred is an indexing judgment made by an algorithm or an indexer, not by you and not by the study’s authors.
Measured (26 August 2026): "sepsis"[majr] returns 116,066 against 158,361 for "sepsis"[mh] — the restriction discards 42,295 records, 27% of the set. [majr] explodes by default too, exactly as [mh] does.
When to use it: scoping work, a rapid orientation to a literature, or building a precise search filter you will validate. When not to: the eligibility-defining concept of a systematic review. An intervention that appears in a trial’s methods but is not the paper’s headline topic will not be starred, and [majr] will drop that trial without telling you.
Mechanic 3: subheadings (qualifiers)
Subheadings narrow a heading to a particular aspect — drug therapy, adverse effects, diagnosis, epidemiology. There are three distinct ways to use them, and they do not behave the same way.
Attached directly to a heading
Documented. “To directly attach MeSH Subheadings, use the format MeSH Term/Subheading, e.g., neoplasms/diet therapy. You may also use the two-letter MeSH Subheading abbreviations, e.g., neoplasms/dh. The [mh] tag is not required, however [majr] may be used, e.g., plants/genetics[majr]. Only one Subheading may be directly attached to a MeSH term.“
That last sentence is the one people get wrong. hypertension/drug therapy/adverse effects is not a valid construction. If you need two aspects, you write two attached-subheading terms and combine them with OR.
Measured (26 August 2026), confirming the equivalences NLM documents: neoplasms/dh[mh], neoplasms/diet therapy[mh] and untagged neoplasms/diet therapy all return 3,199. plants/genetics[majr] returns 21,910.
Free-floating
Documented. “The MeSH Subheading field allows users to ‘free float’ Subheadings, e.g., hypertension [mh] AND toxicity [sh].” Subheadings have their own hierarchy and also explode by default; [sh:noexp] turns that off, and the two-letter abbreviations work here too.
Measured: dh[sh] and diet therapy[sh] both return 61,130. hypertension[mh] AND toxicity[sh] returns 2,661.
The difference between attached and free-floating is not cosmetic. hypertension/toxicity requires the qualifier to have been attached to that heading on the record. hypertension[mh] AND toxicity[sh] only requires both to appear somewhere on the record — the toxicity may have been indexed against an entirely different heading. Free-floating is broader and less exact; use it deliberately, not as a shortcut for syntax you were unsure of.
The double explosion — the trap in this section
Documented, and routinely missed. “For a MeSH/Subheading combination, PubMed always includes the more specific terms arranged beneath broader terms for the MeSH term and also includes the more specific terms arranged beneath broader Subheadings.” NLM’s own worked example: hypertension/therapy also retrieves hypertension/diet therapy, hypertension/drug therapy, hypertension, malignant/therapy, hypertension, malignant/drug therapy, and so on.
So an attached-subheading term explodes along two hierarchies at once. [mh:noexp] switches off both: NLM states that hypertension/therapy [mh:noexp] “turns off the more specific terms in both parts, searching for only the one Subheading therapy attached directly to only the one MeSH term hypertension.”
Measured (26 August 2026) — this is the largest single effect on the page:
| Query | Records |
|---|---|
"hypertension/therapy"[mh] |
117,170 |
"hypertension/therapy"[mh:noexp] |
18,973 |
A 6.2-fold difference between two queries that differ by one modifier. Anyone who assumes hypertension/therapy means “records about treating hypertension, specifically” is reasoning about a set roughly one-sixth the size of the one they will actually get.
How MeSH interacts with Automatic Term Mapping
You cannot reason about a MeSH search without understanding what PubMed does to an untagged term, because the two paths produce different result sets from the same words.
Documented. An untagged term goes through Automatic Term Mapping (ATM), which checks it against the subject translation table, then the journals table, then the author and investigator indexes. If a subject match is found, “the term will be searched as MeSH (that includes the MeSH term and any specific terms indented under that term in the MeSH hierarchy), and in all fields.” NLM’s example: child rearing becomes "child rearing"[MeSH Terms] OR ("child"[All Fields] AND "rearing"[All Fields]) OR "child rearing"[All Fields].
That is why an untagged search often out-recalls a tagged one: ATM has already built the MeSH-OR-free-text hedge for you. Tagging with [mh] discards the free-text half of it. Whichever path a term takes, the expansion is only visible if you go and look at it: expand the row in your search history and read the Search Details translation on the Advanced Search page.
The four things that switch ATM off
All four are documented, and each is a live way to lose the MeSH half of a search without noticing:
- A search field tag. “Search field tags turn off Automatic Term Mapping, limiting your search to the specified term only.”
- Double quotes / phrase searching. “When you enter search terms as a phrase, PubMed will not perform automatic term mapping that includes the MeSH term and any specific terms indented under that term in the MeSH hierarchy.” NLM’s example:
"health planning"retrieves records indexed to Health Planning but not the narrower Health Care Rationing, Health Care Reform or Health Plan Implementation. Quoting a concept to be “more precise” therefore silently un-explodes it. - Wildcards. “Wildcards turn off Automatic Term Mapping and the process that includes the MeSH term and any specific terms indented under that term.” NLM’s example:
"heart attack*"will not map to Myocardial Infarction at all. Truncating a term to catch plurals costs you its entire MeSH mapping. - A hyphen. Hyphenated phrases are handled as phrases. Terms must begin with at least four characters before a wildcard, incidentally, which rules out truncating short stems.
Tagged terms still map — unless you quote them
Documented: “A tagged term is checked against the subject translation table, and then mapped to the appropriate MeSH term(s); entry terms tagged with [mh] also map to the appropriate MeSH term(s)… To search for the exact term only and turn off mapping to multiple MeSH terms, enter the tagged MeSH term in double quotes.”
Measured (26 August 2026), and this is where the documentation’s precision matters:
| Query | Records | Reading |
|---|---|---|
heart attack[mh] |
206,603 | The entry term maps to the descriptor. |
"myocardial infarction"[mh] |
206,603 | Identical — the mapping resolved to exactly this descriptor. |
"heart attack"[mh] |
0 | Quoting suppresses the mapping, and Heart Attack is not itself a descriptor. |
odontalgia[mh] |
3,036 | Same pattern with a second entry-term/descriptor pair. |
"toothache"[mh] |
3,036 | |
"neoplasms/dh"[mh] |
0 | Quoting also suppresses expansion of the two-letter abbreviation, which unquoted returns 3,199. |
NLM’s wording is exact — it says to quote “the tagged MeSH term“. The measured consequence, which the documentation does not spell out, is that quoting an entry term rather than a descriptor returns nothing. Quote only strings you have confirmed are preferred descriptors in the MeSH Database.
There is a documented safety net, and you should not lean on it. When a search whose terms were tagged during ATM retrieves zero results, PubMed “triggers a subsequent search using ‘Schema: all'”, removing the automatically added tags and searching each term in all fields. Observed only: issuing "heart attack"[mh] through the E-utilities esearch API on 26 August 2026 returned a count of 0 rather than a fallback set, so the behaviour is not something to rely on outside the web interface. Either way, a zero-result MeSH search that quietly becomes an all-fields search is not a result you want appearing in a reported strategy.
Why MeSH alone under-retrieves the recent literature
This is the part that turns a syntax problem into a review-quality problem, and it is the reason no competent search strategy is MeSH-only.
PubMed is larger than MEDLINE. A citation enters PubMed when the publisher deposits it and moves through processing stages that NLM exposes as searchable citation status subsets. Documented definitions:
publisher[sb]— “citations recently added to PubMed via electronic submission from a publisher, and are soon to proceed to the next stage”. Bibliographic data not yet reviewed. No MeSH.inprocess[sb]— “MeSH terms will be assigned if the subject of the article is within the scope of MEDLINE.” No MeSH yet.medline[sb]— “citations that have been indexed with MeSH terms, Publication Types, Substance Names, etc.” Has MeSH.pubmednotmedline[sb]— citations that “will not receive MEDLINE indexing” because the journal is not indexed for MEDLINE, the article is out of scope, or it predates the journal’s selection. Will never have MeSH.
Measured (26 August 2026). Across the whole database, all[sb] returns 41,057,533 citations and medline[sb] returns 33,611,819 — so 18% of PubMed carries no MeSH indexing at all. Broken down by publication year, the shape of the problem becomes obvious:
| Publication year | Citations | Indexed for MEDLINE (has MeSH) | Share with MeSH |
|---|---|---|---|
| 2020 | 1,641,259 | 1,193,380 | 72.7% |
| 2025 | 1,882,004 | 1,088,346 | 57.8% |
| 2026 (partial year) | 1,312,874 | 640,585 | 48.8% |
The 2026 residue decomposes exactly: 198,826 at publisher[sb] (15.1%), 110,242 at inprocess[sb] (8.4%), and 363,221 at pubmednotmedline[sb] (27.7%). The four subsets sum to the full 1,312,874.
Be precise about what that means, because it is two different problems. The publisher and inprocess records — 23.5% of the 2026 total — are a genuine lag: they will acquire MeSH and the gap will close as the year matures, which is what the 2020 row shows. The pubmednotmedline records are a permanent absence: no amount of waiting gives them MeSH. Any claim that “MeSH catches up eventually” is only half true, and the permanent half is the larger one.
What that costs a real search
Measured (26 August 2026), one ordinary concept, MeSH-only against MeSH-or-free-text:
| Date limit | "sepsis"[mh] |
"sepsis"[mh] OR "sepsis"[tiab] |
MeSH-only recall |
|---|---|---|---|
| All years | 158,361 | 246,356 | 64.3% |
2026[dp] |
3,922 | 9,778 | 40.1% |
Across the whole database a MeSH-only search on this concept retrieves about two-thirds of what the combined search finds. Restricted to the current year, it retrieves two-fifths. The under-retrieval is concentrated precisely where a review is most exposed — the newest trials, the ones that arrived after your protocol was registered, the ones a peer reviewer will name.
And it is silent. There is no warning, no flag, no partial-coverage notice. You get 3,922 well-ranked, entirely genuine records.
The fix, and how to report it
- Build every concept as a MeSH-OR-free-text block.
("sepsis"[mh] OR sepsis[tiab] OR septic*[tiab]), not"sepsis"[mh]. This is standard practice in Cochrane and PRISMA-conformant reviews for exactly the reason measured above, and it is why Boolean nesting is the load-bearing skill in a search strategy rather than a formality. - Search the title/abstract fields deliberately.
[tiab]covers title and abstract;[tw](text word) is broader still and includes MeSH terms among other fields. - Diagnose the gap before you defend the search. Run your strategy, then re-run it with
NOT medline[sb]appended. Whatever comes back is what a MeSH-only version of that strategy would have missed. If that set is large or contains obviously eligible studies, your free-text arm is too narrow. - Beware the update search. NLM documents that MeSH Date
[mhda]“is initially set to the Entry Date[EDAT]when the citation is added to PubMed; citations added to PubMed more than twelve months after the date of publication have EDAT and MHDA set to the date of publication.” A date-limited update search built on the wrong date field will re-surface old records, miss newly indexed ones, or both. - Report the strategy verbatim, with the date you ran it. PRISMA 2020 and its search extension PRISMA-S expect the full strategy per database; see CASRAI’s guides to PRISMA and systematic review methodology and the systematic review protocol template. Counts here move daily, and so will yours.
Where this fits in a search strategy
MeSH is one half of a pair. This page and CASRAI’s page on Google Scholar’s advanced search operators are complementary halves of the same problem: Google Scholar gives you enormous, loosely structured reach with almost no query language and no controlled vocabulary at all; PubMed gives you a precise, fully specified query language over a vocabulary that only covers the records someone has already indexed. Each fails in the way the other does not, and both fail silently.
Beyond that pairing: the vocabulary problem repeats every time you cross a database boundary — MeSH does not exist in Embase, where Emtree is a separately maintained thesaurus with a different tree and an opt-in explosion default, and ERIC, Web of Science and the rest each have their own. Structure the question first with PICO or a PICOT question; cover what no database indexes with a grey-literature search; and if you are only trying to establish that a source is peer-reviewed rather than run a review, searching for peer-reviewed articles is a different and simpler job. For background on what the underlying database actually is, see the dictionary entry on the MEDLINE database, and for the wider tooling landscape, CASRAI’s research tools hub.
One thing MeSH is not for: choosing the keywords you submit with your own manuscript. Many biomedical journals ask for MeSH terms there, but the goal is discoverability rather than retrieval precision — see how to choose keywords for a research paper.
What we could not verify
- Every count on this page is a snapshot taken on 26 August 2026 via the E-utilities
esearchAPI. PubMed adds citations daily and MEDLINE indexing runs continuously, so all of them will have changed. The ratios should be stable in direction; re-run any figure you intend to cite. - The “Schema: all” fallback is documented by NLM for PubMed searches, but the E-utilities API returned a plain zero count rather than a fallback result set in our test. We did not establish where the boundary between the two behaviours lies.
- Indexing lag is not a published service level. The year-by-year percentages above demonstrate that a lag exists and roughly how it closes, but NLM does not publish a target interval from deposit to MEDLINE indexing, and we do not assert one.
- Retrospective re-indexing when a new descriptor is introduced is not something we could confirm as a general NLM policy from primary documentation. Check the “year introduced” on any descriptor central to your strategy and, where the literature predates it, cover the earlier period with free text.
- The split between algorithmic and human indexing over time is searchable via
indexingmethod_automated/indexingmethod_curated/indexingmethod_manual, but we did not test the year-by-year composition, and NLM’s own technical documentation is the place to go for how those values were applied historically.
Frequently asked questions
What is the difference between [mh] and [majr] in PubMed?
[mh] retrieves records where the heading was assigned at all; [majr] retrieves only those where it was assigned as a major topic, marked with an asterisk on the record. Measured on 26 August 2026, "sepsis"[majr] returned 116,066 against 158,361 for "sepsis"[mh] — a 27% reduction. Both explode by default.
Do MeSH terms explode automatically in PubMed?
Yes. NLM documents that “MeSH terms in PubMed automatically include the more specific MeSH terms in a search.” Use [mh:noexp] to turn it off. Note that this is the reverse of the Ovid convention, where explosion is opt-in with exp.
How do I attach a subheading to a MeSH term?
Write it as Heading/Subheading — for example neoplasms/diet therapy, or with the two-letter abbreviation, neoplasms/dh. The [mh] tag is optional; [majr] may be substituted. Only one subheading can be directly attached to a heading; for two aspects, write two terms and combine them with OR.
Why does hypertension/therapy return so many more records than I expected?
Because an attached-subheading term explodes along both hierarchies at once — the MeSH tree and the subheading tree. NLM states it “always includes” both. Measured 26 August 2026: "hypertension/therapy"[mh] returned 117,170 and "hypertension/therapy"[mh:noexp] returned 18,973, a 6.2-fold difference.
Does putting a MeSH term in quotation marks change the search?
Yes, in two ways. Quoting turns off Automatic Term Mapping, so the term is no longer expanded down the MeSH hierarchy — NLM’s example is that "health planning" excludes the narrower headings. And quoting a tagged term turns off mapping to multiple MeSH terms, which means quoting an entry term rather than a preferred descriptor returns nothing: measured 26 August 2026, heart attack[mh] returned 206,603 while "heart attack"[mh] returned 0.
Can I rely on MeSH terms alone for a systematic review search?
No. MeSH exists only on records already indexed for MEDLINE. Measured 26 August 2026, 48.8% of citations with a 2026 publication date carried MEDLINE indexing, against 72.7% of 2020 citations — and for one ordinary concept, a MeSH-only search retrieved 40.1% of what the same concept retrieved as MeSH OR title/abstract when limited to 2026. Build every concept as a MeSH-OR-free-text block.
Will unindexed PubMed records eventually get MeSH terms?
Some will, some never will. Records at publisher[sb] and inprocess[sb] are awaiting indexing. Records at pubmednotmedline[sb] will not receive MEDLINE indexing at all — because the journal is not indexed for MEDLINE, the article is out of scope, or it predates the journal’s selection. On 26 August 2026 that permanent category accounted for 27.7% of all 2026-dated citations, more than the two temporary categories combined.
How can I tell what my MeSH search is missing?
Append NOT medline[sb] to your strategy. Everything it returns is a record with no MeSH indexing that your free-text arm caught — which is exactly the set a MeSH-only version of the same strategy would have missed. If it is large, or contains obviously eligible studies, widen the free-text arm before you report the search.
Where do I look up the correct MeSH term?
The MeSH Database at ncbi.nlm.nih.gov/mesh, or the index in PubMed’s Advanced Search builder. Check the preferred descriptor, the entry terms, the tree position (which tells you what explosion will pull in), the allowable subheadings, and the year the descriptor was introduced.








