Skip to main content
v2026.11,772 entries · CC-BY 4.0

Subgroup and Sensitivity Analysis in Meta-Analysis: Pre-Specification and Interpretation

How subgroup analysis (does the effect differ across pre-defined groups) and sensitivity analysis (is the pooled result robust to analytic choices) differ, why pre-specification determines credibility, and how to report both honestly.

Written and maintained by CASRAI Editorial Board

Last updated

Subgroup analysis and sensitivity analysis are commonly used interchangeably, and they answer different questions. A subgroup analysis asks whether the treatment effect genuinely differs across pre-defined categories of participants, interventions, or settings — it is a substantive claim about effect modification. A sensitivity analysis asks whether your pooled result is robust to the analytic choices you made along the way — the studies included, the pooling model, the handling of missing data. Confusing the two leads to a common failure mode: presenting an analytic-choice check as if it were a clinical finding, or dismissing a real effect-modification signal as “just a sensitivity check.” Both belong in a well-conducted meta-analysis; neither substitutes for the other.

What each analysis is actually asking

A subgroup analysis splits the included studies into categories — by dose, population, comparator, risk-of-bias rating, or any other study-level characteristic — and asks whether the pooled effect differs meaningfully between them. The question it answers is about the world: does this intervention work differently in older patients, or at a higher dose, or in a particular setting.

A sensitivity analysis re-runs the same pooled analysis under a different analytic decision — a different pooling model, a different set of included studies, a different assumption about missing data — and compares the result to the primary analysis. The question it answers is about your methods: does the conclusion depend on a choice that could reasonably have gone the other way. See CASRAI’s fixed-effect vs. random-effects comparison and the inverse-variance weighting guide for two of the model choices sensitivity analyses most often test.

Why pre-specification is the dividing line for credibility

The single biggest determinant of whether a subgroup finding is believed is whether it was defined before the data were pooled. A subgroup hypothesis chosen after seeing which split produces a significant result is a form of multiple testing dressed up as a finding — test enough subgroups and one will cross p<0.05 by chance alone. This is why a registered protocol matters: a subgroup analysis specified in a PROSPERO registration or a published protocol, with a stated direction of expected effect, carries far more weight than the same split reported for the first time in the results section. The Cochrane Handbook is explicit on this point — subgroup and sensitivity analyses should be pre-specified in the protocol wherever possible, and any analysis added after seeing the data should be clearly labelled as post hoc and interpreted accordingly.

Sensitivity analyses benefit from the same discipline but are held to a slightly different standard: because they are meant to interrogate the robustness of a single primary analysis rather than generate new substantive claims, an unplanned sensitivity analysis is less concerning than an unplanned subgroup analysis — but running dozens of them and reporting only the ones that support the primary conclusion is the same selective-reporting problem in a different guise. Cochrane guidance recommends stating a minimal, pre-specified set of sensitivity analyses rather than an open-ended list assembled after the fact.

Judging whether a subgroup effect is credible

A statistically significant difference between subgroups is not, by itself, evidence that the effect genuinely differs. Methodologists (notably Oxman and Guyatt’s original criteria, and their more recent formalization) have converged on a small set of questions worth asking before trusting a subgroup finding:

  • Was it tested with a formal interaction test, not just separate within-subgroup significance? Two subgroups where one result is “significant” and the other is “not significant” can still show no real difference between them — the comparison that matters is a direct test of the interaction (or the difference in effect sizes with its own confidence interval), not a comparison of two p-values.
  • Was the direction of the effect specified in advance? A hypothesis that merely predicts “the effect will differ” is weaker than one that predicts which subgroup will show the larger effect and why.
  • How many subgroup comparisons were tested? A credible finding among 3 pre-specified subgroups is very different from the same p-value found among 20 exploratory splits.
  • Is there a plausible biological or clinical mechanism? A subgroup difference with a coherent mechanism behind it is more credible than a statistically similar difference with no proposed explanation.
  • Is the difference consistent across studies and settings, or does it depend on one study? A subgroup effect that only appears once a single influential study is included is a fragile finding (see the leave-one-out discussion below).

These criteria were formalized into a structured instrument by Schandelmaier, Briel, and colleagues: ICEMAN (the Instrument to assess the Credibility of Effect Modification Analyses), published in CMAJ in 2020. ICEMAN walks through a consistent set of considerations — along the same lines as the list above — for both randomized trials and meta-analyses, and produces an overall credibility rating from very low to high rather than a bare significance call. It is worth using as a checklist even informally: the value is in forcing the same questions every time, not in the exact scoring mechanics.

Common sensitivity analyses in practice

A handful of sensitivity analyses recur across most meta-analyses because they test the choices reviewers make most often:

  • Leave-one-out (influence) analysis. The pooled estimate is recalculated once per study, each time omitting one study, to see whether any single study is driving the result. It is especially worth running when heterogeneity is high, when there are fewer than ten included studies, or when one study is noticeably larger or more extreme than the rest. Results are best reported as a table of re-estimated pooled effects alongside the primary result, rather than as a separate set of forest plots — that keeps the comparison to the main analysis easy to read.
  • Excluding studies at high risk of bias. Re-running the pooled analysis restricted to studies rated low risk of bias (see CASRAI’s RoB 2 / ROBINS-I guide) tests whether the primary conclusion depends on including weaker evidence.
  • Alternative pooling model. Comparing the fixed-effect and random-effects pooled estimates directly is itself a sensitivity check, particularly when the choice of model was not obvious in advance.
  • Alternative missing-data assumptions. Where studies report incomplete outcome data, comparing the primary result to one obtained under a different assumption about the missing values (see CASRAI’s multiple imputation guide) checks whether the conclusion depends on how missingness was handled.
  • Alternative effect-size metric or analysis population. Repeating the pooled analysis with an alternative but defensible outcome definition, or restricted to a specific comparator, checks whether a result is an artifact of one particular analytic decision rather than a robust signal.

Reporting both without overclaiming

Three habits keep subgroup and sensitivity findings honest in the write-up:

  • Report the interaction test or the between-subgroup difference with its own confidence interval — not just the two within-subgroup p-values side by side.
  • State explicitly whether each analysis was pre-specified or post hoc, and be more cautious in the interpretation of the latter. This is also what a GRADE assessment will weigh when it downgrades certainty for inconsistency.
  • Distinguish “the estimate was robust to X” from “the estimate changed materially under Y, and here is the most plausible reason why” — a sensitivity analysis that changes the conclusion is a genuine, reportable finding about the evidence base, not a result to quietly drop from the manuscript.

PRISMA 2020 expects both subgroup and sensitivity analyses that were actually performed to be reported in the results, whether or not they were pre-specified, precisely so that selective reporting of only the flattering ones cannot happen invisibly. See CASRAI’s guide on detecting and assessing publication bias for the related problem of a whole study, rather than a subgroup or sensitivity check, going unreported. A well-documented statistical analysis plan — see CASRAI’s meta-regression guide for a related pre-specification discussion — is the practical mechanism that makes all of this auditable after the fact.

Frequently asked questions

Is subgroup analysis the same as sensitivity analysis?

No. Subgroup analysis asks whether the true effect differs across categories of studies or participants — a substantive, clinical question. Sensitivity analysis asks whether the pooled result changes under a different analytic choice — a robustness check on your own methods. A finding can be a legitimate sensitivity result without implying anything about effect modification, and vice versa.

How many sensitivity analyses is too many?

There is no fixed number, but Cochrane guidance is to pre-specify a minimal, deliberate set tied to genuine analytic uncertainties (model choice, risk-of-bias exclusions, missing-data handling) rather than running an open-ended list and reporting only the ones that agree with the primary result.

What makes a subgroup effect credible rather than spurious?

A formal interaction test (not separate within-subgroup significance), a pre-specified direction, a small number of tested subgroups, a plausible mechanism, and consistency across studies and settings. The ICEMAN instrument (Schandelmaier et al., CMAJ, 2020) formalizes these considerations into a structured credibility rating for both trials and meta-analyses.

Do sensitivity analyses need to be registered in the protocol?

They should be, wherever the analytic uncertainty is foreseeable at protocol stage — it is the strongest protection against selectively reporting only the sensitivity analyses that support the primary conclusion. An analysis added after seeing the data can still be reported, but should be labelled as post hoc.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about Subgroup and Sensitivity Analysis in Meta-Analysis: Pre-Specification and Interpretation

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.