Written and maintained by CASRAI Editorial Board
Last updated
A Bayesian analysis reports one posterior, but that posterior is a joint product of the data and the prior you chose. Prior sensitivity analysis is the practice of re-running the analysis under a small set of alternative, defensible priors and checking whether the conclusion that actually matters for the paper — the direction of an effect, whether an interval excludes a value of interest, a decision to stop a trial — changes. If it doesn’t, the result is described as robust to prior specification. If it does, that’s not a failure to hide; it’s the honest finding that the data alone don’t settle the question, and the prior is doing real work.
This is the step a reviewer is checking for when a manuscript reports “we used a weakly informative prior” without saying what would have happened under a different one. A single posterior, however carefully computed, is not evidence that the choice of prior was inconsequential — that has to be shown, not asserted.
Why this is different from an ordinary robustness check
In frequentist work, sensitivity usually means re-running an analysis under alternative model specifications — a different covariate set, a different functional form. Bayesian prior sensitivity analysis targets a component that has no equivalent in frequentist inference at all: the prior distribution itself is an input you chose, not something estimated from the data. Two analysts can run identical models on identical data and reach different posteriors purely because they specified different priors — a source of disagreement a p-value calculation cannot produce, because the null in a significance test is a single, fully specified point with no analogous researcher-chosen input.
The sensitivity is largest in exactly the situations many applied studies actually run in: small samples, rare events, or hierarchical models with few groups at the level where a variance component is estimated. As sample size grows, the likelihood dominates and most reasonable priors converge on similar posteriors — which is also why a sensitivity analysis is most necessary precisely when it’s tempting to skip it for time.
What to vary, and what to hold fixed
A useful sensitivity analysis varies one thing at a time against a clearly stated baseline, not everything at once:
- Hyperparameters within the same prior family. If the baseline is a Normal(0, 1) prior on a regression coefficient, re-fit with Normal(0, 0.5) and Normal(0, 2.5) — a narrower and a wider version of the same shape — and compare the posterior mean/median and the width of the credible interval each time.
- The prior family itself. Swap a Normal prior for a Student-t with a few degrees of freedom (heavier tails, less aggressive shrinkage of extreme values), or a Beta prior for a Uniform prior on a probability parameter, and check whether the substantive conclusion holds.
- Informative vs. weakly informative vs. flat. If a genuinely informative prior was used (built from a prior study, an elicited expert range, or a meta-analytic estimate), report the analysis again under a deliberately weaker version of it and under a flat/reference prior, so a reader can see how much the informative choice is contributing to the result on its own.
Hold the likelihood, the data, and every other model component fixed across these runs — the point is to isolate what changes when the prior changes, not to re-run a general robustness check under the same label.
A reportable protocol
- State the baseline prior and its justification before showing any sensitivity result — what it is, and why it was chosen (weakly informative to regularize estimation, informative from a specific external source, or a standard reference/flat choice for a first analysis).
- Define what “the conclusion” means for this analysis in advance of running alternatives — a sign, an interval excluding zero, a probability crossing a decision threshold, a Bayes factor crossing an evidence-strength boundary. Sensitivity analysis without a pre-specified target invites picking whichever framing looks best after the fact.
- Fit the pre-specified alternative priors (typically 2–4 is enough to demonstrate a pattern; more than that usually adds noise rather than information for a reader).
- Report every run, not just the ones that agree with the baseline — a table or a small forest-style plot of the point estimate and interval under each prior, side by side, is the standard presentation and lets a reader see the pattern at a glance rather than trust a one-line summary.
- State the verdict plainly. “The estimated effect and its 95% credible interval were materially unchanged across the four priors tested” is a real robustness claim. If the runs disagree, say so and explain what that disagreement implies for how confidently the result should be read — don’t bury a sensitive result inside a supplementary table without discussing it in the main text.
A worked illustration
Suppose a small pilot trial (n = 40) estimates a treatment effect with a Normal(0, 1) prior on the standardized effect size and reports a posterior mean of 0.42 with a 95% credible interval of [0.02, 0.83] — just clearing zero. A prior sensitivity analysis re-fits the same model with Normal(0, 0.5) (a more skeptical, shrinkage-heavy prior) and Normal(0, 2.5) (a much flatter, less informative prior). If the posterior mean moves only slightly (say 0.35 to 0.48) and every interval still excludes zero, that’s a genuine robustness finding worth stating in those terms. If the narrower prior pulls the interval to include zero while the flatter one doesn’t, the honest conclusion is that this particular result depends materially on how much the prior shrinks small samples toward zero — a finding that should change how the result is described, not one to quietly omit.
Where this fits in a Bayesian workflow
Prior sensitivity analysis is one specific, checkable step inside a broader discipline of transparent Bayesian reporting that also includes prior predictive checks (does the prior alone, before seeing data, generate plausible outcomes?), posterior predictive checks (does the fitted model reproduce features of the observed data?), and convergence diagnostics for the sampler itself. None of these substitute for the others — a model can converge cleanly and still be highly sensitive to the prior, and a prior can look sensible in a prior predictive check while still meaningfully shifting the posterior in a small sample. Treat prior sensitivity analysis as one item on that checklist, reported alongside the others, not as a stand-in for the full set.
Frequently asked questions
How many alternative priors are enough?
There’s no fixed rule, but two to four well-chosen alternatives — spanning a genuinely narrower and a genuinely wider (or a different-family) prior around the baseline — is standard practice and enough to show a reader the pattern. Testing dozens of minor variations doesn’t add information and can read as searching for whichever set makes the result look robust.
Does a sensitive result mean the analysis is wrong?
No. It means the data alone don’t fully determine the conclusion, and the prior is doing real inferential work. That’s a legitimate, reportable finding — the problem is only in not checking for it, or checking and not reporting it.
Is this the same as a prior predictive check?
No. A prior predictive check asks whether the prior, before any data are seen, generates outcomes that are plausible on their own terms (e.g., does a prior on a probability ever imply a probability greater than 1, or a prior on a rate imply implausibly extreme rates). A prior sensitivity analysis asks a different question: after fitting the model to real data, does the posterior conclusion change materially across otherwise-reasonable prior choices. Both are useful and neither replaces the other.
Do I need this for every Bayesian analysis I run?
It matters most exactly where priors have the most leverage: small samples, rare events, and hierarchical models with few groups at the level where a variance component is estimated. In a large-sample analysis where the likelihood clearly dominates any reasonable prior, a brief sensitivity check is still good practice to demonstrate that dominance, but the stakes — and a reviewer’s reasonable skepticism — are lower.
Related CASRAI resources
- Interpreting Bayes Factors: Evidence Scales and What They Mean — Bayes factors are themselves sensitive to the prior in a way p-values are not; this guide covers how to read the resulting evidence scale.
- Bayesian Hierarchical Models — the setting (few groups, shrinkage of a group-variance parameter) where prior sensitivity is typically most consequential.
- Bayesian vs. Frequentist Statistics: What the Interval Actually Means — background on how a credible interval differs from a confidence interval.
- Markov Chain Monte Carlo — convergence diagnostics are a separate, necessary check alongside prior sensitivity, not a substitute for it.
- Statistical Analysis Plan Template — pre-specifying the sensitivity-analysis protocol (which priors, which target conclusion) belongs in the analysis plan, not decided after seeing results.








