Skip to main content
v2026.11,610 entries · CC-BY 4.0

Superiority, Non-Inferiority and Equivalence Trials

Superiority, non-inferiority and equivalence trials ask three different questions and use three different margin structures. This guide distinguishes the three by hypothesis and decision rule, then walks through a fully computed, reproducible worked example of setting a defensible non-inferiority margin using the fixed-margin (preservation-of-effect) method.

Ask about Superiority, Non-Inferiority and Equivalence Trials

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

Every comparative trial protocol makes a design decision before enrollment begins: what would count as a positive result? Three distinct answers are in routine use — superiority, non-inferiority, and equivalence — and they are not stylistic variants of the same test. Each specifies a different null hypothesis, a different margin (or absence of one), and a different rule for what the resulting confidence interval has to do before the trial can claim success. Which one applies is fixed in the protocol and statistical analysis plan; it is a design choice, not something read off the data after the fact.

The three trial-objective types

All three compare a new (or investigational) arm against a control on the same endpoint. What differs is the question being asked and, for two of the three, where a pre-specified margin sits relative to zero.

Design Question asked Null hypothesis (H0) Margin shape A positive result licenses
Superiority Does the new arm work better than the comparator? No difference between arms None — any difference favoring the new arm counts A claim that the new arm is better
Non-inferiority Is the new arm not unacceptably worse than an active control? New arm is worse than control by at least the margin M One-sided: a single lower margin, −M A claim of “not meaningfully worse” only — never, on its own, a claim of superiority
Equivalence Is the new arm neither meaningfully better nor meaningfully worse? True difference lies outside the interval (−Δ, +Δ) Two-sided: symmetric bounds, −Δ and +Δ A claim of “clinically indistinguishable” in either direction

A useful shorthand: superiority tests whether the difference is different from zero; non-inferiority tests whether it is worse than a threshold in one direction only; equivalence tests whether it stays inside a band around zero in both directions. Bioequivalence testing (comparing a generic drug’s pharmacokinetic exposure to a reference product) is a close statistical relative of the equivalence design but follows its own fixed regulatory rule — see the note below — and the two should not be treated as interchangeable terms.

How the decision rule actually works: confidence intervals, not just p-values

All three designs are usually evaluated the same practical way: compute a confidence interval for the between-arm difference, then check where that interval falls relative to the margin.

  • Superiority succeeds if the (typically two-sided 95%) confidence interval for the difference excludes zero in the favorable direction.
  • Non-inferiority succeeds if the upper bound of a one-sided 97.5% confidence interval (numerically identical to the upper bound of a two-sided 95% CI) for how much worse the new arm is stays below the margin M. The lower bound is not part of the decision rule — a non-inferiority trial is not designed to detect or rule out superiority, though a hierarchical test for superiority can be pre-specified as a secondary step if non-inferiority is met first.
  • Equivalence succeeds only if the entire two-sided confidence interval for the difference falls inside (−Δ, +Δ) — both the upper and lower bound have to clear their respective margin. This is why equivalence trials generally need larger samples than a non-inferiority trial with a numerically similar margin: there are two boundaries to clear instead of one, using the same data.

This is also why a “negative” superiority result and a non-inferiority or equivalence finding are not interchangeable. Failing to reject the superiority null only means no difference was detected at the trial’s power — it says nothing about whether the true difference is small enough to fall inside a margin that was never defined, because a superiority design was never powered or analyzed against one.

Worked example: setting a defensible non-inferiority margin

Illustrative composite — not a real trial. The inputs below (historical event rates, sample sizes) are hypothetical assumptions chosen to demonstrate the method; they are not drawn from any actual drug, device, or published study. Every number that follows the inputs is calculated directly from them using the standard normal-approximation formulas shown, not asserted as a result.

The most common defensible approach to setting a non-inferiority margin is the fixed-margin (preservation-of-effect) method: use historical placebo-controlled evidence to estimate how much benefit the active control has over placebo, then set the margin as a fraction of the most conservative (smallest-magnitude) plausible estimate of that benefit — never the point estimate, and never the largest estimate.

Step 1 — Estimate the control’s historical effect over placebo

Suppose historical placebo-controlled trials of the active control (illustrative inputs) recorded a treatment-failure rate of 30% on placebo (n=400) and 18% on the active control (n=400). The risk difference is −12.0 percentage points, with a standard error of 2.99 percentage points computed from those two independent proportions. The resulting 95% confidence interval for the control’s effect is [−17.9, −6.1] percentage points.

Step 2 — Take the conservative bound and apply a preservation fraction

The bound of that interval closest to zero — 6.1 percentage points — is the smallest defensible estimate of what the control actually does, after accounting for historical sampling error. Applying a commonly used 50% preservation fraction (retain at least half of the control’s known effect) gives a margin of M = 3.07 percentage points. A looser preservation fraction (say, 75% retained, i.e. a wider margin) would let a more genuinely inferior treatment “pass”; a stricter one produces a smaller, more defensible margin at the cost of a larger required trial — margin-setting is a real trade-off, not a formality, which is exactly why it is typically negotiated with a regulator before the pivotal trial begins rather than fixed unilaterally by the sponsor.

Step 3 — Size the trial against that margin

Using the standard sample-size formula for two independent proportions under a non-inferiority design (one-sided α = 0.025, 90% power, assuming the new treatment truly matches the control’s historical 18% failure rate), the margin from Step 2 requires 3,292 participants per arm (6,584 total). That is a large trial for a 3.07-point margin — a direct, calculable consequence of setting a tight, defensible margin rather than a loose one. This is the trade-off sponsors are negotiating when they push for a wider margin: a smaller, cheaper trial, at the cost of a claim that tolerates a larger true loss of effect.

Step 4 — What different trial outcomes on this design would actually show

Three illustrative post-trial scenarios on the exact design above, computed the same way the real analysis would be:

  • Scenario A — new and control arms both observe an 18% failure rate. Observed difference 0.0 points, 95% CI [−1.9, +1.9] points. The upper bound (1.9) is below the 3.07-point margin: non-inferiority demonstrated.
  • Scenario B — new arm observes 21% vs. control’s 18%. Observed difference +3.0 points, 95% CI [+1.1, +4.9] points. The upper bound (4.9) exceeds the margin: non-inferiority not demonstrated — despite a difference that many readers would eyeball as “close.”
  • Scenario C — new arm observes roughly 20.6% vs. control’s 18%, deliberately close to the margin itself. Observed difference +2.6 points, 95% CI [+0.7, +4.5] points. The upper bound (4.5) still exceeds the 3.07-point margin: non-inferiority not demonstrated.

Reporting all three honestly matters: a tight, defensible margin means several plausible real-world outcomes — including ones where the new treatment looks only modestly worse — fail to clear it. That is the margin doing its job, not a flaw in the trial. A sponsor tempted to widen the margin after seeing results like Scenario B or C, in order to convert a failed non-inferiority claim into a passing one, is exactly the kind of post-hoc margin manipulation regulatory review and the CONSORT extension for non-inferiority/equivalence trials (Piaggio et al., JAMA, 2012) are designed to catch by requiring the margin and its justification to be stated before the trial, not fitted to the result afterward.

Analysis population: another place the three designs diverge

For a superiority trial, intention-to-treat (ITT) analysis is the standard, conservative primary analysis, because non-adherence and dropout dilute a real effect toward the null — the same bias that makes ITT the cautious choice works against the trial finding a difference that isn’t there. For a non-inferiority or equivalence trial, that logic reverses: dilution toward the null can make two genuinely different treatments look more similar than they are, which biases the trial toward a false non-inferiority or equivalence conclusion. This is why non-inferiority and equivalence trials are generally expected to report both ITT/Full Analysis Set and per-protocol analyses as co-primary, requiring both to agree before the claim is considered robust — see our guide on intention-to-treat analysis and the related ITT vs. per-protocol comparison for the full mechanics.

Common pitfalls

  • Switching hypotheses after seeing the data. Testing for superiority after a non-inferiority trial succeeds is a valid, pre-specifiable hierarchical step. Testing for non-inferiority only after a superiority trial fails, using a margin chosen post hoc, is not a defensible non-inferiority claim under CONSORT or regulatory guidance — the margin and the hierarchy both have to be fixed in the protocol before unblinding.
  • Using placebo as the comparator in a non-inferiority design. The entire logic depends on the active control having a reliable, well-established effect over placebo to preserve a margin against. Without that established effect (the “assay sensitivity” requirement), showing “not unacceptably worse than the control” provides no evidence the new treatment does anything at all.
  • Treating a wide margin as free. A margin is a clinical judgment about how much benefit can be acceptably lost, agreed with a regulator — not a knob to turn to shrink the required sample size.
  • Conflating equivalence trials with pharmacokinetic bioequivalence studies. Bioequivalence testing (comparing a generic’s rate and extent of absorption — AUC, Cmax — to a reference product) uses a fixed regulatory rule, not a negotiated clinical margin: the entire 90% confidence interval for the test-to-reference ratio must fall within 80.00%–125.00%. It shares the two-sided “inside a band” logic of a clinical equivalence trial but is a distinct, narrower framework.

Frequently asked questions

Can one trial test for both non-inferiority and superiority?

Yes, if the hierarchy is pre-specified in the statistical analysis plan. A common design tests non-inferiority first; if the margin is met, the same data can then be tested for superiority, since superiority is the stronger claim and implies non-inferiority. Running the tests in the reverse order — testing non-inferiority only after a superiority claim fails — is a recognized source of bias unless that hierarchy was fixed before unblinding.

Who actually sets the non-inferiority or equivalence margin?

The sponsor proposes it, grounded in historical evidence of the control’s effect (as in the worked example above) and clinical judgment about how much of that effect can be acceptably lost. In practice it is typically negotiated and agreed with the regulator — an FDA Type B/C meeting, for example — before the pivotal trial begins, rather than fixed unilaterally.

Why do equivalence trials usually need more participants than a non-inferiority trial with a similar-sized margin?

Because equivalence requires the entire confidence interval to clear two boundaries (−Δ and +Δ) rather than one. Clearing two bounds with the same confidence level, from the same data, is a more demanding statistical requirement than clearing a single one-sided bound, which generally translates into a larger required sample for a comparably tight margin.

Does a non-inferiority trial ever prove the new treatment is as good as the control?

No — it demonstrates only that any true difference is unlikely to exceed the pre-specified margin, not that the treatments are identical or that no difference exists at all. “Not unacceptably worse, within a defined and justified margin” is a narrower and more precise claim than “equally good,” and framing a non-inferiority result as proof of equal effectiveness overstates what the design can show.

Related CASRAI resources

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.