Skip to main content
v2026.11,610 entries · CC-BY 4.0

Futility Analysis and Conditional Power: When to Stop a Trial for Futility

How conditional power is calculated at an interim look, why the assumed treatment effect changes the answer, and why non-binding futility rules cannot be used to “save” alpha for the efficacy boundaries.

Ask about Futility Analysis and Conditional Power: When to Stop a Trial for Futility

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

A futility analysis is an interim look at a trial’s accumulating data that asks a different question than an efficacy analysis does. An efficacy analysis asks: “is the effect already large enough to stop and declare success?” A futility analysis asks: “given what we’ve seen so far, is it still plausible that this trial will reach a positive result if it runs to completion?” The calculation that answers the second question is conditional power — the probability, computed from the data accumulated at an interim look, that the trial will eventually cross its pre-specified significance threshold at the final analysis, under a stated assumption about the treatment effect for the remainder of the study.

This guide works through the conditional-power calculation itself, its Bayesian counterpart (predictive power), and the alpha-spending consideration that determines whether a futility stop is statistically defensible or just an informal judgment call. For the broader group-sequential and alpha-spending machinery that both efficacy and futility boundaries sit inside, see Adaptive Trial Designs: Types, Alpha-Spending & Pre-Specification; for how the monitoring committee that acts on these numbers is structured and firewalled from the sponsor, see DMC Charter: What It Must Include and How Unblinding Procedures Work.

Efficacy Stopping, Futility Stopping, and Continuing

A group-sequential trial with interim monitoring has three possible outcomes at each look, not two. The upper boundary governs stopping for overwhelming efficacy — see Adaptive Trial Designs for how Pocock, O’Brien-Fleming, and Lan-DeMets alpha-spending functions set that boundary so the cumulative probability of a false-positive stays at the trial’s nominal alpha level across every look. A lower boundary — or, more commonly in practice, a conditional-power threshold rather than a fixed test-statistic boundary — governs stopping for futility: the data so far make a positive final result unlikely enough that continuing no longer serves participants or the sponsor. Between the two, the trial simply continues to the next planned look or to completion.

The asymmetry matters. An efficacy boundary that’s crossed lets the trial claim a significant result — it directly consumes a portion of the trial’s Type I error budget, which is exactly what an alpha-spending function is built to control (see p-value and Type I and Type II Errors for the underlying error-rate concepts). A futility boundary that’s crossed does the opposite: the trial stops without rejecting the null, so on its own it doesn’t spend any of the alpha budget the way an efficacy stop does. That asymmetry is exactly why the alpha-spending consequence of futility stopping is subtler than it looks, covered below.

The Conditional Power Calculation

Conditional power is computed at a specific interim look, characterized by its information fraction t (roughly, the proportion of the trial’s total planned statistical information — often approximated by enrolled/events accrued — collected so far, with t running from just above 0 to 1 at the final analysis). At that look, the trial has an observed standardized test statistic zt. The standard formula (from the group-sequential monitoring literature — Lan, Simon & Halperin’s 1982 conditional-power framework, developed further in the Lan-DeMets alpha-spending approach covered in the adaptive-designs guide above) treats the path of the test statistic as Brownian motion with a drift parameter representing the assumed effect, and gives conditional power as:

CP(θ) = 1 − Φ( [ c·√1 − zt·√t − θ·(1−t) ] / √(1−t) )

where c is the final-analysis critical value, θ is the assumed standardized drift for the remainder of the trial, and Φ is the standard normal cumulative distribution function. The whole calculation hinges on θ, which isn’t observed — it’s an assumption the analyst has to choose, and different choices can give very different answers from the identical interim data. Three choices dominate practice:

  • The original design effect. Use the effect size the trial was originally powered to detect. This answers “if the treatment works exactly as well as we assumed when we designed the trial, how likely is a positive result from here?”
  • The observed trend. Extrapolate the effect implied by the data collected so far (θ = zt / √t) forward at the same rate. This answers “if the treatment keeps performing exactly as it has so far, how likely is a positive result?” — and is usually the more sobering number, since early trends regress toward the true effect as more data accrues.
  • The null. Set θ = 0. This is the lower bound — conditional power under “no further true effect at all” — useful as a floor for how bad the picture already looks.

Illustrative Example (constructed for this guide, not a real trial)

The numbers below are an independently computed, reproducible illustration — not a real trial’s data — built to show how differently the three θ assumptions can answer the same interim look. Design: a two-sided test at α = 0.05 (critical value c = 1.960), originally powered at 80% for its planned effect (design drift θH1 = z0.025 + z0.20 = 1.960 + 0.842 = 2.802). At the interim look, information fraction t = 0.5 (halfway through the planned information) and the observed test statistic is zt = 0.80 — a positive but unimpressive trend, not yet close to either boundary.

θ assumption Value of θ Conditional power
Original design effect 2.802 50.4%
Observed trend (zt/√t) 1.131 12.1%
Null (θ = 0) 0 2.4%

The spread is the point. Under the assumption that the treatment performs as originally hoped, this trial still has roughly a coin-flip’s chance of reaching significance — hardly grounds for stopping. Under the assumption that it keeps performing the way it actually has so far, conditional power drops to 12.1%, inside the range (commonly 10%–30% in published DMC charters) that many pre-specified futility rules would flag for review. Reporting a single conditional-power number without saying which θ produced it is close to meaningless — the same interim data supports numbers 20-fold apart depending on the assumption.

These figures were computed directly from the formula above (verified against a normal-CDF implementation checked at four known reference points) and cross-checked with an independent 2,000,000-trial Monte Carlo simulation of the same Brownian-motion model, which reproduced the analytic 50.4% figure to within 0.03 percentage points.

Predictive Power: The Bayesian Alternative

Conditional power fixes θ at one chosen value and reports a single number. Predictive power — sometimes credited to Spiegelhalter, Freedman and Blackburn’s mid-1980s work on trial monitoring — instead treats the future effect as uncertain, places a prior (or uses the posterior formed from the interim data itself) over plausible values of θ, and averages conditional power over that whole distribution rather than conditioning on a single point guess. The practical difference: predictive power is typically lower and more conservative than conditional power computed at the original design effect, because it doesn’t get to assume the design assumption is still correct — it spreads its bet across the genuine uncertainty about what the true effect actually is. Predictive power is less commonly implemented than conditional power in practice because it requires specifying a defensible prior in the protocol or SAP, which is an extra pre-specification burden sponsors don’t always take on, but it’s the more honest answer when there’s real uncertainty about whether the original design assumption still holds by the time of the interim look.

Non-Binding vs. Binding Futility Rules — the Alpha-Spending Consequence

This is the part of futility stopping that’s easy to get wrong, and it’s the reason the calculation above isn’t just descriptive statistics — it has to be built into the trial’s pre-specified alpha-spending plan to remain statistically valid.

A binding futility rule commits the trial, in the protocol, to stopping if conditional power falls below the pre-specified threshold — there is no discretion to continue. Because the sample space genuinely excludes “continue after crossing the futility boundary,” the final-analysis alpha-spending calculation can formally account for that exclusion, which in principle frees up slightly more of the total alpha budget for the efficacy boundaries than a design with no futility rule at all would have.

A non-binding futility rule treats the conditional-power threshold as a recommendation to the data monitoring committee, not a contractual stopping requirement — the DMC (or sponsor) retains the discretion to continue the trial even after conditional power drops below the pre-specified line, if there’s a compelling clinical reason to. Regulators, including the FDA’s adaptive-designs guidance, generally favor non-binding futility rules for exactly this reason: a binding rule that forces a sponsor to stop a trial that later data might have vindicated is an outcome nobody wants to be locked into. The cost of that flexibility is statistical: because the trial could hypothetically continue past a non-binding futility boundary, the final-analysis test still has to be evaluated as though that possibility exists — you cannot retroactively claim the alpha “savings” a binding rule would have earned, even though in practice the trial almost always does stop once conditional power is low. In effect, a non-binding futility rule buys operational flexibility at zero statistical cost to Type I error control, but it also earns no statistical credit — the efficacy alpha-spending boundaries have to be set exactly as conservatively as if the futility rule didn’t exist. This is the single most common misunderstanding in how sponsors talk about futility stopping: treating a non-binding recommendation as if it had licensed a more aggressive efficacy boundary, when it hasn’t.

Choosing a Futility Threshold in Practice

The conditional-power threshold that triggers a futility recommendation — commonly somewhere in the 10%–30% range in published trial designs, though there’s no universal number — has to be written into the Statistical Analysis Plan and the DMC charter before the trial starts, along with which θ assumption(s) will be used to compute it. Pre-specification matters here for the same reason it matters for the efficacy boundary: a threshold and a θ assumption chosen after seeing how the interim data looks isn’t a valid stopping rule, it’s a rationalized decision dressed up as one.

Two practical cautions follow directly from the worked example above. First, conditional power computed early in a trial (low t) is unstable — with only half the planned information collected, a single interim look’s trend can swing the observed-trend conditional power dramatically once more data arrives, which is why many DMC charters restrict futility looks to after some minimum information fraction has been reached (commonly around 0.3–0.5) rather than allowing arbitrarily early futility analyses. Second, because the choice of θ changes the answer this much, DMC charters that pre-specify a futility rule should also pre-specify exactly which θ assumption computes it — the original design effect, the observed trend, or a blend — rather than leaving that choice to be made in the room when the interim data is unblinded.

Common Mistakes

  • Reporting conditional power without stating the θ assumption. As the worked example shows, the same data can support numbers from roughly 2% to roughly 50% depending only on which effect assumption drives the calculation.
  • Treating a low conditional power as proof the treatment doesn’t work. Conditional power is a probability computed under an assumption, not a hypothesis test of the null — a low number is a signal to review, not a conclusion in itself.
  • Assuming a non-binding futility rule earns back alpha. It doesn’t, for the reason covered above — the efficacy boundaries still have to be set as though the trial could continue past the futility line.
  • Running unplanned futility looks at very low information fractions. Conditional power is most volatile exactly when there’s least data to base it on.
  • Confusing conditional power with the trial’s original statistical power. Statistical power is fixed at the design stage under the design assumption; conditional power is recomputed at each interim look using the data actually observed so far — see Statistical Power Analysis for the design-stage calculation this updates.

Frequently Asked Questions

What’s the difference between a futility analysis and an interim efficacy analysis?

An efficacy analysis asks whether the accumulated data already crosses the (alpha-spending-adjusted) significance threshold, allowing the trial to stop and claim a positive result. A futility analysis asks whether the accumulated data makes it unlikely the trial will ever reach that threshold, supporting a recommendation to stop without a positive result. They use different boundaries and, critically, have a different relationship to the trial’s alpha budget — see the alpha-spending section above.

What conditional-power threshold is typically used to stop a trial for futility?

There’s no universal number; published designs commonly use something in the 10%–30% range, computed under the observed-trend or a blended assumption, but the exact threshold and the θ assumption behind it must be pre-specified in the protocol or SAP rather than chosen after seeing the interim result.

Does stopping a trial for futility require spending alpha?

No — crossing a futility boundary doesn’t reject the null, so it doesn’t itself consume Type I error the way crossing an efficacy boundary does. The alpha-spending consequence is indirect: whether the futility rule is binding or non-binding determines whether the trial’s efficacy boundaries can be set any less conservatively, and for the common non-binding case, they cannot.

Is conditional power the same thing as statistical power?

No. Statistical power is a design-stage quantity: the probability of detecting the assumed effect, computed before any data is collected. Conditional power is computed at an interim look, given the data observed so far, and answers a narrower question — the probability of eventually reaching significance from this specific point onward, under a stated assumption about the remaining data.

Can a trial continue after crossing a futility boundary?

Under a non-binding futility rule — the more common convention and the one regulators generally favor — yes: the data monitoring committee retains discretion to recommend continuation despite low conditional power if there’s a compelling reason to. Under a binding rule, no: the protocol commits the trial to stopping once the threshold is crossed.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.