Skip to main content
v2026.11,772 entries · CC-BY 4.0

Trial Sequential Analysis: Monitoring Boundaries for Cumulative Meta-Analysis

How Trial Sequential Analysis adapts group sequential monitoring boundaries and required information size to cumulative meta-analysis, and what the Cochrane Handbook actually says about using it.

Ask CASRAI · included with Regulatory Radar

Ask about Trial Sequential Analysis: Monitoring Boundaries for Cumulative Meta-Analysis

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Written and maintained by CASRAI Editorial Board

Last updated

Every time a systematic review team updates a meta-analysis with a newly published trial, they are doing something statistically closer to a data monitoring committee reviewing interim results than to a single, one-time significance test. A cumulative meta-analysis re-tests the pooled effect estimate against the same significance threshold each time a study is added — and each of those re-tests is a fresh chance for a false-positive to appear purely by accident, the same repeated-testing problem a clinical trial faces when a data monitoring committee looks at accumulating data more than once before the trial’s planned end. Trial Sequential Analysis (TSA) is the methodology built specifically to correct for this: it adapts the group sequential monitoring boundaries developed for individual randomized trials to the cumulative meta-analysis setting, so that a conclusion reached partway through the evidence base — “yes, this effect is real,” “no, it isn’t,” or “we don’t have enough evidence yet” — is actually protected against the same inflated false-positive and false-negative risk that an uncorrected single trial would face from looking too often.

TSA was developed by the Copenhagen Trial Unit, and its free software is the most common way reviewers apply it in practice. This guide works through the actual mechanism — the required information size, the diversity adjustment for between-trial heterogeneity, and how alpha-spending monitoring boundaries get drawn onto a cumulative meta-analysis — and then covers something most explainers skip: the Cochrane Handbook’s own, fairly cautious, position on when TSA should and shouldn’t be used to draw a review’s main conclusions.

The problem TSA solves: repeated significance testing across meta-analysis updates

A single meta-analysis, computed once from a fixed set of trials and tested against the conventional two-sided alpha of 0.05, controls its false-positive rate exactly the way a single trial’s final analysis does. The complication is that most influential meta-analyses aren’t computed once — they’re updated as new trials publish, and each update effectively re-tests the accumulating evidence against the same threshold. This is directly analogous to the Type I error inflation problem covered in this site’s guide to alpha-spending functions and stopping boundaries: every additional look at growing data is an additional opportunity for the test statistic to cross the significance threshold by chance alone, even when the true underlying effect is null. A meta-analysis updated and re-tested five or six times over a decade, each time at an uncorrected 5% threshold, carries meaningfully more than a 5% chance of a spurious “significant” result somewhere along the way — and because updates tend to stop once a result looks significant, that inflation is not symmetric: it biases toward premature, overconfident conclusions of benefit or harm.

TSA treats each meta-analysis update the same way a group sequential trial design treats each interim look: as a “look” that has to be accounted for in the overall error budget, not tested in isolation.

Required information size: the meta-analysis equivalent of a trial’s sample size

A conventional randomized trial is powered in advance — a required sample size, calculated from the assumed effect size, variance, and desired Type I/II error rates, that the trial needs to reach before its result can be trusted at face value. TSA applies the same logic to the accumulated evidence in a meta-analysis: it calculates a required information size (RIS), the total number of participants across all included trials the pooled evidence needs before a conventional (unadjusted) significance test on the meta-analysis could be considered adequately powered, using the same parameters — the anticipated effect size, the underlying event proportions or variance, and the target Type I and Type II error rates — as a single trial’s sample-size calculation.

Meta-analyses, unlike a single trial, also have to account for between-trial heterogeneity, because a pool of small, variable trials contains less independent statistical information than the same total number of participants recruited into one uniform trial. TSA adjusts for this with a diversity measure, denoted D², producing a diversity-adjusted required information size (DARIS): the RIS inflated by however much heterogeneity is estimated to be reducing the pooled evidence’s effective information content. D² is closely related to, but not identical to, the I² statistic covered elsewhere on this site — both quantify how much of the total variation across trials is due to real between-trial differences rather than chance, but D² is defined specifically within the TSA framework to adjust the information-size calculation, not as a general-purpose heterogeneity descriptor. Because both the information-size calculation and the diversity adjustment depend on which pooling model produced the point estimate, TSA has to be run on top of an explicit choice between fixed-effect and random-effects pooling, using whichever inverse-variance weighting approach (DerSimonian-Laird, REML, or a Hartung-Knapp adjustment to the confidence interval) the review has already committed to for its primary analysis — TSA is not a separate statistical model, it’s a correction layered on top of the pooling model already in use.

Monitoring boundaries: adapting a single trial’s stopping rules to cumulative evidence

Once the required (diversity-adjusted) information size is set, TSA constructs monitoring boundaries across the cumulative Z-curve — the running plot of the pooled test statistic each time a new trial is added — using the same Lan-DeMets alpha-spending construction, built on the earlier O’Brien-Fleming and Pocock boundary shapes, that a single trial’s data monitoring committee uses to decide whether an interim result is convincing enough to act on. The mechanics carry over directly: an O’Brien-Fleming-shaped boundary is conservative early (it takes an extreme result to cross it while little information has accumulated) and relaxes toward the conventional threshold as the evidence approaches the required information size, so that the total probability of a false-positive crossing across every trial-addition “look” still sums to the nominal alpha rather than being spent all at once on the first update.

TSA constructs the same two-sided family of boundaries a single group sequential trial can carry: an efficacy (benefit) boundary on each side of the null, and, where specified in advance, a futility boundary constructed the same way conditional power calculations are used within a single trial — a signal that the accumulating evidence is unlikely to ever cross the efficacy boundary even with the full required information size reached, so continuing to test for benefit is no longer worthwhile. A cumulative meta-analysis’s Z-curve crossing a monitoring boundary before reaching the required information size is read the same way a trial crossing an interim stopping boundary is read: as evidence strong enough, adjusted for the number of looks taken, to support a conclusion without waiting for more trials. A Z-curve that never crosses either boundary and has not yet reached the required information size is read as inconclusive — not as evidence of no effect, and not as license to treat the current pooled estimate as final.

Reading a TSA chart against a conventional forest plot

A TSA output is not a replacement for the conventional forest plot a systematic review already reports — it’s an additional chart layered on top of the cumulative Z-curve, showing the same running pooled estimate a standard cumulative meta-analysis would produce, but now bounded by the monitoring boundaries and a vertical marker for the required information size. A pooled effect that looks “significant” on a conventional forest plot (its confidence interval excludes the null) can still fall inside the TSA monitoring boundaries and short of the required information size — which is exactly the scenario TSA exists to flag: a result that clears the unadjusted significance bar but has not yet accumulated enough adjusted evidence to be trusted as final. This is also the direct meta-analysis analogue of what publication bias assessment does for a different threat to a pooled estimate’s validity — both ask whether an apparently conclusive result actually is one, from two different angles.

Cochrane’s actual position on Trial Sequential Analysis

TSA is a real, actively used and actively debated methodology — but it is not something the Cochrane Handbook for Systematic Reviews of Interventions recommends applying by default to a standard review update. Chapter 22 of the current Handbook, on prospective approaches to accumulating evidence, states its position plainly: formal sequential meta-analysis approaches, including TSA, are discouraged for updated meta-analyses in most circumstances within the Cochrane context, and should not be used for the main analyses or to draw a review’s main conclusions.

The Handbook does carve out two situations where sequential methods are considered appropriate:

  • A prospectively planned series of trials, where the meta-analysis itself functions as the primary, pre-specified analysis for a coordinated group of trials whose investigators plan the data collection and pooling together from the outset — closer to a true multi-site group sequential trial than to an opportunistic, retrospective pooling of independently run studies.
  • A pre-planned secondary analysis, specified in the review’s protocol in advance, with its effect-size, heterogeneity, and error-rate assumptions explicitly justified rather than chosen after seeing how the evidence has accumulated.

The practical implication for a review team considering TSA: it is on strongest methodological footing when it’s part of the review’s pre-specified protocol from the start, used to characterize how conclusive the accumulated evidence genuinely is, rather than retrofitted onto an already-published meta-analysis as a way to add a veneer of extra statistical rigor to a conclusion the team has already reached. Reviewers should also be aware TSA is applied inconsistently in the published literature — a systematic methodological review of TSA use in published meta-analyses (the METSA review) specifically set out to catalogue common mistakes and misapplications, which is itself a signal that running the software correctly is not the same as applying the method correctly.

When TSA is and isn’t the right tool

TSA is most useful where a meta-analysis is genuinely being treated as an accumulating, living body of evidence — a research question where new trials keep publishing, a review is updated periodically, and the team wants to know not just “what does the pooled estimate say right now” but “is the pooled estimate conclusive, or are we still accumulating toward enough adjusted information to trust it.” It is a poor fit for a one-off meta-analysis of a fixed, closed set of trials with no expectation of future updates, where a conventional pooled estimate with an honestly reported heterogeneity assessment already answers the question the review is asking. Because the required information size calculation depends on assumptions about the anticipated effect size and variance, an analysis run with those assumptions chosen after seeing the data — rather than pre-specified — loses much of the protection TSA is meant to provide, which is precisely the concern behind Cochrane’s own cautious framing above.

Frequently asked questions

Is Trial Sequential Analysis the same as a cumulative meta-analysis?

No. A cumulative meta-analysis is simply the practice of re-running the pooled estimate each time a new trial is added, without any statistical correction for the number of times that re-test has happened. TSA is a specific correction applied on top of that cumulative process — monitoring boundaries and a required information size that account for the repeated-testing problem a plain cumulative meta-analysis ignores.

Does TSA replace GRADE certainty-of-evidence assessment?

No. TSA addresses random error from repeated significance testing specifically; it does not evaluate risk of bias, inconsistency, indirectness, imprecision from other sources, or publication bias, all of which GRADE assessment covers separately. A meta-analysis can reach its required information size and still warrant a lower GRADE certainty rating for other reasons.

What is diversity (D²) in TSA, and how is it different from I²?

Both quantify between-trial heterogeneity as a proportion of total variability, but I² (covered in this site’s heterogeneity guide) is a general-purpose descriptive statistic reported alongside any meta-analysis, while D² is defined specifically within the TSA framework as the adjustment applied to the required information size calculation to account for that heterogeneity’s effect on the evidence’s effective information content.

Can TSA be applied retrospectively to an existing meta-analysis?

It can be run on existing data, but Cochrane’s own guidance is that sequential methods are best used for a review’s main conclusions only when planned prospectively — either as the primary analysis for a coordinated series of trials, or as a protocol-specified secondary analysis with justified assumptions. A retrospective TSA run on an already-completed, already-interpreted meta-analysis is weaker evidence than one planned in advance.

For the underlying single-trial statistical machinery TSA builds on, see Group Sequential Designs: Alpha-Spending Functions and Stopping Boundaries, Interim Analyses in Clinical Trials, and Futility Analysis and Conditional Power. For the meta-analysis methodology TSA sits on top of, see Heterogeneity in Meta-Analysis, Inverse-Variance Weighting in Meta-Analysis, Fixed-Effect vs. Random-Effects Meta-Analysis, and Meta-Regression in Meta-Analysis.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.