Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Stepped-Wedge and Cluster-Randomized Trial Designs: When Sponsors Use Them and Regulatory/Ethical Considerations

A practical guide for research administrators and IRBs to stepped-wedge and cluster-randomized trial designs: why sponsors choose them, how they differ from individually randomized trials, and the CONSORT, Common Rule, and Ottawa Statement considerations that apply.

Stepped-wedge and cluster-randomized trial (CRT) designs randomize groups — clinics, hospital wards, schools, primary care practices, geographic regions — rather than individual participants. Sponsors and investigators reach for them when the intervention under study is delivered, or can only realistically be delivered, at the level of an organization or community rather than to one person at a time. For research administrators, IRB/REC staff, and sponsors, these designs raise distinct methodological, regulatory, and ethical questions that a standard individually randomized controlled trial (RCT) protocol does not — this guide covers both.

Cluster-randomized vs. stepped-wedge: the core distinction

A parallel cluster-randomized trial allocates each cluster to either the intervention or the control arm for the duration of the study, the same way a standard RCT allocates individuals, except the unit of randomization is the group. A stepped-wedge cluster-randomized trial (SW-CRT) is a specific variant in which every cluster starts in the control condition and crosses over to the intervention condition at a randomly assigned time point, in a staggered sequence, until all clusters have received the intervention by the end of the trial. Visually, the design resembles a set of stairs (or a wedge) when clusters are plotted against time and treatment condition — hence the name.

Both designs are usually analyzed at the individual level (outcomes measured on people within clusters) while accounting statistically for within-cluster correlation, but they answer somewhat different practical questions: a parallel CRT is closer in logic to a standard two-arm trial; a stepped-wedge design is closer to a phased, sequential rollout that happens to be randomized.

Why sponsors and investigators choose these designs

Three recurring justifications appear across the methodological literature and in real trial protocols:

  • The intervention is inherently delivered at the group level. Policy changes, staff training programs, clinical decision-support tools, and quality-improvement interventions typically cannot be randomized to individual patients within the same clinic without contaminating the control condition — a nurse trained on a new protocol will apply it to every patient she sees, not just the ones assigned to the intervention arm. Randomizing at the cluster level avoids this contamination.
  • Logistical or practical feasibility. A stepped-wedge design is often chosen when an intervention is going to be rolled out to every site anyway (for operational, funding, or policy reasons) and simultaneous rollout across all sites is not logistically possible — the staggered rollout schedule doubles as a randomization schedule, so a program that was already going to happen sequentially becomes a trial largely for free.
  • Ethical appeal of eventual universal access. Because every cluster eventually receives the intervention, sponsors and ethics committees sometimes view a stepped-wedge design as more acceptable than a parallel design when there is a reasonable expectation the intervention is beneficial, since no site is permanently denied it for the study’s duration. A 2019 methodological review in the International Journal of Epidemiology notes this ethical framing is common in trial justifications, while also cautioning that it is not automatically a stronger justification than a well-designed parallel trial and needs to be argued on its merits case by case, not assumed.

These designs are especially common in implementation science, health-services research, quality-improvement evaluation, and public-health policy evaluation — contexts where the object of study is a program, protocol, or system change rather than a drug or device administered to individuals.

Design mechanics research administrators should know

  • Complete vs. incomplete designs. In a complete stepped-wedge design, every cluster is measured at every time period. Incomplete designs measure only a subset of clusters at some time periods, usually to reduce data-collection burden; this trades some statistical efficiency for lower cost and logistical simplicity.
  • Cross-sectional vs. cohort designs. A cross-sectional SW-CRT samples different individuals within a cluster at each time point (e.g., a new set of clinic patients each period); a cohort SW-CRT follows the same individuals across the whole study. This choice has major implications for consent procedures, attrition, and the statistical model used.
  • Intracluster correlation (ICC). Because outcomes within the same cluster tend to be more similar to each other than to outcomes in a different cluster, sample size and power calculations for both parallel CRTs and SW-CRTs must account for the ICC — a naive individual-level sample size calculation will understate the true sample size needed and is a common protocol-review error.
  • Sequence and step allocation are themselves randomized. Which cluster switches to intervention at which step is determined by randomization, not by convenience or by which site is ‘ready first’ — preserving randomization integrity here is what distinguishes a stepped-wedge trial from an unrandomized phased rollout evaluation.

Reporting standard: the CONSORT extensions

The core CONSORT statement was developed for individually randomized parallel-group trials and does not adequately cover cluster-level randomization on its own, so two extensions apply:

  • The CONSORT extension for cluster randomised trials (Campbell et al., 2012 update) addresses parallel CRTs specifically — reporting the number of clusters (not just participants) randomized to each arm, the method used to identify and recruit clusters, and how clustering was accounted for in the analysis.
  • The CONSORT extension for stepped-wedge cluster randomised trials (Hemming et al., published 2018) adds design-specific reporting items on top of the cluster-trial extension — critically, item 2a requires investigators to explicitly justify why a stepped-wedge design was chosen over a parallel design or an individually randomized design. Reviewers assess whether that justification is substantive (for example, the intervention could not feasibly be withheld from any site, or simultaneous rollout was operationally impossible) rather than a post-hoc rationalization for a design decision that was actually driven by convenience.

Sponsors preparing a protocol or manuscript should build against both extensions from the outset — retrofitting cluster-level reporting into a document drafted against the standard CONSORT checklist is a common source of late-stage rework.

Regulatory and IRB/ethical considerations

Cluster-level randomization complicates several assumptions that individual-participant human-subjects review is built around, and this is the area sponsors and administrators most often underestimate.

Who counts as a research subject?

In a standard trial, the research subject is unambiguous: the person who is randomized, intervened upon, and from whom data are collected. In a cluster trial, those three roles can separate — a hospital ward might be randomized and its staff trained (intervention delivered to staff), while outcome data are collected from patients who never consented to anything and may not even be aware a trial is underway. Determining who is a human research subject for regulatory purposes, and therefore whose consent (or waiver of consent) is required, is one of the first questions an IRB or REC has to resolve, and the answer differs by design.

The Ottawa Statement

The most widely cited ethics framework built specifically for this problem is the Ottawa Statement on the Ethical Design and Conduct of Cluster Randomized Trials (Weijer, Taljaard, Grimshaw, Edwards, and Eccles, 2012; the full statement and a shorter precis for researchers and ethics committees were published in Trials and PLOS Medicine respectively). It sets out 15 recommendations across seven domains, including: justifying the choice of a cluster design, when research ethics committee review is required, how to identify who counts as a research participant at the cluster and individual level, appropriate consent models, the role of gatekeepers who consent on behalf of a cluster, assessment of benefits and harms at both the cluster and individual level, and protection of vulnerable clusters and individuals. A 2025 citation analysis in Research Integrity and Peer Review found the Ottawa Statement remains the dominant reference framework in the field but also identified gaps that newer guidance has not yet fully closed — sponsors relying on it should treat it as the authoritative starting point, not the final word, and check for design-specific supplementary guidance from their funder or national research-ethics body.

Consent models

Because a whole cluster is allocated to a study condition, individual informed consent to randomization itself is often not meaningful or even possible — an individual patient cannot consent to whether their clinic is randomized to the intervention arm. Ethics frameworks and IRBs typically distinguish several consent models: individual consent to specific research procedures (still required where the trial involves data collection, sampling, or interventions applied directly to individuals); gatekeeper consent, where an authorized representative of the cluster (a clinic director, health authority, or community leader) consents on behalf of the cluster to its participation, which does not substitute for individual consent to research procedures but can be appropriate for the cluster-level allocation decision itself; and a waiver or alteration of consent, which US IRBs may grant under the Common Rule (45 CFR 46.116(f)) when the research involves no more than minimal risk, could not practicably be carried out without the waiver, and the waiver will not adversely affect participants’ rights and welfare — a common fact pattern for practice-level quality-improvement or implementation trials where the intervention is a change to how care is organized rather than a direct clinical intervention on the patient.

What this means for IRB submissions

  • Identify explicitly, in the protocol, who is a human subject at each phase of the trial (the individuals delivering the intervention, the individuals it is delivered to, and anyone contributing outcome data) — do not assume this is self-evident to reviewers.
  • Propose and justify a specific consent model per participant category rather than a single blanket consent approach; a stepped-wedge trial in particular may need different consent handling for early-phase (control) versus later-phase (intervention) participants within the same cluster.
  • Address gatekeeper authority explicitly — who has legitimate authority to consent on behalf of a cluster, and what happens if a cluster’s gatekeeper agrees but individual members object.
  • Assess and report benefits and harms separately at the cluster level (e.g., organizational burden, reputational risk to a clinic) and the individual level (e.g., risk to a patient), since the two are not interchangeable in a cluster design.
  • If citing minimal-risk status to support a consent waiver, document that assessment specifically against the trial’s own procedures — minimal-risk determinations do not transfer automatically from one cluster trial to another just because both are ‘practice-level.’

Limitations and when not to use these designs

Stepped-wedge and cluster designs are not a free substitute for individual randomization. They generally require larger total sample sizes than an individually randomized trial to achieve equivalent statistical power, because of the ICC penalty described above. Stepped-wedge designs in particular are vulnerable to time-varying confounding — because clusters cross over at different calendar times, any secular trend unrelated to the intervention (seasonal effects, a concurrent policy change, a pandemic) can be mistaken for an intervention effect if the analysis does not adequately model time. A 2024 systematic review of stepped-wedge trials in high-impact journals found meaningful inconsistency across published trials in how well design and analysis choices were justified and reported, reinforcing why the CONSORT extension’s explicit-justification requirement exists rather than being a formality.

Frequently asked questions

Is a stepped-wedge trial the same as a phased program rollout?

Only if the order and timing of the rollout is determined by randomization. A phased rollout where sites are sequenced by readiness, convenience, or administrative priority is not a stepped-wedge trial and does not support the same causal inference, even if it superficially looks similar on a timeline.

Does a cluster-randomized trial need IRB approval if individuals never consent directly?

Generally yes — cluster-level allocation does not exempt a study from human-subjects review; it changes what the IRB is being asked to approve (the consent model, the identification of subjects, and often a waiver request) rather than removing the review requirement. See the Ottawa Statement discussion above.

How is sample size different for a cluster or stepped-wedge trial compared with an individually randomized trial?

Both require inflating the individual-level sample size to account for the intracluster correlation coefficient (ICC); the exact inflation factor depends on the design (parallel vs. stepped-wedge), the number of clusters and cluster size, and, for stepped-wedge designs, the number of sequences/steps. Statistical input at the protocol-design stage, not after the fact, is standard practice for this reason.

Which CONSORT extension applies to my trial?

Parallel cluster-randomized trials should report against the CONSORT extension for cluster randomised trials (Campbell et al.); stepped-wedge cluster-randomized trials should report against the CONSORT extension for stepped-wedge cluster randomised trials (Hemming et al., 2018), which builds on the cluster-trial extension with stepped-wedge-specific items.

For related methodology, see CASRAI’s guides to the CONSORT statement and justifying sample size in a manuscript, and the dictionary entries for pragmatic trial and equivalence trial. For the broader clinical-research operations context, see the clinical research pillar.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →