Skip to main content
v2026.11,610 entries · CC-BY 4.0

Crossover Trial Design and Carryover

How crossover trials work, why the carryover effect threatens the comparison, how washout periods are justified, and when an irreversible outcome or long-acting treatment requires a parallel design instead.

Ask about Crossover Trial Design and Carryover

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

A crossover trial is a design in which every participant receives more than one of the interventions under comparison, in sequence, rather than being randomized to receive only one. In the simplest version — a two-period, two-treatment (2×2) crossover — each participant is randomized to one of two sequences, AB or BA: one group takes Treatment A first and then, after a break, Treatment B; the other takes them in the opposite order. Because each participant contributes data under both treatments, that participant effectively serves as their own control, and the comparison of interest becomes a within-subject difference rather than a between-subject one.

That single change is what makes crossover designs attractive: between-subject variability — age, baseline severity, genetics, lifestyle, anything that makes one person’s response differ from another’s regardless of treatment — is removed from the treatment comparison entirely, because it cancels out when a subject is compared against themselves. The trade-off is a design-specific threat that a parallel-group design does not have to manage at all: carryover. This guide covers how crossover mechanics work, why carryover threatens the comparison, how a washout period is supposed to neutralize it, and the specific situations where a parallel design is required instead because no washout can fix the problem.

How Crossover Design Works

A crossover trial has three structural elements that a parallel-group trial doesn’t need to distinguish: sequence (the order of treatments a given subject receives, e.g. AB or BA), period (the time block — first, second, sometimes third or later — during which a given treatment is administered and outcomes measured), and treatment (the intervention itself). Randomization is applied to sequence, not directly to treatment: a subject doesn’t get randomized to “Treatment A,” they get randomized to a sequence that includes both A and B in some order.

This structure is what makes the crossover design statistically efficient. In a parallel two-arm trial, the treatment effect is estimated by comparing the mean outcome of the A-arm subjects against the mean outcome of the B-arm subjects — two different groups of people, so all the person-to-person variability shows up in the noise term. In a crossover trial, the treatment effect for each subject is estimated as that subject’s own A-period outcome minus their own B-period outcome, and the average of those individual differences is the trial’s estimate. Because within-subject variability is typically much smaller than between-subject variability for most physiological and symptom-based outcomes, the same sample size yields more statistical power — or, equivalently, a smaller sample size yields the same power — than a parallel design would need. The N-of-1 trial pushes this logic to its extreme, applying the same multi-period crossover logic to a single patient to determine what works for that individual specifically, rather than to a cohort to determine what works on average; it’s a genuinely different use case from the trial-level crossover design this guide covers, but the underlying mechanic — randomized, repeated within-subject comparison — is the same.

The Carryover Effect: The Central Threat to Crossover Validity

Carryover is the residual effect of a treatment given in an earlier period that persists into a later period and contaminates the measurement of whatever treatment is being given at that later point. If a subject in the AB sequence still has pharmacologically active drug A in their system — or a lingering physiological, symptomatic, or even psychological effect of having taken it — when period B’s outcome is measured, that measurement no longer reflects Treatment B alone. It reflects some mixture of B’s true effect and A’s leftover effect.

Carryover is easy to conflate with a period effect (a systematic difference between period 1 and period 2 that has nothing to do with treatment order — practice effects on a cognitive task, seasonal change in a chronic-symptom trial, disease progression over calendar time), but the two are statistically distinguishable and require different fixes. A period effect biases both sequences the same direction and can be adjusted for in the analysis. Carryover is sequence-specific — it only affects the sequence where the carrying-over treatment came first — which is exactly what makes it dangerous: it can masquerade as a genuine treatment-by-period interaction and bias the treatment effect estimate itself, not just add noise.

The classical approach (Grizzle’s two-stage test) tested for a sequence effect first and used a parallel-groups-only analysis of period 1 data alone if that test was significant, on the logic that a detected sequence effect signals carryover. This approach is now widely regarded as underpowered and unreliable for this purpose in a typical crossover-sized trial — the test for a sequence effect has poor power to actually detect real carryover, and “no significant carryover on this test” is not the same claim as “no carryover occurred.” The methodological consensus that has replaced it is to design the washout to make carryover implausible on pharmacological or biological grounds *before* the trial starts, rather than trying to detect and statistically correct for it after the fact. If a design genuinely cannot rule out carryover in advance, that’s a signal the crossover design itself may be the wrong choice, not a problem to solve at the analysis stage.

Washout Periods: Justifying the Interval, Not Just Including One

A washout period is the interval between treatment periods during which no study intervention is given, intended to let the previous treatment’s effects fully dissipate before the next period’s outcomes are measured. The washout has to be long enough that carryover becomes implausible, but the correct length is a judgment call specific to what’s actually being measured, not a fixed default.

For a drug with a straightforward pharmacokinetic profile, the standard planning heuristic ties washout duration to the drug’s elimination half-life: after five half-lives, roughly 97% of the drug is eliminated, which is conventionally treated as pharmacologically negligible. The table below is computed directly from that rule for a drug with an 8-hour half-life:

Half-lives elapsed Washout duration Drug remaining
1 8 hours 50.00%
2 16 hours 25.00%
3 24 hours 12.50%
4 32 hours 6.25%
5 40 hours 3.13%
6 48 hours 1.56%

(Figures computed directly from the residual-fraction formula (1/2)^n for n half-lives; substitute the actual elimination half-life of the study drug to get the real washout duration for a specific protocol — this table illustrates the method, not a universal number.)

The half-life rule only covers pharmacokinetic carryover. Two other carryover mechanisms need a different justification entirely: pharmacodynamic/physiological carryover, where a drug is cleared from the body quickly but its downstream effect (e.g., a receptor downregulation, a healed tissue change, an altered biomarker) persists long after the drug itself is gone; and non-pharmacological carryover in behavioral, dietary, device, or psychological interventions, where “washout” means the practice effect of having learned a skill, the psychological effect of having tried a therapy, or a dietary/behavioral adaptation wearing off — none of which follows a half-life curve at all. For these, the washout has to be justified from the specific mechanism (a validated recovery timeline from prior studies of the same intervention) rather than borrowed from a pharmacokinetic formula.

Worked Example: What the Within-Subject Design Actually Buys

The efficiency gain from a crossover design over a parallel design depends on how strongly correlated a given subject’s two period outcomes are (denoted rho). The table below is computed from the standard two-sided, continuous-outcome sample-size formulas — parallel: n per arm = 2(z_alpha/2 + z_beta)^2*sigma^2/delta^2; crossover: n per sequence = (z_alpha/2 + z_beta)^2*2*sigma^2*(1-rho)/delta^2 — for a 1-SD effect size, alpha=0.05 two-sided, power=0.80:

Design Total N required Reduction vs. parallel
Parallel-group (2 arms) 32
Crossover, rho = 0.3 (weak within-subject correlation) 22 31.3% fewer participants
Crossover, rho = 0.5 (moderate) 16 50.0% fewer participants
Crossover, rho = 0.7 (strong) 10 68.8% fewer participants

The pattern holds generally: the more correlated a subject’s own repeated measurements are — true for most stable physiological and chronic-symptom outcomes — the more a crossover design shrinks the sample size needed for the same statistical power. That efficiency gain is the entire reason to accept the added complexity of managing carryover in the first place; if rho is low (the outcome is noisy or unstable within a person over time), the efficiency case for crossing over weakens substantially and the design’s main advantage largely disappears.

When You Need a Parallel Design Instead

A crossover design is only valid when a subject can plausibly return to a comparable baseline state between periods. Several situations make that assumption untenable, and no washout period — however long — fixes them:

  • Irreversible outcomes. Death, cure, surgical outcomes, permanent organ damage, or any endpoint that, once it happens, cannot be “washed out” and re-measured under the other treatment in a later period. A subject who is cured, or who dies, cannot then cross over to the comparator.
  • Long-acting or depot treatments with uncertain elimination. Some biologics, depot injections, and disease-modifying agents have effects that persist for months after the drug itself is undetectable, or a washout duration that isn’t reliably established — the pharmacokinetic-only heuristic above simply doesn’t apply, and no defensible washout length can be specified.
  • Progressive or unstable conditions. If the condition being studied changes materially between periods regardless of treatment — a degenerative disease, a condition that resolves on its own, a rapidly evolving acute illness — the “return to baseline” assumption crossover relies on breaks down, and period effects become impossible to separate cleanly from treatment effects.
  • Curative or one-shot interventions. Vaccines, single-administration gene therapies, and surgical procedures aim to produce a lasting or permanent state change; there is no meaningful “second period” to cross over into.
  • High expected dropout across periods. Crossover designs require each subject to complete multiple periods; any within-subject attrition (someone who completes period 1 but withdraws before period 2) removes that subject from the within-subject comparison entirely and can introduce a dropout pattern correlated with treatment — a risk a parallel design, where each subject only needs to complete one period, doesn’t share to the same degree.

When any of these apply, a parallel-group design is the methodologically required choice, not a fallback taken only for convenience — the within-subject efficiency gain a crossover offers is not worth trading for a biased or uninterpretable estimate.

Design Variants Beyond the Simple 2×2

The AB/BA design is the simplest case; several established variants extend the same logic:

  • Latin square / Latin square designs extend crossover logic to three or more treatments and three or more periods, using a balanced square that ensures every treatment appears in every sequence position an equal number of times, which helps separate treatment effects from period effects when there are more than two conditions.
  • N-of-1 trials apply repeated crossover cycles within a single patient rather than across a cohort, answering “does this work for this specific person” instead of “does this work on average.”
  • Case-crossover designs look similar on the surface — each subject serves as their own control — but are a fundamentally different, observational (not experimental) design used to study transient exposures and acute outcomes retrospectively, comparing a case period against the same person’s own earlier reference period rather than randomizing prospective treatment sequences. Don’t conflate the two: this guide covers the prospective, randomized experimental crossover trial.

Reporting a Crossover Trial

Because a crossover trial’s validity depends so heavily on design choices that are easy to under-report — the washout rationale, whether carryover was assessed, how period and sequence effects were modeled, how dropouts during the crossover were handled — CONSORT has published a dedicated reporting extension specific to randomized crossover trials, sitting alongside the base CONSORT 2010 checklist used for parallel-group randomized controlled trials. A protocol or manuscript that doesn’t state and justify the washout duration, and doesn’t describe how carryover and period effects were handled in the analysis, is missing information a reader needs to evaluate whether the crossover assumption actually held.

Frequently Asked Questions

What’s the difference between a period effect and a carryover effect?

A period effect is a systematic difference between the first and second (or later) measurement occasions that has nothing to do with which treatment came first — it affects both sequence groups equally and can be statistically adjusted for. Carryover is sequence-specific: it’s the residual effect of a treatment given in an earlier period bleeding into the outcome measured in a later period, and it only affects the sequence where that treatment came first, which is what makes it a threat to the treatment effect estimate itself rather than just added noise.

How long should a washout period be?

For treatments with a well-characterized pharmacokinetic elimination profile, five elimination half-lives is a common planning heuristic, leaving roughly 3% of the original drug remaining. For treatments with pharmacodynamic effects that outlast elimination, or non-pharmacological interventions (behavioral, dietary, device, psychological), the half-life heuristic doesn’t apply — the washout has to be justified from evidence about how long that specific intervention’s effect persists, not borrowed from a formula.

Can I just test for carryover statistically and use a crossover design anyway if the test is non-significant?

This was the classical approach (Grizzle’s two-stage test), but it’s now widely considered unreliable because the test for a sequence effect typically has low power in a normally-sized crossover trial — failing to detect carryover statistically is not strong evidence that carryover didn’t occur. The stronger approach is to design the washout to make carryover biologically implausible before the trial starts.

Is a crossover design always more efficient than a parallel design?

Only when within-subject correlation (rho) between a subject’s repeated measurements is reasonably high. The efficiency gain scales with rho: strong within-subject correlation can cut the required sample size by roughly two-thirds versus a parallel design for the same power, but weak correlation erodes that advantage substantially, and it disappears entirely once carryover, irreversibility, or condition instability rule out crossing over at all.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.