Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us
Dictionary termTrack Proposedv2026.1

Statistical Analysis Plan (SAP)

A Statistical Analysis Plan (SAP) is the pre-specified document that translates a clinical trial protocol's objectives and endpoints into the exact statistical methodology that will be used to analyze the trial's data — which populations are analyzed (e.g. intent-to-treat/full analysis set vs. per-protocol set), which statistical tests and models are applied to each endpoint, how missing data and multiplicity are handled, and what any interim or subgroup analyses will consist of. A document only functions as a genuine SAP if it is finalized and version-locked before database lock and unblinding — that is, before anyone with access to outcome data can see how the results turn out. A statistical methods section written or substantively revised after the analysts already know (or could infer) the results is not a valid SAP in the ICH E9 sense, regardless of what it is called, because it no longer prevents the analytic flexibility (selective reporting, post-hoc subgroup mining, outcome-switching) that pre-specification exists to close off.

ByCASRAI Editorial Board
· Last updated 23 Jul 2026

Examples

Worked examples

  • Is an instance

    A Phase 3 trial's protocol names the primary endpoint (e.g. change from baseline in a symptom score at 12 weeks) and states the trial is randomized and double-blind. A separate SAP document, finalized and signed by the lead biostatistician before database lock, then specifies: the primary analysis population (Full Analysis Set, per ICH E9), the exact statistical model (e.g. a mixed model for repeated measures with baseline as covariate), how missing visits will be handled (multiple imputation vs. last-observation-carried-forward, and why), the multiplicity-adjustment method if there are multiple secondary endpoints, and the content of any planned subgroup or sensitivity analyses. Because these choices are locked in advance, an analyst running the primary analysis after unblinding is executing a pre-agreed recipe, not choosing the model that produces the most favorable result.

  • Is an instance

    ClinicalTrials.gov's results-reporting requirements (under FDAAA 801 / the Final Rule, 42 CFR Part 11) require that the statistical analysis performed matches what was planned; sponsors of applicable clinical trials commonly upload the SAP itself as a study document alongside the protocol, and many trial registrations amend the registry record to reference SAP version history so reviewers can see when the plan was locked relative to enrollment completion and database lock.

Counter-examples

Looks similar, but isn't

  • Not an instance

    A brief 'Statistical Methods' paragraph in a journal manuscript, written after the study has already been analyzed and the results are known, describing what analysis was in fact performed. This is a retrospective methods description, not a SAP — a SAP is a prospective, dated, version-controlled document that exists and is finalized before anyone sees outcome data, and its role is to constrain the analysis, not summarize it after the fact.

  • Not an instance

    An investigator deciding, after seeing an interesting pattern in unblinded interim data, to add a new subgroup analysis or switch the primary endpoint to one that now looks more favorable. Even if this is written up formally afterward, it is post-hoc analysis, not SAP-governed pre-specified analysis — the defining feature of a SAP is that the analytic decisions it documents were locked before that knowledge existed.

Editorial commentary

Why the SAP exists as a document separate from the protocol

A clinical trial protocol establishes the trial’s design, eligibility criteria, interventions, and endpoints, and typically includes a summary of the planned statistical approach. The Statistical Analysis Plan (SAP) exists as a separate, more granular document because ICH E9 (“Statistical Principles for Clinical Trials”) requires that the full technical details of the planned analysis be specified and documented before anyone with access to the data can see results that might reveal treatment effects — a level of detail (exact model specifications, handling of missing data, multiplicity strategy, definitions of analysis populations) that would make a protocol unwieldy if embedded directly, and that in practice is often finalized somewhat later than the protocol itself, closer to database lock. FDA’s Guidance for Industry on E9 and ICH E9(R1)’s estimands addendum both describe the SAP as the vehicle for this pre-specification.

Timing: why finalization before database lock and unblinding is the whole point

The SAP’s methodological value depends entirely on when it is finalized relative to two events: database lock (the point after which the trial dataset is frozen and can no longer be modified) and unblinding (the point at which treatment-arm assignment becomes visible to those analyzing the data). A SAP finalized and version-locked before both events documents analytic decisions made in ignorance of outcome data. A SAP written, or materially revised, after unblinding — even if it accurately describes reasonable statistical choices — cannot rule out that those choices were influenced by seeing which analysis produced a more favorable result. This is the direct mechanism by which SAP pre-specification prevents outcome-driven p-hacking and selective subgroup reporting: it removes the opportunity, not just the intent. For trials with a Data Safety Monitoring Board (DSMB) conducting interim looks, the SAP additionally pre-specifies the interim-analysis schedule and any statistical stopping boundaries, since the DSMB’s unblinded access to comparative data makes this the same integrity concern at an earlier stage.

What a SAP typically specifies

  • Analysis populations — how the Full Analysis Set / intent-to-treat population and, where relevant, a per-protocol population are defined; see Intent-to-Treat vs. Per-Protocol Analysis for how these populations differ and why the choice affects estimated treatment effect.
  • Endpoint definitions and statistical models — the exact test or model for each primary, secondary, and exploratory endpoint, not just which endpoints exist.
  • Missing-data handling — the imputation or sensitivity-analysis approach for dropouts and missing assessments, specified in advance rather than chosen once the pattern of missingness in the actual data is known.
  • Multiplicity control — how Type I error is protected when there are multiple endpoints, comparisons, or interim looks.
  • Subgroup and sensitivity analyses — which subgroups will be examined and how, distinguishing pre-specified (confirmatory-adjacent) subgroup analyses from any later exploratory ones, which must be reported as exploratory.
  • Mock shells/tables — many SAPs are accompanied by table, listing, and figure (TLF) shells showing the exact output format the analysis will populate, reinforcing that the analysis was planned in structure, not just described in prose after the fact.

Relationship to the protocol, to registration, and to results reporting

The SAP is downstream of, and must be consistent with, the protocol’s stated objectives and endpoints — a SAP that introduces a new primary endpoint not named in the protocol, or changes the primary analysis population, is a protocol-inconsistency finding that reviewers and journal editors are trained to flag. Prospective trial registration on ClinicalTrials.gov (see also Prospective Clinical Trial Registration) is the public record against which a completed SAP can later be checked: because the registry captures the primary and secondary outcome measures before the trial completes, any mismatch between the registered endpoints and what the SAP (or the eventual publication) actually analyzes is detectable by any reader, which is precisely the accountability mechanism prospective registration is designed to create. Under ClinicalTrials.gov’s results-reporting requirements (FDAAA 801 / the Final Rule, 42 CFR Part 11), sponsors of applicable clinical trials frequently upload the SAP itself as a study document, and many registrations track SAP version history so the timing of finalization relative to enrollment completion and database lock is independently verifiable rather than taken on trust. For the broader design choices the SAP builds on — endpoint selection, sample-size and power calculation, and randomization scheme — see Designing a Clinical Trial: Endpoints, Sample Size, Randomization, SAP.

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="Statistical Analysis Plan (SAP)"
      vocab-term-identifier="https://casrai.org/dictionary/term/statistical-analysis-plan-sap" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/statistical-analysis-plan-sap",
  "name": "Statistical Analysis Plan (SAP)",
  "identifier": "https://casrai.org/dictionary/term/statistical-analysis-plan-sap",
  "description": "A Statistical Analysis Plan (SAP) is the pre-specified document that translates a clinical trial protocol's objectives and endpoints into the exact statistical methodology that will be used to analyze the trial's data — which populations are analyzed (e.g. intent-to-treat/full analysis set vs. per-protocol set), which statistical tests and models are applied to each endpoint, how missing data and multiplicity are handled, and what any interim or subgroup analyses will consist of. A document only functions as a genuine SAP if it is finalized and version-locked before database lock and unblinding — that is, before anyone with access to outcome data can see how the results turn out. A statistical methods section written or substantively revised after the analysts already know (or could infer) the results is not a valid SAP in the ICH E9 sense, regardless of what it is called, because it no longer prevents the analytic flexibility (selective reporting, post-hoc subgroup mining, outcome-switching) that pre-specification exists to close off.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/clinical-research#set",
  "url": "https://casrai.org/dictionary/term/statistical-analysis-plan-sap",
  "sameAs": [],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "dateModified": "2026-07-23T08:13:01",
  "inLanguage": "en"
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →