Skip to main content
v2026.11,772 entries · CC-BY 4.0

Sample Size Justification for a Grant Application or IRB Protocol

How to justify a sample size to a grant reviewer or IRB, not just calculate one: the NIH Approach criterion, 45 CFR 46 risk/benefit review, the pilot-study rule of thumb when no effect size exists, and the attrition math reviewers expect to see.

Ask CASRAI · included with Regulatory Radar

Ask about Sample Size Justification for a Grant Application or IRB Protocol

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Written and maintained by CASRAI Editorial Board

Last updated

A sample size justification for a grant application or IRB protocol is a different document from a power calculation, and reviewers can tell the difference. A power calculation is arithmetic: given an effect size, an alpha, and a target power, it solves for N. A justification is the argument built around that arithmetic — why this effect size, why this alpha, why this many participants given the attrition you expect to lose, and, for a genuine pilot or feasibility study, why a formal power calculation may not even be the right tool. This page covers that argument: how NIH and NSF reviewers actually score it, what an IRB protocol’s “number of subjects” section needs that a manuscript’s Methods section doesn’t, and the specific case — a pilot study with no prior effect-size estimate to build a calculation from — where the honest answer is a defensible rule of thumb, not a fabricated power analysis. It does not cover how to run the calculation itself; for the statistical mechanics (effect size, alpha, power, and the four-quantity relationship among them), see CASRAI’s Power Analysis and Sample Size Calculation guide. It also does not cover how to report a completed study’s sample size after the fact for a journal’s Methods section — see Justifying Sample Size in a Manuscript for that, prospective and retrospective justification read very differently to a reviewer, and confusing the two produces a submission that answers the wrong question.

Why a Grant or IRB Reviewer Reads This Differently Than a Journal Reviewer Does

A journal’s Methods section is read after the fact: the study already happened, and CONSORT and STROBE ask authors to report how the sample size was actually determined so a reader can judge whether a null result reflects a true absence of effect or an underpowered study. A grant application or IRB protocol is read before anything happens, and the question a reviewer is answering is different: should this specific number of participants, at this specific cost and risk, be approved to go forward at all?

For an NIH application, the sample-size justification lives inside the Approach criterion — one of NIH’s five core scored review criteria alongside Significance, Investigators, Innovation, and Environment. A reviewer scoring Approach is asking whether the proposed methods can actually answer the proposed question, and an unsupported or circular sample-size number is one of the fastest ways to depress that score, because it signals the applicant hasn’t thought rigorously about feasibility. NSF’s Merit Review criteria (Intellectual Merit and Broader Impacts) don’t name a sample-size line item as explicitly, but a reviewer evaluating whether the “proposed activities suggest and explore creative, original, or potentially transformative concepts” with a “well-reasoned” plan reads an unjustified N the same way — as a gap in the plan, not a formality.

For an IRB protocol, the number of subjects is a distinct, regulated question under the Common Rule (45 CFR 46). Risk/benefit assessment — one of the Belmont Report’s three core applications, alongside informed consent and selection of subjects — requires the IRB to weigh the number of people exposed to a study’s risks against the value of the knowledge the study can produce. An unjustifiably small sample risks approving a study that cannot answer its own question while still exposing participants to risk for no scientific return; an unjustifiably large one exposes more participants than the question requires. Both are reasons an IRB can require revision before approval, independent of whether a funder ever sees the application at all.

Justification and Calculation Are Not the Same Requirement

Every proposal that collects data needs a sample-size justification — a stated, reproducible reason for the number chosen. Not every proposal needs a formal power calculation to produce that reason, and treating the two as interchangeable is a common way applicants either over-claim precision they don’t have or, worse, back-calculate a fake effect size to make a power calculation output the N they had already decided to use.

A genuine pilot or feasibility study — the kind NIH funds through R21 and R34 mechanisms specifically because no adequate prior estimate exists yet — is the clearest case where a formal power calculation is the wrong tool, not a shortcut around a harder one. NIH guidance explicitly cautions against treating a small pilot’s own effect-size estimate as a reliable input to a downstream power calculation, because a pilot is itself too underpowered to estimate an effect size precisely; using it anyway just launders a shaky number through an arithmetic process that looks more rigorous than it is. The accepted alternative, when there is no prior effect size to power against, is a rule-of-thumb sample size justified on grounds of feasibility and estimation precision rather than statistical power — the most cited version is Julious’s recommendation of roughly 12 participants per group for a pilot study (Julious SA. “Sample size of 12 per group rule of thumb for a pilot study.” Pharmaceutical Statistics. 2005;4(4):287-291), defended on three grounds: it gives a reasonable estimate of the variability needed to design the full-scale study, it is large enough to assess feasibility (recruitment rate, protocol adherence, retention) meaningfully, and it is small enough to be defensible on cost and participant-burden grounds for a study that is explicitly not yet trying to detect an effect. Citing that rationale, by name, is a stronger justification than presenting a power calculation built on borrowed numbers.

Building the Justification When You Do Have a Basis for an Effect Size

When the proposal is for a full-scale, confirmatory study rather than a pilot, the calculation itself belongs in the guide linked above; what belongs here is what a grant or protocol reviewer specifically expects the surrounding paragraph to state, in order:

  • The effect size and where it came from, in the order of defensibility reviewers actually credit: a systematic review or meta-analysis of comparable studies first, a single closely comparable prior study used cautiously second, your own pilot data with its NIH-flagged limitation disclosed third, and a named convention (Cohen’s small/medium/large benchmarks, or the smallest effect of genuine clinical or practical interest) only as a last resort when no field-specific estimate exists at all.
  • Alpha and target power, stated as numbers (conventionally 0.05 and 0.80, though clinical and high-stakes designs increasingly use 0.90) rather than left implicit — a reviewer should not have to infer them from software output pasted into an appendix.
  • The specific statistical test the calculation was built around — a power calculation for a two-sample t-test does not transfer to a chi-square test or a mixed model, and citing the wrong one is a documented, common reviewer objection.
  • An explicit attrition or non-response adjustment, not folded silently into a rounded-up final number. The standard form is straightforward: Nenrolled = Nneeded / (1 − expected dropout rate). A protocol that calculates N = 64 per arm, expects 15% attrition, and proposes to enroll 75 per arm without showing that arithmetic is asking a reviewer to trust a number they cannot check; showing 64 / 0.85 ≈ 76 removes the trust requirement entirely.

A justification that skips straight to a final N without naming these four elements in order is the pattern reviewers most often flag as circular or unsupported — not because the underlying math was necessarily wrong, but because the applicant made the reviewer do the verification work the applicant should have done.

Where This Goes in the Application

The placement differs by document, and using the wrong template signals unfamiliarity with the mechanism as much as a weak justification does:

  • NIH R01/R21/R34-style Research Strategy: sample-size justification belongs in the Approach section, typically within or immediately following the statistical analysis plan subsection — see CASRAI’s NIH Specific Aims guidance for how this connects back to what each Aim is actually designed to detect. A pilot mechanism (R21, R34) should say explicitly that it is a pilot and cite a feasibility-and-precision rationale (Julious or equivalent) rather than presenting a power calculation dressed up as more definitive than the mechanism actually requires.
  • IRB protocol: sample-size justification is typically its own numbered item under “Number of Subjects” or “Subject Selection,” separate from the risk/benefit and consent sections, even though all three inform the same underlying review — see CASRAI’s IRB Protocol Worked Example for how a full submission sequences these sections.
  • Budget justification: sample size is also a cost driver, not just a statistical input — per-participant compensation, procedure, and site-management costs scale directly with N, so a sample-size justification and a budget justification should agree with each other exactly. See CASRAI’s Budget Justification: A Worked Example and Budget Justification Narrative for a Grant Proposal for how that number should reappear in the budget narrative.

Frequently Asked Questions

Does a pilot study need a power calculation at all?

Generally no. A pilot or feasibility study’s sample size is usually justified on grounds of feasibility, recruitment/retention estimation, and precision around variability estimates for the future full-scale study — not on statistical power to detect an effect, since the pilot is not designed to detect one. Julious’s rule of thumb (roughly 12 per group) is the most commonly cited anchor for this rationale; state it and its reasoning explicitly rather than presenting a power calculation the pilot cannot actually support.

What if I have no prior data and no relevant literature at all?

This is rare in practice — most fields have at least a loosely comparable study or a field-convention benchmark — but when it is genuinely true, say so, and justify the sample size against the smallest effect that would be scientifically or clinically meaningful rather than against an effect size you cannot support with any citation. Reviewers credit an honest “no comparable prior estimate exists, so we powered against the minimum clinically important difference of X” far more than a fabricated citation to a marginally related study.

Can I use the same sample-size text in my grant application, IRB protocol, and eventual manuscript?

The underlying calculation and its inputs should be identical across all three — a reviewer or auditor who finds the grant and the IRB protocol citing different effect sizes for the same study has found a real problem. The framing around it should differ: the grant version argues the study is worth funding at this size, the IRB version argues the number of subjects is ethically justified, and the eventual manuscript (see CASRAI’s Justifying Sample Size in a Manuscript guide) reports what was actually determined, in the past tense, per CONSORT/STROBE reporting requirements.

How does attrition inflation interact with a stratified or clustered design?

The same enrollment-inflation logic applies, but attrition should be estimated per stratum or cluster where dropout is expected to vary (a multi-site trial with sites of different retention history, for example), rather than applied as one blended rate across a design where the risk is not actually uniform — see CASRAI’s Statistical Analysis Plan Template for where this level of detail belongs in the SAP itself.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.