Reviewers and editors expect a Methods section to answer one question about sample size before they will trust anything that follows: why this many participants, sites, or specimens, and not some other number? This guide covers the writing task of answering that question on the page — how to summarize a power calculation in prose a non-statistician reader can follow, how to state the effect-size assumption a calculation depended on, and how to write an honest, defensible justification when the sample was constrained by feasibility rather than set by a clean a priori calculation. It does not cover how to run the power calculation itself; for the statistical mechanics of choosing a sample size, see CASRAI’s Designing a Clinical Trial: Endpoints, Sample Size, Randomization, SAP guide. This page is about turning that calculation, or the constraints that prevented one, into text a reviewer will accept.
Why This Belongs in the Methods Section, and Why Reviewers Look for It
A sample-size justification is not a formality tacked onto the Methods section — it is part of how a reader evaluates whether the study was capable of answering its own question. A trial or study that was too small to detect a meaningful effect can produce a misleadingly reassuring null result (an under-powered study reporting no significant difference is not the same evidence as an adequately powered study reporting no significant difference), and a study that recruited far more participants than needed raises its own concerns about resource use and, for clinical research, the ethics of exposing more participants than necessary to a study’s risks and burdens.
Reporting guidelines reflect this directly. CONSORT 2010, the reporting standard for randomized trials, makes sample-size justification its own checklist item: item 7a requires authors to report how sample size was determined, including the target difference the calculation was built around, and item 7b requires reporting any interim analyses or stopping guidelines that affected the final number. STROBE, the equivalent guideline for observational studies, requires item 10: an explanation of how the study size was arrived at — either the formal calculation, if one was performed, or the practical considerations that determined it (for example, a fixed available sample or a defined recruitment window) when it was not. Both guidelines explicitly warn against a specific failure mode: justifying a study’s size after the fact using the effect actually observed, or presenting a retrospective (post hoc) power calculation as if it were a real design decision. A journal’s methods reviewers, and increasingly automated pre-submission checks such as SciScore, are checking for exactly this section — write it as a real account of a decision made in advance, not as an afterthought added during revision.
The Two Situations This Guide Covers
Most sample-size justifications fall into one of two situations, and they call for different, but equally honest, writing:
- A formal a priori power calculation was performed. The task is to summarize it clearly and completely, in prose, without dumping the full statistical derivation into the Methods section.
- No formal power calculation determined the final number — the sample was set by a fixed budget, a recruitment window, a rare condition’s limited eligible population, or a pilot/feasibility design. The task is to say so plainly and explain the constraint, rather than construct or imply a calculation that did not actually drive the enrollment target.
Both are legitimate. What is not legitimate, in either case, is a paragraph that reads as though a rigorous a priori calculation set the number when it did not, or a hollow calculation reverse-engineered to match a sample size the investigators had already committed to for other reasons.
Writing the Power-Calculation Summary
When a formal calculation was performed, the Methods section needs enough information for a reader to understand the logic and, in principle, reproduce the number — without reproducing the full statistical workup. A complete summary typically states, in a short paragraph:
- The primary outcome the calculation was built around. Sample size should be justified against the study’s primary outcome, not a secondary or exploratory one — a calculation for anything else does not actually establish that the study is adequately powered for the question it is designed to answer.
- The effect size (or target difference) assumed, and where that number came from — a prior study, a pilot, a clinically meaningful threshold such as a minimal clinically important difference, or a discipline convention (e.g., Cohen’s small/medium/large benchmarks in behavioral research). State the number, not just “a moderate effect was assumed.”
- The significance level (alpha) and power (1−beta) used — conventionally 0.05 and 0.80 or 0.90, but state whatever was actually used rather than assuming the reader will infer it.
- The statistical test or model the calculation assumed, since the formula (and therefore the required sample) differs by test — a two-sample t-test, a chi-square test of proportions, a log-rank test for survival data, and a mixed-effects model for clustered or repeated-measures data all require different inputs and produce different numbers for the same target effect.
- The resulting required sample size, and any inflation applied on top of it — most commonly an anticipated dropout or attrition rate, and for cluster-randomized designs, a design effect derived from the intracluster correlation coefficient.
- The software or method used to run the calculation (e.g., G*Power, PASS, nQuery, or a named R package), which lets a statistical reviewer sanity-check the number.
A worked example of how this reads in practice: “Sample size was calculated using G*Power 3.1, assuming a between-group difference of 10 points on the [named outcome measure] (SD = 15, based on [prior study/pilot]), a two-sided alpha of 0.05, and 90% power, using a two-sample t-test. This yielded a required sample of 96 participants per arm. Anticipating 15% attrition, the enrollment target was set at 113 per arm (226 total).” Every input is stated, the source of the effect-size assumption is named, and the final recruitment number is traceable back through the calculation rather than simply asserted.
Keep the summary to a short paragraph. Full derivations, intermediate formulas, and sensitivity analyses across multiple effect-size assumptions belong in a supplementary file or the statistical analysis plan, not the Methods section itself — the Methods section needs to establish that the calculation was sound and reproducible, not walk the reader through it step by step.
Writing an Honest Feasibility-Based Justification
Not every study is built around a formal power calculation, and reviewers generally accept that — provided the manuscript says so directly instead of obscuring it. Common, legitimate reasons a sample size was set by constraint rather than calculation include:
- Rare-condition or rare-population studies, where the total eligible population is small enough that a target based on a conventional power calculation is not achievable within the study’s timeframe.
- Pilot and feasibility studies, whose purpose is explicitly to inform the design (including the sample size) of a future, adequately powered study — not to detect the effect itself. Common conventions for pilot-study sample sizes (for example, a fixed number per arm sufficient to estimate recruitment and variability parameters) exist precisely because a full power calculation is not the right tool for a study whose primary aim is feasibility, not effect estimation.
- Fixed-resource or fixed-timeline constraints — a defined grant budget, a defined recruitment window, or a defined biobank/registry cohort that already exists at a known size.
- Secondary analyses of existing datasets, where the sample size is simply however many eligible cases exist in the dataset being analyzed.
The writing task in each case is the same: state the actual constraint plainly, explain why it determined the number, and — where relevant — report what the achieved sample was capable of detecting, framed as a limitation rather than a hidden justification. For example: “Given the rarity of [condition] (estimated incidence of X per 100,000), a target sample size was not set by a priori power calculation; instead, all eligible patients presenting at participating sites during the 24-month study period were enrolled. A post hoc sensitivity analysis indicated the achieved sample of 42 patients provided 80% power to detect only a large effect (Cohen’s d ≥ 0.9); the study was therefore not powered to detect smaller, clinically plausible differences, a limitation discussed further below.” This is honest in a way that both STROBE and CONSORT’s guidance explicitly ask for: it does not disguise a constrained sample as a calculated one, and it does not lean on a retrospective power number as though it retroactively validated the design — it uses that number only to characterize what the completed study can and cannot support, and flags it as a limitation.
What to Avoid
- Post hoc power calculations presented as design justification. Calculating power using the effect size actually observed in the completed study is close to circular — a non-significant result will, almost by construction, appear “under-powered” for the effect it happened to observe. STROBE’s own guidance explicitly discourages this practice; where an achieved-power figure is reported at all, it belongs in the limitations discussion, framed honestly, not presented as though it were part of the original design.
- An effect-size assumption with no stated source. “We assumed a medium effect size” without saying where that number came from (a cited prior study, a pilot, a clinically meaningful threshold) is not verifiable and is one of the most common reasons methods reviewers flag a power-calculation summary as incomplete.
- Reverse-engineering the calculation to match a sample already decided on. If the number of participants was set first by budget or recruitment feasibility and a calculation was performed afterward to match it, the manuscript should describe the process that actually happened — a feasibility constraint — not an a priori calculation that did not really drive the number.
- Conflating statistical significance with practical importance in the justification. A sample size chosen only to detect statistical significance at some effect, without regard to whether that effect is clinically or practically meaningful, invites exactly the kind of scrutiny a well-chosen target difference (grounded in a minimal clinically important difference or an established field convention) is meant to preempt.
- Burying the justification. Reviewers look for this in a predictable location — at the end of the participants/design description, immediately before or after the statistical analysis subsection. Scattering pieces of it across the manuscript makes it harder to evaluate and easier to flag as incomplete.
Where This Fits in the Manuscript
The sample-size justification is conventionally placed within the Methods section, either as its own short subsection or as the final paragraph of the participants/design description, immediately before the statistical analysis plan. It should not appear for the first time in the Discussion as a defense against a reviewer’s critique — if a limitation genuinely applies (an achieved sample smaller than the calculated target, for instance), the Methods section should state the original target and the Discussion’s limitations paragraph should acknowledge the shortfall, rather than the sample size being explained only reactively. For guidance on structuring the rest of the Methods section and the manuscript around it, see CASRAI’s How to Write a Research Paper guide; for the section that follows and where effect sizes are actually reported against the sample described here, see Writing the Results Section of a Research Paper.
Frequently Asked Questions
Do I need a power calculation for every study?
No. Formal a priori power calculations are expected for studies designed to test a specific hypothesis with inferential statistics, particularly randomized trials, where CONSORT makes it a checklist item. Descriptive studies, many qualitative studies, pilot/feasibility studies, and secondary analyses of existing datasets do not require one — but they do require the manuscript to say what determined the sample size instead, per STROBE’s item 10 guidance for observational research.
Is it acceptable to report a post hoc power calculation at all?
Only as a way of characterizing what an already-completed study was capable of detecting, framed explicitly as a limitation — not as a justification for the original design. STROBE and most methodologists advise against presenting retrospective power calculated from the observed effect as though it validates the sample size chosen; it does not, because it is calculated from the same data the study collected.
What if my calculation used effect sizes from more than one prior study?
State which one the final calculation used, and briefly explain the choice (e.g., the most methodologically similar prior study, the most conservative/smallest effect among several, or a meta-analytic pooled estimate). If several assumptions were tested, a sensitivity analysis across them can be summarized in one sentence in the Methods section, with the full table in supplementary material.
How is this different from the statistical analysis plan?
The sample-size justification explains how many participants were needed and why; the statistical analysis plan (SAP) specifies how the collected data will be analyzed once gathered — the models, covariates, handling of missing data, and pre-specified subgroup or sensitivity analyses. The two are related (the SAP’s primary analysis is usually the same test the power calculation assumed) but serve different purposes in the manuscript. See CASRAI’s trial design guide for how the two connect in a clinical trial specifically.
This guide covers the writing task of presenting and defending a sample size that has already been determined. It does not provide statistical guidance on how to calculate power or choose an effect size — consult a biostatistician or the CASRAI clinical-research cluster’s trial-design content for that.







