Written and maintained by CASRAI Editorial Board
Last updated
The question that decides everything is not “how many measurements do I have?” but “how many things could independently have been assigned to a different treatment group?” Those two numbers are frequently different, and reporting the first where the analysis requires the second is pseudoreplication. It is the single most common statistical defect in in vivo and cell biology papers, it is almost always invisible in the figure, and it manufactures statistical significance out of nothing.
The one-line test for the experimental unit
ARRIVE 2.0 defines the experimental unit as the biological entity subjected to an intervention independently of all other units, such that it is possible to assign any two experimental units to different treatment groups. That last clause is the operational test. Walk through your protocol and ask, at each level of the hierarchy: could these two have gone into different groups?
- Two mice injected separately with drug or vehicle? Yes — the mouse is the experimental unit.
- Two mice sharing a cage that receives a medicated diet? No — the cage is the experimental unit, however many mice it holds.
- Two pups from a dam dosed during gestation? No — the dam (litter) is the experimental unit; the pups are subsamples.
- Fifty neurons imaged in one mouse? No — ARRIVE 2.0 is explicit that “measurements from 50 individual cells from a single mouse represent n = 1 when the experimental unit is the mouse.”
- Two wells in the same plate, thawed from the same vial on the same day? Usually no. Cells in a well are not independent of one another either: as Lazic and colleagues point out, they “form cell-to-cell connections, release signalling molecules, and compete for the same nutrients.”
The entity you measure is the observational unit. The entity you randomise is the experimental unit. Pseudoreplication is what happens when the analysis silently substitutes the first for the second. This is the animal-research instance of a general problem covered in our guide to the unit of analysis and the mismatch error.
What it actually costs you
The damage is not cosmetic. Because dependent observations carry overlapping rather than unique information, treating them as independent shrinks the standard error below its true value, which inflates the test statistic and drives the p-value down. The effective sample size — the number of genuinely independent information units — stays where it was.
Aarts and colleagues simulated two groups with no real difference between them and analysed the nested data with a conventional t-test. At a nominal alpha of 0.05, the actual type I error rate reached 0.80: four out of five null experiments would have produced a “significant” result. The rate climbs steadily as you add observations per cluster and as the unexplained intracluster correlation (ICC) rises. Simulations in the Journal of General Physiology reached the same conclusion at more modest scale — three animals with five cells each, with equal between- and within-animal standard deviations, produced a 29% false-positive rate, and even two cells per animal pushed it above 10%.
The rodent-litter case is the best quantified. Using real body-weight data, Holson and Pearce showed that three treated and three control litters with two offspring each (12 pups total) gives a 20% false-positive rate instead of 5%; raise it to 12 offspring per litter (72 pups) and the false-positive rate reaches 80%.
The counterintuitive second cost is lost power. Litter-to-litter or cage-to-cage variation that is left unmodelled becomes unexplained noise in the residual, which can mask real effects. Pseudoreplication therefore does not simply trade rigour for sensitivity — it degrades both.
How common is it?
Consistently, across four decades and four disciplines, somewhere between a third and nine-tenths of the affected literature gets it wrong:
- Hurlbert’s original 1984 survey of 176 experimental studies in ecology found pseudoreplication in 27% of them — 48% of the studies that used inferential statistics at all.
- Aarts and colleagues found that at least 53% of 314 papers sampled from Science, Nature, Cell, Nature Neuroscience and Neuron contained nested data; only 7% reported cluster-based summary analyses.
- Lazic, Clarke-Williams and Munafò surveyed animal experiments published 2011–2016 in which an intervention was applied to parents and effects measured in offspring. Only 22% (95% CI 17–29%) replicated the correct entity–intervention pair; 46% were pseudoreplicated; 32% gave too little information to judge.
- In the valproic acid model of autism specifically, Lazic and Essioux found that 3 of 34 studies (9%) correctly treated the litter as the experimental unit. In the same corpus, litter effects accounted for up to 61% of the variation in behavioural outcomes — larger than the treatment effects being reported.
Hurlbert’s three flavours — still the most useful taxonomy
Simple pseudoreplication
One experimental unit per treatment, with multiple samples drawn from it and analysed as replicates. One tank per condition, one cage per diet, one dish per transfection. Treatment effect and unit identity are completely confounded; no statistical test can separate them.
Temporal pseudoreplication
Hurlbert’s definition: this “differs from simple pseudoreplication only in that the multiple samples from each experimental unit are not taken simultaneously but rather sequentially over each of several dates.” Dates are then treated as replicated treatments. Because successive samples from one unit are so obviously correlated, he notes, “the potential for spurious treatment effects is very high with such designs.” Repeated measurement itself is fine; treating the timepoints as independent replicates is not.
Sacrificial pseudoreplication
The subtlest and, in Hurlbert’s survey, only slightly less common than the simple kind. The design does have true replication, but the data are pooled across replicates before analysis, or the several measurements per unit are entered as independent rows. The between-replicate variance exists in the raw data and is then confounded or thrown away — hence “sacrificial.” This is the failure mode most likely to survive peer review, because the methods section can honestly say the experiment was replicated.
A study can have more than one experimental unit
This is where careful teams still trip. In a prenatal-exposure design where dams are dosed and then weaned offspring are individually randomised to a therapeutic compound or placebo, the dam is the experimental unit for the prenatal factor and the individual animal is the experimental unit for the postnatal factor. That is a split-plot design with two error strata, and it needs two different denominators. Analysing the whole thing with a single one-way ANOVA at the animal level will be anticonservative for the prenatal comparison. If your protocol has any treatment applied at a level above the individual animal — room, cage, rack, litter, tank, batch, surgery day — assume a split-plot structure until you have proved otherwise. See our guide to design of experiments for the general blocking and factorial vocabulary.
Three defensible ways to analyse it
1. One observation per experimental unit
Randomly select a single offspring per litter, or a single field per animal, and use standard methods. Clean, unimpeachable, and wasteful of the data you already collected. Useful mainly as a design decision, not a rescue.
2. Summarise to the experimental unit, then test
Average the cells within each animal, or the pups within each litter, and run the test on those summaries with n equal to the number of units. Lazic and colleagues note the summary need not be a mean — “other numeric summaries such as the slope or area under the curve may capture a feature of interest better.” Hurlbert reached the same conclusion in 1984, arguing that the honest move is to analyse at the level of the experimental unit and skip formal analysis of the subsamples, because fancier nested analyses “will not be any more powerful in detecting treatment effects, but will be more susceptible to calculation and interpretation error.” The cost: summary-statistic analysis loses power relative to a multilevel model, and the loss is largest when the number of clusters is small.
3. Model the hierarchy explicitly
Fit a multilevel or mixed-effects model with the clustering variable (animal, litter, cage, experimental run) as a random effect. This uses the effective rather than the observed sample size, protects the type I error rate, and gives you an estimate of how large the litter or cage effect actually is — information the summary approach discards. It also handles unbalanced designs, which averaging handles badly. Our guide to mixed-effects models covers what the random part does and how to specify it. In practice, when the design is balanced, summarising and modelling will often give the same answer.
What does not work: reporting “n = 300 cells from 3 independent experiments” and testing on 300. Lord and colleagues make the correction concrete — in their SuperPlot examples, “P values were calculated using an n of three, not 300.” Their recommended figure superimposes the per-replicate means, colour-coded by biological replicate, on the full cell-level swarm, so the reader can see both the cell-to-cell spread and whether the effect actually recurred across runs.
The one legitimate defence: check your scope of inference
Not every set of correlated measurements is pseudoreplication. Jordan’s response in PLOS Biology makes the point that whether a study is pseudoreplicated depends on its scope — each measurement on an individual is independent if the experiment aims to understand treatment effects for that individual only, because the measurements are then samples from the population of possible responses of that one animal.
This is a narrow defence, and it must be declared in advance rather than invoked after a reviewer objects. If your abstract generalises to the strain, the model or the disease, you have inferred beyond the individual, and the individual-level defence is unavailable. Deciding the target population before the analysis — not after — is the whole discipline here.
Where regulators and reporting standards have already ruled
For developmental studies the question is settled in guideline text, not just in the methods literature. OECD Test Guideline 426 (Developmental Neurotoxicity Study, adopted 2007) states at paragraph 49 that developmental studies using multiparous species where multiple pups per litter are tested “should include the litter in the statistical model to guard against an inflated Type I error rates. The statistical unit of measure should be the litter and not the pup. Experiments should be designed such that littermates are not treated as independent observations.” It adds that any endpoint repeatedly measured in the same subject must be analysed with models that account for the non-independence of those measures. The guideline’s recommended group size — approximately 20 litters per dose level — is a count of litters precisely because the litter is the unit.
ARRIVE 2.0 places the experimental unit in the Essential 10, the minimum set of items that must appear in any paper reporting animal research: authors should clearly indicate the experimental unit for each experiment so that sample sizes and statistical analyses can be properly evaluated. Conflating experimental units with subsamples, ARRIVE notes, “underestimates the true variability in a study, which can lead to false positives and invalidate the analysis and resulting conclusions.”
The compliance consequence is practical. A power calculation performed at the wrong level produces the wrong animal number, and the animal number is the thing your ethics committee approves. If your justification counts pups but your valid analysis counts dams, either the study is underpowered for its stated aim or you have requested animals that cannot answer the question — a Reduction problem, not merely a statistics problem. Settle the experimental unit before you write the IACUC protocol, not after the data are in. The NC3Rs Experimental Design Assistant is a free tool built for exactly this stage, providing structured support for randomisation, blinding and sample size calculation plus feedback on the experimental plan.
Red flags in a methods section
- Two different n’s in the same sentence — “n = 45 neurons from 3 animals” followed by a test on 45.
- Vanishingly small error bars on a biological measurement, with a p-value several orders of magnitude below 0.05. As Lord and colleagues put it, if your p-value looks too good to be true, it probably is.
- An intervention delivered through the environment — diet, drinking water, bedding, ambient temperature, gas anaesthesia in a shared chamber — with n reported as animals.
- “Three independent experiments” with cell-level statistics, and no statement of how the replicates were made independent (separate passages? separate thaws? separate days?).
- Repeated measures analysed as independent rows, or timepoints treated as replicates.
- No statement of the experimental unit at all — the modal case in every survey above, and the reason a third of studies could not even be classified.
Frequently asked questions
Is n the number of animals or the number of cells?
Neither by default. It is the number of entities that were independently assigned to a treatment. If you dosed each mouse separately and then imaged 50 cells in each, n is the number of mice. If you dosed a cage of five mice through the drinking water, n is the number of cages. Count the randomisations, not the measurements.
Does averaging my subsamples throw away data?
It discards the within-unit variability from the test, but that variability was never evidence about the treatment. What averaging does cost you is power relative to a mixed model, and that penalty grows as the number of clusters shrinks. With a balanced design and a decent number of units, summarising and modelling typically agree; with few clusters or unbalanced subsample counts, fit the model.
Is a technical replicate ever an experimental unit?
No. Technical replicates — re-reading the same sample, duplicate wells from one lysate, two sections from one block — estimate measurement precision. They tell you how reliably you measured one unit, not how the treatment behaves across units. Averaging them before analysis is the right move; counting them in n is not.
Are littermates always pseudoreplication?
Only when the treatment was applied at or above the litter level. If dams were dosed in pregnancy, the litter is the unit and pups are subsamples. If weaned littermates were individually randomised to drug or vehicle, the animal is the unit for that comparison — and using littermates is actively good design, because it blocks out between-litter variation. The trap is the study that does both and analyses it as one flat comparison.
Can I fix pseudoreplication after the data are collected?
Sometimes. If you recorded which animal, litter, cage or run each measurement came from, you can reanalyse correctly — the information is there. If you pooled before recording, or ran a single cage per treatment, no analysis can recover the treatment replication that was never built in. This is why the identifiers matter: record the cluster membership of every observation as a matter of routine, even when you expect not to need it.
Does this apply to cell culture, or only animals?
It applies to any hierarchy. Cells within wells, wells within plates, plates within runs, runs within passages, and passages within a cell line all generate dependence. The question is always the same: at what level was the treatment independently applied, and how many such units are there? A single cell line handled on one day is n = 1 for inference about the biology, regardless of how many wells were plated.
What ICC should I assume when planning?
There is no universal value — it is outcome- and model-specific, and depends on the relative variability between and within clusters. Lazic and Essioux found litter effects explaining up to 61% of variation in some behavioural outcomes but well under half in others. Estimate it from your own pilot or historical data for the specific endpoint; borrowing a number from a different assay is a guess, not a plan. What is general is the direction: adding more observations per unit buys little once dependence is appreciable, whereas adding more independent units buys real precision.
References
- Percie du Sert N, et al. ARRIVE 2.0 Item 1b: Study design — experimental unit (explanation). ARRIVE Guidelines.
- Percie du Sert N, et al. Reporting animal research: Explanation and elaboration for the ARRIVE guidelines 2.0. PLOS Biology 2020;18(7):e3000411.
- Hurlbert SH. Pseudoreplication and the design of ecological field experiments (PDF). Ecological Monographs 1984;54(2):187–211.
- Lazic SE, Clarke-Williams CJ, Munafò MR. What exactly is ‘N’ in cell culture and animal experiments? PLOS Biology 2018;16(4):e2005282.
- Lazic SE, Essioux L. Improving basic and translational science by accounting for litter-to-litter variation in animal models. BMC Neuroscience 2013;14:37.
- Aarts E, Verhage M, Veenvliet JV, Dolan CV, van der Sluis S. A solution to dependency: using multilevel analysis to accommodate nested data. Nature Neuroscience 2014;17:491–496.
- Eisner DA. Pseudoreplication in physiology: More means less. Journal of General Physiology 2021;153(2):e202012826.
- Lord SJ, Velle KB, Mullins RD, Fritz-Laylin LK. SuperPlots: Communicating reproducibility and variability in cell biology. Journal of Cell Biology 2020;219(6):e202001064.
- Jordan CY. Population sampling affects pseudoreplication. PLOS Biology 2018;16(10):e2007054.
- OECD. Test No. 426: Developmental Neurotoxicity Study. OECD Guidelines for the Testing of Chemicals, adopted 16 October 2007 (see paragraphs 11, 14 and 49).
- NC3Rs. Experimental Design Assistant (EDA).
- Lazic SE. Genuine replication and pseudoreplication. Nature Reviews Methods Primers 2022;2:114.








