“Replication” is easy to define in the abstract and hard to picture concretely. Psychology is the field where the concept has been tested most publicly, largely because psychology ran the largest coordinated replication effort of any discipline and published its results openly. This guide walks through two real, well-documented psychology replication examples in enough detail to show what a replication study actually looks like in practice, what “success” and “failure” mean in this context, and what a research administrator or RCR trainer can take from them.
For the formal definition and criteria, see CASRAI’s replication study entry. This guide assumes that definition and focuses entirely on worked examples.
Example 1: The Reproducibility Project: Psychology (a large-scale, multi-study replication)
The single most-cited psychology replication example is the Reproducibility Project: Psychology (RP:P), coordinated by the Center for Open Science and published as Open Science Collaboration, “Estimating the Reproducibility of Psychological Science,” Science 349(6251), 2015.
- What they did. A large team of researchers selected 100 studies published in three well-regarded psychology journals (Psychological Science, Journal of Personality and Social Psychology, and Journal of Experimental Psychology: Learning, Memory, and Cognition) and attempted a direct replication of each one, using methods and materials as close to the original as they could obtain, often with input from the original authors.
- What they found. 97% of the original studies had reported a statistically significant result. Among the replications, only 36% reached statistical significance. Where an effect did replicate, the replicated effect size was, on average, about half the magnitude of the originally reported effect.
- Why this counts as a replication example, not just a statistic. Each of the 100 individual attempts is itself a replication study in the CASRAI sense: a study whose primary purpose was to re-test a specific prior claim using closely matched methods. RP:P is instructive precisely because it aggregates 100 individual replication studies rather than being one itself, which is why it is cited as evidence about the reliability of a field rather than about any single finding.
RP:P is widely credited with turning “the replication crisis” from an informal concern into a documented, field-wide finding, and it remains the reference point most institutional research-integrity and rigor training points to. See CASRAI’s reproducibility crisis entry for how the term is used more broadly across disciplines.
Example 2: The facial feedback hypothesis Registered Replication Report (a single, focused replication)
Where RP:P shows replication at scale, the facial feedback hypothesis Registered Replication Report shows what one specific, well-designed replication study looks like end to end.
- The original finding. Strack, Martin, & Stepper (1988) asked participants to hold a pen in their mouth in one of two ways: gripped with the teeth (which incidentally activates smiling muscles) or held with the lips (which suppresses them), while rating the funniness of cartoons. Participants in the teeth condition rated the cartoons as funnier, a 0.82-point difference on a 10-point scale — offered as evidence that facial muscle activity itself can influence emotional experience, not just express it.
- The replication. In 2016, a Registered Report-format Registered Replication Report (Wagenmakers et al., “Registered Replication Report: Strack, Martin, & Stepper (1988),” Perspectives on Psychological Science) coordinated 17 independent laboratories, all running the same preregistered protocol on more than 1,800 participants combined, with video recording added to confirm participants actually held the pen correctly.
- The result. The meta-analytic effect across the 17 labs was 0.03 points, with a confidence interval spanning zero — essentially no detectable effect, in sharp contrast to the original 0.82-point difference.
- Why it’s a useful contrast to RP:P. This is a single, transparent replication of one specific claim, run under a preregistered, peer-reviewed protocol before data collection began, which is precisely the Registered Report format CASRAI documents separately. It shows a replication study can be null even when it is large, well-designed, multi-site, and methodologically stronger than the original — a null replication result is itself a publishable, informative finding, not simply “no result.”
What these examples illustrate about replication in psychology specifically
- A replication can fail to confirm an effect without anyone having done anything wrong. Both the original and replicating researchers in the facial feedback case followed accepted practice for their eras; the gap reflects, among other things, changes in preregistration norms, sample sizes, and awareness of publication bias, not misconduct on either side.
- Direct replication and conceptual replication are different tools. RP:P and the facial feedback RRR are both direct replications — they tried to match the original method as closely as possible. A conceptual replication instead tests the same underlying hypothesis with a deliberately different method, which is a different (and separately useful) way of building confidence in a claim; see the replication study entry for how CASRAI distinguishes the two.
- Scale changes what a replication can tell you. A single well-powered Registered Replication Report can rule out a specific effect with high confidence; a 100-study program like RP:P is needed to say something about a field’s overall reliability rather than about any one finding.
- Preregistration is what makes a replication’s result credible either way. Both examples used a protocol fixed in advance, which is what allows a null result to be reported as a real finding rather than dismissed as a failed experiment. See CASRAI’s pre-registration entry and the Center for Open Science / OSF guide for the infrastructure most preregistered replications actually run on.
Using these examples in RCR or rigor training
Research administrators building responsible-conduct-of-research or rigor-and-reproducibility training material often need one clear, verifiable illustration rather than an abstract statistic. RP:P and the facial feedback RRR both work well for this because they are fully public, well-documented, and not attributed to any institution CASRAI would need to name — the citations above point directly to the primary sources, so training material can cite them without relying on a secondhand summary. Because both are published, peer-reviewed studies, they can be referenced by name in policy documents and slide decks without the sourcing concerns that come with anecdotal “cases.”
Frequently Asked Questions
What is a real example of a replication study in psychology?
The Reproducibility Project: Psychology (Open Science Collaboration, 2015) is the most-cited example: 100 published psychology studies were independently re-run, and only 36% of the replications reached statistical significance versus 97% of the originals. A more focused single example is the 2016 Registered Replication Report of Strack, Martin, & Stepper’s (1988) facial feedback study, in which 17 labs found essentially no effect where the original had reported a substantial one.
What percentage of psychology studies replicate?
There is no single field-wide figure, but the most widely cited data point is from the Reproducibility Project: Psychology, where 36% of 100 replication attempts reached statistical significance, compared with 97% of the corresponding original studies. Replication rates vary by subfield, effect size, and study design, so this figure should be treated as one influential estimate rather than a fixed constant.
Is a failed replication the same as research misconduct?
No. A replication that does not reproduce an original effect is a normal, expected outcome of doing science under uncertainty, and is not evidence of misconduct on its own. Both examples in this guide involved standard, good-faith methods on both sides; misconduct is a separate, much narrower category involving fabrication, falsification, or plagiarism. See CASRAI’s reproducibility crisis and research-misconduct content for that distinction.
What is the difference between a direct and a conceptual replication?
A direct replication tries to match the original study’s methods and materials as closely as possible, as both examples in this guide did. A conceptual replication instead tests the same underlying hypothesis using deliberately different methods or measures. Both are legitimate replication study types; see CASRAI’s replication study entry for the full definition.







