Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

The Psychology Replication Crisis: The 2015 Reproducibility Project and What Changed

In 2015, a 270-author collaboration attempted to replicate 100 published psychology studies and found only 36% held up. This guide covers what the Reproducibility Project: Psychology actually found, why psychology was especially exposed, and the preregistration, Registered Report, and open-data reforms that followed.

In August 2015, the journal Science published the results of the largest coordinated replication effort psychology had ever attempted. The Open Science Collaboration, a consortium of roughly 270 researchers coordinated through the Center for Open Science, had spent several years trying to reproduce the results of 100 studies published in three leading psychology journals. The headline finding — that only about 36% of the replications produced a statistically significant effect, down from 97% in the original studies — became the single most cited data point in what researchers now call the reproducibility crisis. This guide covers what the study actually measured, why psychology specifically was so exposed, and the methodological reforms — preregistration, Registered Reports, and open-data policy — that followed directly from it.

What the Reproducibility Project: Psychology Actually Did

The study, formally titled “Estimating the Reproducibility of Psychological Science” (Open Science Collaboration, Science, 2015, DOI: 10.1126/science.aac4716), sampled 100 experimental and correlational studies published in 2008 in three journals: Psychological Science, the Journal of Personality and Social Psychology, and the Journal of Experimental Psychology: Learning, Memory, and Cognition. For each study, a separate team obtained the original materials wherever possible and ran a new, high-powered replication attempt using the original design.

The results, by several converging measures:

  • 97% of the original studies reported a statistically significant result (p < .05); only 36% of the replications did.
  • The mean effect size in the replications was roughly half the magnitude of the original reported effect.
  • Subjective ratings by the replicating teams of whether a finding “replicated” produced a similar figure, around 39%.
  • Replication success varied by original journal and subfield — cognitive psychology studies replicated at a notably higher rate than social psychology studies.

The paper’s authors were explicit that a failed statistical replication does not, on its own, prove the original finding was false: a single non-significant replication is also just one data point, and effect sizes can vary for reasons unrelated to whether the underlying phenomenon is real. What made the result significant wasn’t any single replication outcome but the aggregate pattern across 100 independent attempts, which was far too consistent to be explained by chance alone.

Why Psychology Specifically Was So Exposed

The 2015 result didn’t happen because psychologists were unusually careless. It surfaced structural incentives and methodological habits that were common across the field and, at the time, largely unremarked upon.

Small samples and low statistical power

Many of the original studies used sample sizes designed to detect large effects cheaply, not to reliably detect the more modest effects typical of real psychological phenomena. A study that is underpowered to detect a true effect will, by definition, fail to replicate that effect at a high rate even when the effect is genuine — and when an underpowered study does return a significant result, that result is more likely to be an inflated overestimate of the true effect size, which is itself consistent with the roughly 50% shrinkage the OSC team observed.

Researcher degrees of freedom and p-hacking

Simmons, Nelson, and Simonsohn’s influential 2011 paper “False-Positive Psychology” demonstrated how ordinary, undisclosed analytic flexibility — which covariates to include, when to stop collecting data, which outcome measure to report — could push the false-positive rate for a null effect from the nominal 5% up past 60%. This class of practice, now catalogued under terms like researcher degrees of freedom, forking paths, and p-hacking, was rarely intended as fraud — it was simply how exploratory data analysis was normally practiced before it had a name.

HARKing

HARKing (Hypothesizing After Results are Known) describes presenting a post-hoc, data-driven finding in a paper as if it had been predicted in advance. Because it collapses the distinction between confirmatory and exploratory analysis, HARKing makes an exploratory finding look far more statistically robust than it actually was — and therefore far less likely to hold up under an independent, genuinely pre-specified replication attempt.

Publication bias and the file-drawer problem

Journals in psychology, like most fields, historically favored novel, statistically significant, positive results over null findings. That publication bias means the published literature is a filtered, upward-biased sample of the studies actually run — the “file drawer” of unpublished null results was never part of the record researchers were building on, which inflated the apparent reliability of the field’s published effects relative to its true underlying replication rate.

How the Field Responded

The 2015 result didn’t just document a problem; it catalyzed a set of methodological reforms that are now standard expectations in large parts of psychology and have since spread to other empirical social and biomedical sciences.

Preregistration and the Open Science Framework

Pre-registration — publicly time-stamping a study’s hypotheses, design, and planned analysis before data collection begins — directly targets p-hacking and HARKing by making the confirmatory/exploratory distinction verifiable after the fact. The Center for Open Science’s OSF Registries became the dominant infrastructure for this in psychology; see CASRAI’s guide to preregistration for how the process works in practice.

Registered Reports

The Registered Report format goes a step further than plain preregistration: the study’s introduction, hypotheses, and planned analysis are peer-reviewed and provisionally accepted for publication before the results are known, removing the editorial incentive to reward significant results over null ones. First proposed for psychology around 2012-2013 (Chambers’ 2013 Cortex paper is the commonly cited founding description), Registered Reports are now offered by several hundred journals across multiple fields.

Open Science Badges and journal policy change

Several journals began publicly marking articles whose data, materials, or preregistrations were made openly available, most visibly Psychological Science, which introduced Open Science Badges in 2014. The intent was to make existing open practices visible and reward them reputationally, rather than requiring a heavier editorial mandate.

The TOP Guidelines

The Center for Open Science’s Transparency and Openness Promotion (TOP) Guidelines, first published in 2015, gave journals and funders a modular framework — covering citation standards, data and code transparency, preregistration, and replication, each adoptable at increasing levels of stringency — for formalizing exactly which of these open-science practices they require, rather than leaving adoption to individual editors’ discretion.

Has Anything Actually Changed Since 2015?

Direct before/after replication-rate comparisons are hard to run cleanly, but several downstream indicators point the same direction: preregistration rates in major psychology journals rose sharply through the late 2010s, sample sizes in newly published studies have generally increased, and open-data/open-materials sharing is now common in journals that adopted badges or TOP-aligned policies. The debate over whether “crisis” or “credibility revolution” is the more accurate framing for this period is itself ongoing in the field — but both framings point to the same set of concrete reforms. It’s also worth being precise about terminology here: strictly, the Reproducibility Project used new analyses of newly collected data under the original design, which the field now more often calls a replication study rather than reproducibility in the narrower computational-reproducibility sense (rerunning the same analysis code on the same original data). For a broader look at how psychology’s experience fits into the empirical study of scientific practice generally, see CASRAI’s metascience entry.

Frequently Asked Questions

What percentage of psychology studies failed to replicate in the 2015 study?

Of the 100 studies attempted, about 36% of replications reached statistical significance (p < .05), compared with 97% of the original studies. A subjective replication-success rating by the research teams produced a similar figure, around 39%.

Was the Reproducibility Project: Psychology only about psychology?

Yes — it specifically sampled studies from three psychology journals published in 2008. Similar coordinated replication efforts have since been run in other fields (for example, cancer biology and experimental economics), with broadly comparable findings of lower-than-expected replication rates, but the 2015 Science paper itself covered psychology only.

Is a failed replication proof the original finding was wrong?

Not on its own — a single non-significant replication is one additional data point, not a refutation. The 2015 result was notable because of the consistency of the pattern across 100 independent attempts, not because any one replication conclusively disproved its original study.

What is the difference between reproducibility and replicability?

Reproducibility generally refers to obtaining the same results by re-running the original analysis on the original data; replicability refers to obtaining a consistent finding in a new study, with new data, testing the same question. The 2015 project tested replicability in this stricter sense, even though its title uses “reproducibility” — a terminology inconsistency that is common in the literature and worth watching for when reading related sources.

Did the replication crisis extend beyond psychology?

Yes. Comparable, if less extensively publicized, concerns have been documented in preclinical cancer biology, economics, and other fields, and the term reproducibility crisis is now used broadly across the biomedical and social sciences rather than as a psychology-specific label.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →