The replication crisis is the recognition, widely dated to the early 2010s, that a substantial share of published findings across psychology, biomedicine, and other empirical fields fail to hold up when other researchers try to reproduce or replicate them under similar conditions. It is not a single event but a cumulative body of evidence — probabilistic arguments, large coordinated replication projects, and internal field surveys — that converged on the same conclusion from different directions. This guide covers where the crisis came from, how it has played out across disciplines, what causes have been identified, and how research data management and open science practice have responded. For CASRAI’s audience specifically, the throughline is that a large share of the practical response lives in research data management: whether the data, code, and materials behind a result are documented and shared well enough for anyone to check it at all.
Reproducibility, replicability, and “the replication crisis”: three different things
The phrase “replication crisis” is often used loosely to cover a cluster of related but distinct concepts, and the imprecision causes real confusion. The CASRAI Dictionary’s Reproducibility entry follows the National Academies of Sciences, Engineering, and Medicine’s (NASEM) 2019 report Reproducibility and Replicability in Science in drawing a specific line:
- Reproducibility means obtaining consistent results using the original data and the original code or methods — essentially, can someone else run your analysis and get your numbers. It is a property of the study’s record and workflow, not of the underlying scientific claim.
- Replicability (sometimes called replication) means obtaining consistent results in a new study that collects new data addressing the same question. A study can be perfectly reproducible — the original numbers check out exactly — and still fail to replicate, if the underlying effect does not re-emerge with a fresh sample.
- The replication crisis (also referred to as the reproducibility crisis, the two terms are used interchangeably in most of the literature despite the technical distinction above) refers specifically to the historical episode: the accumulation, from roughly 2005 onward, of evidence that replication failure was happening at a rate high enough to be a structural problem for how several fields produce and validate knowledge, rather than isolated, expected scientific disagreement.
This guide uses “replication crisis” in that historical sense throughout, consistent with how the CASRAI Dictionary’s Reproducibility crisis term defines it.
Origins: how the crisis surfaced, field by field
2005 — the probabilistic argument
The commonly cited starting point is John Ioannidis’s 2005 paper in PLOS Medicine, “Why Most Published Research Findings Are False.” Ioannidis did not run an empirical audit of existing papers; instead he built a probabilistic model showing that under realistic conditions — small sample sizes, small true effect sizes, a large number of tested relationships, flexibility in study design and analysis, and fields with strong financial or other interests in a positive result — the majority of published “positive” findings in a research area could be expected to be false, purely as a consequence of how significance testing and publication incentives interact. The paper was aimed primarily at biomedical research but its argument generalized, and it became the theoretical foundation the rest of the crisis narrative built on.
2011 — psychology’s early warning signs
Psychology’s reckoning arrived in a concentrated burst in 2011. Daryl Bem published a paper in a top social-psychology journal reporting statistically significant evidence for precognition, which many researchers treated as a reductio ad absurdum: if standard methods could produce “significant” support for an effect most scientists considered implausible, something was wrong with the methods, not just with one study. The same year, Joseph Simmons, Leif Nelson, and Uri Simonsohn published “False-Positive Psychology,” demonstrating that ordinary, common analytic flexibility — choosing among several plausible ways to collect, exclude, or analyze data after seeing the data — could push a true false-positive rate as high as 60% under a nominal 5% significance threshold. This behavior became known as p-hacking, and the closely related practice of presenting a post-hoc, data-driven finding as if it had been the study’s original hypothesis became known as HARKing (Hypothesizing After the Results are Known). Later the same year, a large-scale fraud case (Diederik Stapel, Tilburg University) came to light, in which dozens of published papers were found to rest on fabricated data — a separate problem from p-hacking, but one that added to the same sense that psychology’s normal quality-control mechanisms weren’t catching serious problems.
2015 — the Reproducibility Project: Psychology
The single most-cited data point in the replication crisis came in August 2015, when the Center for Open Science’s Open Science Collaboration — roughly 270 contributing researchers — published the results of a coordinated attempt to replicate 100 studies from three leading psychology journals (Open Science Collaboration, “Estimating the Reproducibility of Psychological Science,” Science, 2015, DOI: 10.1126/science.aac4716). Of the original studies, 97% had reported a statistically significant effect; only 36% of the replications did, and the average replicated effect size was roughly half the original. CASRAI covers this specific study, why psychology was especially exposed, and the reforms that followed directly from it, in a dedicated guide: The Psychology Replication Crisis: The 2015 Reproducibility Project and What Changed. This page does not repeat that detail; it situates the 2015 result within the broader, cross-disciplinary pattern.
Beyond psychology: cancer biology, economics, and other fields
The pattern was not unique to psychology. The Center for Open Science ran a parallel, multi-year effort, the Reproducibility Project: Cancer Biology, attempting to replicate experiments from 53 high-impact preclinical cancer papers published 2010–2012. Documentation gaps, missing reagents and protocols, and other practical barriers meant only 50 experiments across 23 of the original papers could actually be attempted; where replications were completed, effect sizes were on average substantially smaller than the originals, and positive original findings replicated at a markedly lower rate than original null findings (Errington et al., eLife, December 2021). Similar replication and reanalysis efforts have surfaced comparable concerns in experimental economics, preclinical drug-target research more broadly, and parts of biomedicine — enough independent, discipline-specific evidence that by the mid-2010s “replication crisis” had become a cross-field term rather than a psychology-specific one. A 2016 Nature survey of over 1,500 scientists (Baker, “1,500 scientists lift the lid on reproducibility”) found that more than 70% had tried and failed to reproduce another scientist’s results at least once, and researchers across chemistry, biology, physics, engineering, medicine, and earth science all reported experiencing this — evidence that the underlying problem was structural to how incentives and reporting norms worked, not confined to one field’s methods.
What caused it
The literature that grew out of these findings converges on a fairly consistent set of contributing causes, none of which requires assuming individual researchers acted in bad faith:
- Publication bias. Journals have historically favored novel, statistically significant, “clean” results over null findings or replications, which means the published literature is a selectively filtered, optimistic sample of all the research actually conducted. See CASRAI’s Publication Bias term.
- Researcher degrees of freedom / p-hacking. Without a pre-specified analysis plan, ordinary and often unconscious flexibility in how data is collected, cleaned, and analyzed can turn a null result into a “significant” one. See P-hacking.
- HARKing and other questionable research practices. Presenting an exploratory, data-driven finding as though it had been predicted in advance removes the statistical protection that hypothesis testing is supposed to provide. See Questionable Research Practices.
- Low statistical power. Many published studies, particularly in psychology and preclinical biomedicine, have historically been underpowered to detect the effect sizes they claim, which inflates both the false-positive rate and the exaggeration of effect sizes among studies that do reach significance.
- Incentive structures. Hiring, tenure, and funding decisions have historically rewarded publication volume and novel positive results far more than rigor, replication, or negative findings — the “publish or perish” dynamic that made the practices above rational responses to career pressure rather than aberrant behavior.
- Insufficient data, code, and materials sharing. A result that cannot be checked because the underlying data, analysis code, or study materials were never made available cannot be reproduced even in principle, regardless of whether the original finding was correct. This is the point at which the replication crisis becomes a research data management problem, not only a statistics or incentives problem.
“Crisis” or “credibility revolution”?
Not every methodologist accepts “crisis” as the right framing. An alternative, increasingly common description — sometimes called the “credibility revolution” — treats the same body of evidence as a sign the scientific community’s internal quality-control mechanisms were working roughly as intended: problems were identified, named, studied empirically, and are now being addressed through concrete methodological reform, rather than covered up or ignored. Both framings agree on the underlying facts; they differ on emphasis, and either framing motivates the same practical reforms described below. CASRAI’s Dictionary entry for the reproducibility crisis notes this framing debate directly.
The open science response
The methodological reform movement that grew out of the replication crisis is now generally referred to as open science, and it targets the causes above fairly directly:
- Preregistration. Publicly time-stamping a study’s hypotheses, design, and analysis plan before data collection begins closes off the p-hacking and HARKing pathways after the fact. See Pre-registration and CASRAI’s guide to preregistering a study protocol.
- Registered Reports. A publishing format in which a journal peer-reviews and conditionally accepts a study’s introduction, methods, and analysis plan before results are known, removing outcome-based publication bias entirely. See Registered Report.
- Open data and materials. Making underlying data, code, and study materials available lets independent researchers check reproducibility directly, and is a core commitment of the FAIR principles CASRAI covers throughout its research data management content.
- Replication studies as a legitimate output. Journals and funders increasingly treat well-designed replication attempts, including crowdsourced replication efforts, as publishable, citable contributions rather than a lesser form of research. See Replication study and Replicability.
- Funder and journal policy. The NIH’s rigor-and-reproducibility expectations, journal-level reporting checklists, and infrastructure such as the Center for Open Science’s Open Science Framework have institutionalized many of these practices rather than leaving them to individual researcher discretion.
The research data management angle
Read against CASRAI’s own remit, the replication crisis is, in large part, a data stewardship failure mode: results that cannot be checked because the data, code, or provenance behind them were never adequately documented or preserved. This is why reproducibility sits inside CASRAI’s RDM cluster as its own sub-area rather than only as a publishing-ethics topic. Two distinctions matter here for research administrators and data stewards specifically:
- Computational reproducibility is an infrastructure and documentation problem, not just a statistics problem. Getting the same numbers out of the same data requires that the code, software environment, and workflow used to produce a result are captured well enough to be re-run years later — the concern the ACM’s Artifact Review and Badging framework and provenance standards like PROV-O were built to address.
- A data management plan that specifies where data and code will live, and under what license, is a direct, practical hedge against irreproducibility — not an abstract compliance requirement. A result backed by a dataset in a certified, persistent repository with clear documentation is verifiable in a way one backed by “data available on request” is not.
For a deeper treatment of this distinction and the AI-era data-provenance questions layered on top of it, see CASRAI’s Research Data Management pillar and the Reproducibility and Replicability dictionary entries.
Frequently asked questions
Is the replication crisis over?
No single study or policy change ended it, and replication rates in the affected fields have not been comprehensively re-measured on the scale of the 2015 psychology project to confirm a clear before-and-after improvement. What has changed is infrastructure and norms: preregistration, Registered Reports, open-data requirements, and funder rigor policies are now mainstream practice in the fields most affected, rather than fringe proposals, which is generally treated as the practical measure of progress rather than a single “resolved” milestone.
Which fields are most affected?
Psychology (particularly social psychology) and preclinical biomedicine, including cancer biology, have the most extensively documented replication failures, largely because they were also the first fields to run large, coordinated replication projects. Economics, other social sciences, and parts of the biomedical and life sciences more broadly have surfaced comparable concerns through smaller-scale replication efforts and reanalyses. The 2016 Nature survey found researchers across essentially every discipline reporting failed replication attempts, which suggests the underlying incentive and reporting problems are not confined to any one field, even where the empirical documentation is uneven across fields.
What is the difference between the replication crisis and p-hacking?
P-hacking is one specific practice — exploiting analytic flexibility to manufacture a statistically significant result — identified as one of several contributing causes of the replication crisis. The replication crisis is the broader, field-level pattern of low replication rates; p-hacking is one mechanism, among several (publication bias, low power, HARKing, insufficient data sharing), that helps explain why that pattern exists.
Did the replication crisis start with the 2015 psychology study?
No. The 2015 Reproducibility Project: Psychology is the most-cited single data point, but the underlying argument dates at least to Ioannidis’s 2005 probabilistic paper, and the concrete warning signs in psychology (the Bem precognition paper, the Simmons/Nelson/Simonsohn p-hacking demonstration, and the Stapel fraud case) all surfaced in 2011, four years before the OSC study was published.
Related CASRAI terms and guides
- The Psychology Replication Crisis: The 2015 Reproducibility Project and What Changed
- Reproducibility crisis
- Reproducibility
- Replicability
- Replication study
- Crowdsourced replication
- P-hacking
- Questionable research practices
- Publication Bias
- Pre-registration
- Registered Report
- Preregistration of a Study Protocol
- Open Science Framework (OSF)
- Research Data Management pillar







