Skip to main content
v2026.11,610 entries · CC-BY 4.0

Natural Experiments: How to Argue a Shock Is Genuinely Exogenous

What makes a natural experiment’s shock genuinely exogenous rather than correlated with unobserved confounders, the pre-trend and placebo checks that establish credibility, and named real-world examples.

Ask about Natural Experiments: How to Argue a Shock Is Genuinely Exogenous

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

A natural experiment relies on a stroke of luck that most quasi-experimental designs have to engineer: something outside anyone’s control — a policy change with an arbitrary boundary, a lottery, a storm, a company’s back-office decision — splits a population into a treated group and a comparison group as if by accident. When that split really is accidental with respect to the outcome you’re studying, it does the same job random assignment does in a trial. When it isn’t — when the “shock” is actually correlated with something else driving the outcome — the whole design collapses back into an ordinary confounded comparison wearing a natural experiment’s clothing. Reviewers know this, which is why “is this shock actually exogenous?” is the first question a natural-experiment paper has to survive, not an afterthought.

This guide is about answering that question credibly: what “exogenous” means in this context, the specific checks — pre-trend testing and placebo checks chief among them — that let you argue for it rather than assert it, and how natural experiments relate to the other quasi-experimental designs that share the same underlying logic.

What “exogenous” actually means here

A shock is exogenous with respect to your outcome when its timing, magnitude, and who it hits are unrelated to the unobserved factors that also drive the outcome. The threat is never “does the shock affect the outcome” — that’s the whole point of studying it. The threat is that the shock is correlated with something else that affects the outcome, so what looks like the shock’s effect is partly or wholly that something else.

Two failure patterns account for most rejected natural-experiment claims:

  • Selection into exposure. The “natural” event isn’t randomly distributed across units — people, firms, or regions with certain unobserved traits are systematically more or less likely to be exposed. A state doesn’t raise its minimum wage at random; it raises it when its economy and politics look a certain way, and those same conditions can independently affect employment trends.
  • Anticipation and sorting around the shock. Even a genuinely exogenous event can be contaminated if agents see it coming and adjust behavior beforehand — firms stockpile inventory ahead of an announced tariff, workers change jobs ahead of an announced plant closure. The pre-shock period stops being a clean baseline.

Neither failure means the shock is useless. It means exogeneity is a claim you have to defend with evidence about the shock’s origin and with data-based checks, not a property you get automatically because the word “natural” is in the design’s name. See endogeneity: the three sources, and the remedy that matches each for how this same problem shows up across observational designs generally.

The credibility checklist

These are the checks a methods reviewer runs through, roughly in the order they’ll raise them.

1. Trace the shock to a source outside the units under study

State plainly where the shock came from and argue why that source couldn’t have been responding to the outcome you’re measuring. A state legislature’s minimum-wage vote is not exogenous to state-level labor-market conditions in general, but a specific vote’s timing relative to a particular firm’s hiring decisions can still be argued as outside that firm’s control. A water utility’s decision to relocate its intake pipe is exogenous to the health outcomes of the households it serves. A federal court ruling that redraws eligibility for a program is exogenous to any individual applicant’s characteristics. The strongest natural experiments name the specific institutional, legal, or physical process that generated the shock, not just the shock itself.

2. Rule out selective sorting around the cutoff or event

Check whether units could self-select into or out of exposure once the shock was foreseeable. If eligibility is defined by a date or geographic boundary, check whether people or firms could relocate, delay, or reclassify to land on the favorable side. This is the same manipulation concern regression discontinuity design tests with a density check at the cutoff — a natural experiment with a sharp boundary should get the same scrutiny.

3. Test for parallel pre-trends

If your design compares a treated group’s before/after change against a comparison group’s (the standard difference-in-differences setup underlying most natural-experiment analyses), the comparison group has to be trending the same way the treated group would have absent the shock. You can’t observe the counterfactual directly, but you can check whether the two groups moved together before the shock. Plot both series (or estimate an event-study specification with lead terms) over several pre-periods: if the treated group was already diverging from the comparison group before the shock hit, the post-shock gap is at least partly a continuation of a pre-existing trend, not evidence of the shock’s effect. A flat, statistically insignificant set of pre-shock leads is the standard evidence researchers present; a visibly diverging pre-trend is close to disqualifying.

4. Run placebo checks

A placebo check applies your exact research design somewhere it should find nothing, and treats a null result as support for the design and a non-null result as a warning sign. Common versions for natural experiments:

  • Placebo outcome. Test the same shock against an outcome it has no plausible channel to affect. If a minimum-wage increase “predicts” changes in an unrelated outcome like local rainfall or an outcome measured before the policy existed, something in the design — not the policy — is generating the result.
  • Placebo timing. Re-run the analysis as if the shock happened at a fake date, usually one or two periods before the real one. A design that finds an “effect” at the fake date too is picking up something other than the actual shock — often exactly the pre-trend problem in point 3.
  • Placebo group. Re-run the comparison using a group that was never exposed to the shock at all, checking that no spurious “effect” appears where none should exist.

Placebo and pre-trend checks are close cousins — a pre-trend test is really a placebo-timing check restricted to the actual pre-period — and reviewers generally want to see both because they catch slightly different failure modes.

5. Check robustness to specification choices

Vary the comparison group, the window of time around the shock, and the control variables, and confirm the estimate doesn’t swing wildly. A result that only survives under one specific bandwidth, one specific comparison group, or one specific set of controls is fragile evidence for exogeneity even when every individual check above passes in isolation.

Named examples worth knowing

These are the natural experiments most often cited as the reference cases for what a credible design looks like — useful both as illustrations and as citations in a methods section.

Card and Krueger’s New Jersey minimum-wage study (1994)

David Card and Alan Krueger compared fast-food employment in New Jersey, which raised its minimum wage from $4.25 to $5.05 in April 1992, against employment just across the state line in eastern Pennsylvania, where the minimum wage didn’t change. The design’s exogeneity argument rests on adjacent counties sharing a regional labor market and economic trend, with the state border creating the only systematic difference in minimum-wage exposure. Contrary to the standard prediction that a wage floor increase would reduce employment, the study found no employment decrease in New Jersey relative to Pennsylvania. The finding was contested at the time (a notable critique by David Neumark and William Wascher used payroll records rather than the original phone survey and reported different results), which is itself a useful illustration of why robustness checks against alternative data sources matter. Card shared the 2021 Nobel Memorial Prize in Economic Sciences (with Joshua Angrist and Guido Imbens) in substantial part for this body of natural-experiment work.

Angrist’s Vietnam-era draft lottery study (1990)

Joshua Angrist used the United States’ Vietnam-era draft lottery — which assigned induction priority by randomly drawn birth dates — as a natural experiment to estimate the effect of military service on later earnings. Because draft-eligibility was determined by lottery number rather than by any characteristic of the individual, comparing draft-eligible and draft-ineligible men with similar birth-year cohorts approximates a randomized comparison, sidestepping the obvious selection problem in simply comparing veterans to non-veterans (people who enlist or get drafted differ systematically from those who don’t). Angrist found that veteran status was associated with roughly 15% lower subsequent earnings. The lottery’s randomization is also the textbook example of an instrumental variable: draft eligibility instruments for actual veteran status because it’s related to service but has no direct effect on earnings except through service itself.

John Snow’s South London water-company comparison (1854)

Predating the term “natural experiment” by roughly a century, John Snow’s investigation of London’s 1854 cholera epidemic compared households served by two water companies that drew from the River Thames at different points: the Southwark and Vauxhall Company, which drew water downstream of London’s sewage discharge, and the Lambeth Company, which had relocated its intake upstream in 1852. Because the two companies’ customers were interspersed street by street across the same neighborhoods, the water-source assignment functioned as if random with respect to any other determinant of cholera risk — households didn’t choose their supplier based on health status, and the companies’ pipe layouts predated the epidemic. Snow found sharply higher cholera death rates among Southwark and Vauxhall households, evidence for the (at the time still contested) waterborne theory of cholera transmission. The design is a standard reference point in epidemiology for how a genuinely exogenous, pre-existing institutional arrangement can substitute for randomization.

How natural experiments relate to IV, RD, and DiD

“Natural experiment” describes the source of variation — an exogenous, real-world event rather than a researcher-designed intervention. It’s usually paired with one of a small number of estimation strategies once you have that variation, and the choice between them depends on how the shock split the sample:

  • Difference-in-differences when you have a treated group and a comparison group observed before and after the shock — the Card and Krueger design.
  • Regression discontinuity when the shock creates a sharp cutoff on some continuous variable — an eligibility threshold, a boundary line, a ranking cutoff.
  • Instrumental variables (see endogeneity: the three sources, and the remedy that matches each) when the natural experiment affects the outcome only indirectly, through its effect on a specific endogenous variable — the draft lottery affecting earnings only through military service.
  • Synthetic control when a single treated unit (one state, one country, one firm) needs a weighted combination of untreated units as its counterfactual, rather than a single comparison group.

Whichever estimator you pair it with, the exogeneity argument for the underlying shock is the load-bearing piece — a well-executed difference-in-differences estimate built on a shock that turns out to be correlated with unobserved trends is still a biased estimate, just one with a more sophisticated-looking standard error. See experimental vs. quasi-experimental design for where natural experiments sit relative to a randomized trial on the credibility spectrum, and internal vs. external validity for the companion trade-off: a natural experiment’s exogeneity argument is chiefly about internal validity, and a highly local or historically specific shock (one draft lottery, one state’s policy) can still leave open questions about how far the estimate generalizes.

Common reviewer pushback, and how to pre-empt it

  • “Why this comparison group?” — justify the choice on institutional grounds (shared labor market, shared regulatory environment, geographic adjacency), not just on convenience or on the comparison group looking similar on observables.
  • “Show me it wasn’t already diverging.” — have the pre-trend plot or event-study leads ready before this is asked, not as a response to it.
  • “What if people saw this coming?” — address anticipation directly: was the shock announced in advance, and if so, could units have adjusted behavior in the announcement window? If yes, either exclude that window or treat the announcement itself, not the implementation date, as the actual shock.
  • “Does this hold with a different comparison group or window?” — report at least one alternative specification, even if it’s relegated to an appendix or supplementary table.

Frequently asked questions

What’s the difference between a natural experiment and a quasi-experiment?

The terms overlap heavily and are often used interchangeably. Where a distinction is drawn, “natural experiment” usually refers specifically to variation generated by an event genuinely outside any researcher’s or any study unit’s control (a law, a lottery, a geological event), while “quasi-experiment” is the broader category that also includes designs a researcher deliberately constructs without full randomization — a matched comparison group built after the fact, for instance. Every natural experiment is a quasi-experiment; not every quasi-experiment is a natural experiment. See experimental vs. quasi-experimental design for the fuller comparison.

Can a natural experiment ever be as credible as a randomized controlled trial?

It can approach that credibility for the specific comparison it makes, but it rests on an assumption — exogeneity — that a well-run randomized trial doesn’t need to argue for, because randomization guarantees it by construction. That’s exactly why the checklist in this guide exists: a natural experiment has to earn, through pre-trend and placebo evidence, what randomization gets for free. See what is a control group for how randomized assignment solves this problem directly.

What sample size do I need for a pre-trend test to be informative?

Enough pre-shock periods to distinguish a flat trend from a sloped one with reasonable power — two pre-periods rarely give a convincing test, since almost any two points can look “parallel.” Most published natural-experiment designs report at least three to five pre-shock periods where the underlying data allow it, precisely so a single noisy period doesn’t drive the conclusion either way.

Does a significant placebo result always kill the design?

Not automatically, but it demands an explanation. A placebo check that turns up significant is evidence the design is picking up something beyond the shock itself — the right response is to investigate what, not to drop the placebo check from the writeup. Reporting a placebo result honestly, including an inconvenient one, is itself part of what makes a natural-experiment paper credible to reviewers who’ve seen the alternative.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.