Examples
Worked examples
- Is an instance
A researcher with a 200-variable survey dataset runs pairwise correlations between every variable and an outcome of interest — roughly 20,000 comparisons — then writes up the handful that cross p < 0.05 as the study’s central finding, without disclosing how many comparisons were run or correcting the threshold for them.
- Is an instance
Genome-wide association studies (GWAS) test a trait against hundreds of thousands to millions of genetic variants in a single analysis. Because that scale of search would produce large numbers of false positives at the conventional p < 0.05 threshold, the field adopted a much stricter, commonly cited genome-wide significance threshold of p < 5×10⁻⁸, calibrated specifically to correct for the number of variants tested.
Counter-examples
Looks similar, but isn't
- Not an instance
A researcher pre-registers a single primary hypothesis and analysis plan before collecting data, then separately reports several additional subgroup analyses explicitly labeled as exploratory and interpreted cautiously rather than as confirmatory findings. Examining many variables is not itself data dredging — the disclosure and the exploratory label make the difference.
- Not an instance
A study applies a Bonferroni or false-discovery-rate correction across all comparisons it performed and reports which relationships remain significant after that correction. Because the number of tests is disclosed and the significance threshold is adjusted for it, this is standard multiple-comparisons practice, not data dredging.
Editorial commentary
Data dredging is the practice of searching a dataset for statistically significant relationships — across many variables, subgroups, or model specifications — without a hypothesis specified in advance, and then reporting whichever pattern crosses a significance threshold as though it had been the study’s original object of investigation. The term was introduced by Hanan Selvin and Alan Stuart in their 1966 paper ‘Data-Dredging Procedures in Survey Analysis’ (The American Statistician, Vol. 20, No. 3), which described the practice as ‘hunting’ for relationships in survey data and warned that a variable retained after such a search cannot validly be evaluated with the same statistical procedures used for a variable specified before looking at the data.
References
- Selvin H, Stuart A. ‘Data-Dredging Procedures in Survey Analysis.’ The American Statistician. 1966;20(3):20–23.
- Simmons JP, Nelson LD, Simonsohn U. ‘False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant.’ Psychological Science. 2011;22(11):1359–1366.
- Ioannidis JPA. ‘Why Most Published Research Findings Are False.’ PLOS Medicine. 2005;2(8):e124.
Why it inflates false positives: the multiple comparisons problem
A single statistical test run at the conventional p < 0.05 threshold has, by construction, roughly a 5% chance of a false positive even when there is no real underlying effect. That risk compounds with every additional comparison run against the same dataset: testing 20 unrelated variables against an outcome, with no real effects present, would on average produce about one false positive purely by chance. Data dredging is what happens when that arithmetic is not disclosed — a researcher runs many comparisons, but reports only the one (or few) that happened to reach significance, presenting it as if it were the single, pre-specified test the study set out to run. The reported p-value is then wrong in a specific, quantifiable way: it describes the probability of that one result under the null hypothesis, not the much higher probability that some result among everything tested would have crossed the threshold by chance.
Data dredging vs. p-hacking vs. HARKing
These three terms describe overlapping but structurally distinct failure modes, and CASRAI’s dictionary keeps each as its own entry because the fix for each is not identical:
- Data dredging is a search problem: it happens at the stage of deciding which relationships to test in the first place, typically by scanning many variables or subgroups in an existing dataset without a hypothesis driving the search.
- P-hacking is an analysis-flexibility problem: it happens once a specific relationship is already the focus, by adjusting covariates, exclusion criteria, or model specification until that one relationship crosses significance.
- HARKing (hypothesising after results are known) is a reporting problem: it happens after a result is already in hand, by presenting it in the write-up as though it had been the a priori hypothesis rather than something found along the way.
In practice the three often chain together: a data-dredging search across many variables turns something up, the analysis is then massaged (p-hacked) to sharpen the effect, and the write-up HARKs by presenting the finding as the study’s original confirmatory hypothesis. CASRAI’s researcher degrees of freedom and garden of forking paths entries describe the underlying space of undisclosed analytic choices that data dredging, p-hacking, and HARKing all draw on.
Data dredging vs. legitimate exploratory analysis
Searching a dataset broadly is not itself the problem — exploratory data analysis is a legitimate, well-established part of research, and large observational or secondary datasets are often collected precisely so multiple questions can be asked of them over time. What separates exploratory analysis from data dredging is disclosure and framing, not the number of variables examined. An analysis is exploratory, not data dredging, when the researcher explicitly labels it as hypothesis-generating, reports how many comparisons were run (or applies a correction for them), and does not present the resulting pattern as though it had confirmed a hypothesis that predated the search. A pre-analysis plan or formal pre-registration makes this distinction verifiable to a reader, rather than something the reader has to take on trust, by fixing which analysis is confirmatory before the data are examined.
How research fields guard against it
- Pre-registration and pre-analysis plans. Committing to a primary hypothesis and analysis plan before looking at outcome data removes the opportunity to dredge, since any relationship found outside the registered plan is then reportable only as exploratory, not confirmatory.
- Registered Reports. Because a journal reviews and conditionally accepts the protocol before results exist, a Stage 1-approved study cannot substitute a dredged post-hoc finding for its registered primary analysis without that substitution being visible to reviewers.
- Multiple-comparisons correction. Statistical procedures such as the Bonferroni correction or false discovery rate (FDR) control adjust the significance threshold to account for the number of tests actually run, so that the field-wide false-positive rate stays close to the nominal level even when many comparisons are performed.
- Field-wide significance conventions calibrated to test volume. Some fields have institutionalized a stricter threshold specifically because their standard analyses involve searching across enormous numbers of variables. Genome-wide association studies (GWAS), which routinely test a trait against hundreds of thousands to millions of genetic variants in a single analysis, conventionally use a much stricter significance threshold (commonly cited as p < 5 × 10⁻⁸) than the p < 0.05 used for a single pre-specified test, precisely to correct for the scale of the search.
- Replication. A relationship that survived data dredging in one dataset is, by construction, less likely to hold up in a fresh, independent dataset than a genuine effect — which is why an unreplicated finding from an exploratory search is treated as substantially weaker evidence than a preregistered, replicated one. See CASRAI’s replication study entry.
Frequently asked
Is data dredging the same as p-hacking? No, though the two are closely related and often occur together. Data dredging is about searching broadly across a dataset for a relationship to test in the first place; p-hacking is about manipulating the analysis of a relationship already chosen. A study can p-hack a single pre-specified hypothesis without any data dredging having occurred, and a data-dredging search can turn up a finding that is then reported honestly as exploratory rather than p-hacked.
Does using a large dataset automatically mean a study is data dredging? No. A large dataset with many variables only becomes a data-dredging problem when the search across those variables is undisclosed and the resulting finding is presented as if it had been the pre-specified object of the study. The same dataset, searched openly and reported as exploratory or corrected for multiple comparisons, is legitimate research.
Can data dredging happen by accident, without intent to mislead? Yes, and this is widely considered the more common case. A researcher can genuinely believe a pattern found through exploration is meaningful and report it in good faith as a confirmatory result, simply because the distinction between exploratory and confirmatory analysis was never made explicit during the analysis itself — which is exactly why structural safeguards like pre-registration matter more than researcher intent.
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="Data Dredging"
vocab-term-identifier="https://casrai.org/dictionary/term/data-dredging" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/data-dredging",
"name": "Data Dredging",
"identifier": "https://casrai.org/dictionary/term/data-dredging",
"description": "Searching a dataset for statistically significant relationships across many variables, subgroups, or model specifications without a hypothesis specified in advance, then reporting whichever pattern crosses a significance threshold as if it had been the study’s pre-specified object of investigation, without correcting for the number of comparisons actually performed.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/reproducibility#set",
"url": "https://casrai.org/dictionary/term/data-dredging",
"sameAs": [],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"dateModified": "2026-07-17T21:26:46",
"inLanguage": "en"
}






