Skip to main content
v2026.11,610 entries · CC-BY 4.0

Statistical Significance: What the Verdict Means and Doesn’t Mean

A decision-framework guide to statistical significance: what the verdict asserts, how the alpha threshold is chosen, why sample size distorts it, the multiple-comparisons problem, and the research-integrity risks (p-hacking, HARKing) around the 0.05 line.

Ask about Statistical Significance: What the Verdict Means and Doesn’t Mean

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Statistical significance is a verdict, not a measurement. When a result is described as “statistically significant,” the claim being made is narrow and specific: the observed data would be sufficiently unlikely under a stated null hypothesis that the null hypothesis is rejected, according to a threshold that was fixed before the data were seen. That verdict says nothing on its own about whether an effect is real, how large it is, whether it matters in practice, or whether it will replicate. Confusing the decision rule with any of those four things is the single most common misuse of statistical significance in the research literature, and it is the reason this page exists as a companion to CASRAI’s guide to what a p-value is rather than a restatement of it.

\n\n

What “Statistically Significant” Actually Asserts

\n

A significance test starts with a null hypothesis (H₀) — typically “no effect” or “no difference” — and an alternative hypothesis. A p-value is calculated: the probability of observing a result at least as extreme as the one obtained, if H₀ were true. Before the data are collected, the researcher fixes a significance level, α (commonly 0.05). If the p-value falls at or below α, the result is declared statistically significant and the null hypothesis is rejected at that threshold.

\n

That is the entire content of the claim. It is a binary decision — reject or fail to reject — produced by comparing one number (the p-value) against another pre-set number (α). It is not a statement that the null hypothesis is false, not a probability that the alternative hypothesis is true, not a measure of effect size, and not a prediction about what a future study will find. Each of those is a distinct question with its own methods (effect sizes, confidence intervals, replication studies, Bayesian posterior probabilities) — a significance test does not answer any of them.

\n\n

The Significance Level (α): A Convention, Not a Law of Nature

\n

The 0.05 threshold has no basis in probability theory that forces it to be exactly 0.05. Its widespread use traces to the statistician Ronald Fisher, who in the 1920s suggested a p-value below 0.05 as a convenient, informal benchmark for a result worth a second look — not as a universal rule for accepting or rejecting a hypothesis. Over the following decades, 0.05 hardened from a rule of thumb into a near-automatic publication and decision threshold across many fields, a shift Fisher himself did not originally intend.

\n

Other thresholds are used deliberately in specific contexts:

\n

    \n

  • α = 0.01 or 0.001 — used where a false positive is especially costly, or where a field expects to run many tests and wants a stricter bar per test (see multiple comparisons, below).
  • \n

  • α = 5×10⁻⁸ (genome-wide significance) — used in genome-wide association studies, where hundreds of thousands to millions of statistical tests are run simultaneously and a 0.05 threshold would produce overwhelming numbers of false positives.
  • \n

  • α = 0.10 — occasionally used in exploratory or pilot work where missing a real effect (a Type II error) is considered more costly than a false positive.
  • \n

\n

What matters more than the specific number is a procedural rule: α must be chosen before looking at the data, and it must not change once results are known. Selecting or adjusting a threshold after seeing results — even informally, even unconsciously — breaks the logic of the test and inflates the real false-positive rate far above the stated α. This is the same principle underlying the research-integrity concerns covered later on this page.

\n\n

How Significance Relates to the P-Value

\n

Significance is the decision that follows from comparing a p-value to α — the p-value is the input, the significance verdict is the output. For the full mechanics of what a p-value measures, the four most common misinterpretations of it, and how to read one correctly, see CASRAI’s dedicated guide, What Is a P Value? This page assumes that definition and focuses on what happens once a threshold decision has been made from it.

\n\n

Statistical Significance vs. Practical (Clinical) Significance

\n

A p-value is a function of both the size of an effect and the size of the sample. That single fact is the source of most real-world confusion about what “significant” means, and it cuts in both directions:

\n

    \n

  • Large samples make trivial effects “significant.” Given enough observations, even a tiny, practically meaningless difference — a fraction of a point on a scale, a rounding-level shift in a rate — will eventually produce a p-value below 0.05, because the test is detecting that the true difference is unlikely to be exactly zero, not that it is large enough to matter.
  • \n

  • Small samples make important effects “non-significant.” A genuinely large, clinically or practically important effect can fail to cross the significance threshold simply because the study did not enroll or measure enough subjects to detect it reliably.
  • \n

\n

In other words, statistical significance is partly a statement about sample size, not purely about the strength or importance of a finding. A result can be statistically significant and practically irrelevant, or practically important and statistically non-significant, in the same dataset depending on how large the study was. For the full treatment of this distinction — including the Minimal Clinically Important Difference (MCID) concept used in clinical trials — see CASRAI’s comparison, Statistical Significance vs. Clinical Significance.

\n\n

What a Non-Significant Result Does Not Mean

\n

Failing to reach significance is routinely, incorrectly, reported as “no effect” or “no difference.” A non-significant result supports neither conclusion on its own. It means only that the observed data were not sufficiently inconsistent with the null hypothesis to reject it at the chosen α — which can happen because there truly is no effect, or because the study lacked the statistical power to detect a real effect that exists.

\n

This is the “absence of evidence is not evidence of absence” problem. An underpowered study — one with too small a sample, too much measurement noise, or too small an expected effect relative to its sample size — can produce a non-significant result even when a real, meaningful effect is present. That result is uninformative about the question, not evidence against it. Whether a given non-significant result is a meaningful null finding or simply an underpowered one depends on the study’s statistical power, which is normally assessed and reported through a power analysis conducted at the design stage, not after the fact.

\n\n

Type I and Type II Errors, and the Power Trade-off

\n

Every significance-testing decision carries two possible kinds of error:

\n

    \n

  • Type I error (false positive) — rejecting a true null hypothesis; declaring a result significant when there is actually no effect. The probability of a Type I error is α itself, by construction.
  • \n

  • Type II error (false negative) — failing to reject a false null hypothesis; missing a real effect. Its probability is denoted β.
  • \n

\n

Statistical power is 1 − β: the probability of correctly detecting a real effect of a given size. For a fixed sample size, lowering α (a stricter significance threshold) reduces the Type I error rate but increases the Type II error rate — the trade-off runs in both directions, and the only way to reduce both simultaneously is to increase the sample size or the precision of the measurement. For the full comparison of causes, consequences, and mitigation strategies for each error type, see CASRAI’s comparison, Type I and Type II Errors.

\n\n

The Multiple Comparisons Problem

\n

α describes the false-positive rate of a single test. Run many tests on the same data — comparing several outcome measures, several subgroups, several time points — and the probability that at least one test crosses the significance threshold purely by chance (the familywise error rate) climbs well above the nominal α for any individual test. With 20 independent tests at α = 0.05, the chance of at least one false positive somewhere in the set is roughly 64%, not 5%.

\n

Standard corrections address this in different ways:

\n

    \n

  • Bonferroni correction — divides α by the number of tests (e.g., 0.05 / 20 = 0.0025 per test). Simple and conservative; controls the familywise error rate but can be overly strict when many tests are run, increasing Type II error risk.
  • \n

  • Holm-Bonferroni method — a step-down procedure that applies the Bonferroni logic sequentially to ranked p-values, controlling the same familywise error rate with somewhat more power than the plain Bonferroni correction.
  • \n

  • False discovery rate (FDR) / Benjamini-Hochberg procedure — instead of controlling the probability of any false positive at all, controls the expected proportion of false positives among the results called significant. This is a less conservative standard than familywise error control, which is precisely why FDR methods are the default in fields that run thousands to millions of simultaneous tests — genomics (differential gene expression), neuroimaging, and other high-dimensional data settings — where a strict familywise correction would leave almost no findings detectable at all.
  • \n

\n

Choosing which correction to apply — or whether to correct at all — should be decided as part of the analysis plan, before results are seen, for the same reason α itself must be pre-set.

\n\n

Research Integrity: How the Significance Threshold Gets Gamed

\n

Because crossing a fixed threshold has outsized consequences for what gets published, reported, or acted on, the significance threshold creates a direct incentive to manipulate the path to p < α, whether deliberately or through unrecognized “researcher degrees of freedom.” CASRAI treats this as a research-integrity issue, not just a statistical one:

\n

    \n

  • P-hacking — trying multiple analyses, subgroups, outcome definitions, or exclusion criteria until one combination crosses the significance threshold, then reporting only that result as if it were the pre-planned analysis.
  • \n

  • HARKing (Hypothesizing After the Results are Known) — presenting a hypothesis discovered by exploring the data as though it had been specified in advance, which converts an exploratory finding into a false confirmatory one. See CASRAI’s comparison, P-hacking vs. HARKing, for how the two practices differ and overlap.
  • \n

  • Optional stopping — repeatedly checking results as data accumulate and halting data collection as soon as significance is reached, which inflates the true false-positive rate well above the nominal α even though no single analysis appears improper in isolation.
  • \n

  • Selective outcome reporting — measuring several outcomes but reporting only the ones that reached significance, omitting the pre-specified primary outcome if it did not.
  • \n

  • Publication bias / the file-drawer problem — significant results are more likely to be submitted and accepted for publication than non-significant ones, so the published literature over-represents significant findings relative to what was actually studied, distorting the apparent weight of evidence on a question.
  • \n

\n

These practices fall under what CASRAI’s dictionary catalogs as questionable research practices — behaviors that fall short of outright fabrication or falsification but that distort the evidentiary value of a significance verdict.

\n

The standard remedies work by removing the researcher’s ability to select, after the fact, which analysis or outcome gets reported as the significant one:

\n

    \n

  • Preregistration — publicly time-stamping the hypothesis, primary outcome, sample size, and analysis plan before data collection begins, so the pre-specified analysis is verifiable and any deviation is visible.
  • \n

  • Registered reports — a publication format in which the introduction, hypotheses, and methods are peer-reviewed and provisionally accepted before data are collected, removing the incentive to alter the analysis to reach significance, since the paper’s acceptance no longer depends on the result.
  • \n

  • Transparent, complete outcome reporting — reporting every pre-specified outcome and every planned analysis, significant or not, rather than a subset chosen after seeing the results.
  • \n

\n\n

The Reform Debate: Is 0.05 the Right Bar, or the Wrong Question?

\n

The reliance on a single significance threshold has been the subject of sustained methodological debate. In 2016, the American Statistical Association published a formal statement on p-values, cautioning against reducing scientific conclusions to whether a p-value crosses 0.05 and stating explicitly that “smaller p-values do not necessarily imply the presence of larger or more important effects, and larger p-values do not imply a lack of importance or even lack of effect” (Wasserstein & Lazar, The American Statistician, 70(2), 2016).

\n

Since then, two distinct reform proposals have gained traction, and they point in different directions:

\n

    \n

  • Redefine the threshold. A 2018 proposal led by Daniel Benjamin and dozens of co-authors, published in Nature Human Behaviour, argued that the default threshold for claiming a new discovery should move from p < 0.05 to p < 0.005, on the reasoning that 0.05 is too lenient a bar given how many “significant” findings in fields like psychology and biomedicine have failed to replicate.
  • \n

  • Abandon the label. A competing camp, prominent in a widely-signed 2019 commentary, argued that the problem is dichotomization itself — treating p = 0.049 and p = 0.051 as categorically different results — and called for retiring the phrase “statistically significant” altogether in favor of reporting effect sizes, confidence intervals, and exact p-values without a bright-line verdict.
  • \n

\n

The counterargument for keeping some form of pre-set threshold is practical rather than statistical: regulators, journals, and clinical decision-makers need a reproducible, pre-agreed rule for when a result is treated as actionable, and a purely descriptive report without any decision rule can shift the same p-hacking-style flexibility from “which analysis to run” to “which result to call convincing” after the fact. CASRAI does not take a position on which camp is right; a research administrator or investigator should expect this debate to keep shaping journal and funder reporting requirements, and should look for the specific reporting standard a target journal or funder actually requires rather than assuming 0.05 is universal.

\n\n

What to Report Instead of — or Alongside — a Significance Verdict

\n

Across nearly all sides of the reform debate, the same practical recommendations recur for what a results section should contain beyond a bare “p < .05”:

\n

    \n

  • Effect sizes with confidence intervals — the magnitude of the observed effect (see CASRAI’s guide to effect size), reported together with a confidence interval that conveys the precision of the estimate, rather than a p-value alone.
  • \n

  • Exact p-values — reporting the precise value obtained (e.g., p = 0.031) rather than only the inequality relative to a threshold (p < .05), so a reader can judge strength of evidence rather than a binary pass/fail.
  • \n

  • Pre-specified, published analysis plans — whether through preregistration, a registered report, or a trial protocol, a plan fixed before the data were seen is what gives any subsequent significance verdict — or effect-size estimate — its evidentiary weight in the first place.
  • \n

\n\n

Frequently Asked Questions

\n

Is a statistically significant result always a “real” effect?

\n

No. By construction, a significance test at α = 0.05 will produce a false positive in roughly 5% of cases where the null hypothesis is actually true, purely by chance. Significance is evidence against the null hypothesis at a stated error rate, not proof that an effect exists.

\n

What p-value counts as “significant”?

\n

Whatever value is at or below the pre-set α for that analysis — most commonly 0.05, though stricter thresholds (0.01, 0.001, or far stricter in high-dimensional fields like genomics) are used where appropriate. There is no p-value that is universally “significant” independent of the threshold chosen for that study.

\n

Can a result be statistically significant but not important?

\n

Yes. This is the central practical significance vs. statistical significance distinction covered above — a large enough sample can make a trivially small, practically unimportant effect statistically significant.

\n

Does a non-significant result prove there is no effect?

\n

No. It means the data did not provide sufficient evidence to reject the null hypothesis at the chosen threshold — which can reflect a true null effect or simply an underpowered study. See “What a Non-Significant Result Does Not Mean,” above.

\n

Why do some fields use a much stricter threshold than 0.05?

\n

Fields that run very large numbers of simultaneous statistical tests — genome-wide association studies are the clearest example — use far stricter thresholds (or false discovery rate control) specifically because 0.05 applied to hundreds of thousands of tests would produce overwhelming numbers of false positives. See “The Multiple Comparisons Problem,” above.

\n

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →