Cohen’s d of 0.5 is not automatically “medium.” The 0.2/0.5/0.8 thresholds that
almost every textbook and calculator site repeats come from a single 1988 rule of thumb that Jacob
Cohen himself described as arbitrary and intended only for use when a field has no better basis for
judging effect sizes. In practice, “medium” in one field is enormous in another. A d of 0.5 in a
medical trial data set is a big deal; the same d of 0.5 in a personality-psychology study sits near
the average effect the field routinely publishes. Reading a d value against the wrong yardstick is
one of the more common interpretation errors in a results section — this guide gives you the
field-calibrated yardsticks, a worked example of the sentence a researcher actually writes from a d
value, and when to report Hedges’ g instead.
What Cohen’s d actually measures
Cohen’s d is a standardized mean difference: the difference between two group means, expressed in
units of the pooled standard deviation, rather than in the outcome’s original units. That
standardization is the entire point — it lets you compare an effect measured in milliseconds to one
measured in points on a Likert scale, or compare effects across two studies that used different
instruments.
The basic formula for independent groups is:
d = (M₁ − M₂) / SDpooled
where SDpooled is the pooled standard deviation of the two groups. A d of 1.0 means the
two group means are one pooled standard deviation apart. Related indices — confidence interval around d, and the p-value from the underlying test — answer a different question (is the difference likely due to chance) and should always be reported alongside d, never in place of it. See how to report p-values for that companion piece.
A worked interpretation
Suppose a study compares a coaching intervention against a waitlist control on a 40-point
self-efficacy scale:
| Intervention (n=42) | Control (n=40) | |
|---|---|---|
| Mean | 28.6 | 25.1 |
| SD | 5.9 | 6.2 |
Pooled SD works out to about 6.05, giving d = (28.6 − 25.1) / 6.05 ≈ 0.58, 95% CI
[0.14, 1.01]. The sentence a researcher would write from that is something like: “Participants in
the coaching condition scored higher on self-efficacy than waitlist controls (Mdiff = 3.5,
d = 0.58, 95% CI [0.14, 1.01]), a medium-to-large effect by conventional benchmarks, though the wide
interval reflects the modest sample and is consistent with anywhere from a small to a very large true
effect.” Notice what that sentence does: it reports the point estimate, gives it a size
description, and — critically — lets the confidence interval qualify the certainty rather than letting
the point estimate alone imply more precision than 82 participants can support. A bare “d = 0.58” with
no interval is an incomplete report by current APA and CONSORT-aligned standards.
Why 0.2 / 0.5 / 0.8 is field-dependent
Cohen’s own benchmarks were calibrated loosely against behavioral-science research of the 1960s-80s.
Later empirical reviews of published effect sizes, discipline by discipline, show the actual
distribution of “typical” effects varies enough that applying one fixed scale across fields routinely
misclassifies results — a genuinely large, field-beating effect in medicine can fall well under the
generic “small” cutoff, and a routine, unremarkable effect in social psychology can clear the generic
“large” one.
| Field / literature | Small | Medium | Large | Note |
|---|---|---|---|---|
| Cohen’s original (1988) generic convention | 0.2 | 0.5 | 0.8 | Explicitly offered as a fallback only, not a rule |
| Psychology (published-literature average) | <0.2 | ~0.4 | >0.8 | Mean published effect is closer to d = 0.4 than to Cohen’s “medium” of 0.5 |
| Education research | 0.2 | 0.4 | 0.6 | Commonly cited alternative scale shifted down from Cohen’s |
| Gerontology | 0.15 | 0.40 | 0.75 | Field-specific benchmarks proposed to match the discipline’s typical effect distribution |
| Clinical medicine | 0.05–0.2 can be meaningful | — | — | A “small” effect on Cohen’s generic scale can still represent a clinically important, life-saving difference at population scale |
The practical rule: report the raw effect (the mean difference in original units, e.g. “3.5 points
on the 40-point scale”) alongside d, cite the specific convention you’re using and attribute it, and
where possible compare your d to the distribution of effects already published in your own
sub-literature rather than to a single borrowed threshold. Presenting 0.2/0.5/0.8 as a fixed rule
rather than the convenience convention it was designed to be is itself a common and avoidable
reporting error.
Hedges’ g: the small-sample correction
Cohen’s d has a known small-sample upward bias — it tends to slightly overestimate the true
population effect size when group sizes are small, roughly under 20 per group. Hedges’ g applies a
correction factor to remove that bias:
g = d × J, where J = 1 − 3 / (4df − 1)
with df the degrees of freedom for the pooled standard deviation (n₁ + n₂ − 2 for
two independent groups). The correction factor J is always slightly less than 1, so g is always
slightly smaller in magnitude than d. For the worked example above (df = 80), J ≈ 0.991 — the
correction moves d = 0.58 to g ≈ 0.575, a negligible shift because the sample isn’t small. Run the
same correction on a d = 0.58 from two groups of n = 8 each (df = 14), and J ≈ 0.946 — g drops to
about 0.549, a difference large enough to matter when the result is small and n is small. As a rule of
thumb: report Hedges’ g instead of Cohen’s d whenever either group has fewer than about 20
observations, and report g by default in a meta-analysis, since study-level bias otherwise
accumulates and skews the pooled estimate. See PRISMA and systematic review methodology for how effect-size synthesis fits into that broader workflow.
Assumptions and when Cohen’s d isn’t the right index
- Roughly normal, continuous outcome. Cohen’s d assumes an approximately normal,
interval- or ratio-scaled outcome in each group. For ordinal data (ranks, Likert items treated as
ordinal) or heavily skewed distributions, a rank-based effect size (e.g. rank-biserial correlation) is
more appropriate — check normality of the distribution
and skewness before defaulting to d. - Unequal variances between groups. The pooled-SD version of d assumes the two
groups have similar variance. When variances differ substantially, Glass’s delta (which divides by the
control group’s SD alone) is the more defensible choice. - Paired or repeated-measures designs. A within-subject or pre/post design needs a
different denominator convention (dz, standardized against the SD of the difference scores,
or dav, against the average of the two SDs) — plugging pre/post means into the independent-
groups formula above overstates or understates the effect depending on which convention silently gets
used. - Small, unrepresentative samples. A large d from a small, non-random sample is a
precision problem, not a strength — report the confidence interval, not just the point estimate, and
prefer Hedges’ g as above. - Correlated or clustered data. Cohen’s d as defined here assumes independent
observations. Clustered designs (students within classrooms, patients within clinics) need a design-
adjusted effect size or the clustering will inflate apparent precision.
Frequently asked questions
Is a Cohen’s d of 0.5 a “good” effect?
It depends entirely on the field and what the outcome measures. By Cohen’s generic convention, 0.5
is labeled “medium.” In psychology’s published literature, where the average effect size is closer to
0.4, a 0.5 is somewhat above average. In clinical medicine, where meaningful population-level effects
are often well under 0.2, a 0.5 would be considered a notably large effect. Always state which
convention you’re using and, where possible, compare against effects already published in your
specific sub-literature rather than relying on the label alone.
What’s the difference between Cohen’s d and Hedges’ g?
Hedges’ g is Cohen’s d with a small-sample bias correction (the J factor above) applied. For group
sizes of roughly 20 or more per group the two are nearly identical; for smaller samples, g is the more
accurate estimate of the true population effect and is generally preferred, especially in meta-analysis
where small-study bias otherwise compounds across many pooled effects.
Can Cohen’s d be negative?
Yes — the sign simply depends on which group’s mean is subtracted from which and which direction is
treated as “improvement.” A negative d is not an error; it means the difference runs opposite to
whatever direction was defined as M₁ − M₂. Report the direction in words alongside
the number so the sign is unambiguous to a reader.
How large a sample do I need to detect a given Cohen’s d?
That’s a power-analysis question, not a d-interpretation one — the required sample size depends on
d, your chosen significance level, and your target power (conventionally 80%), and is calculated
before data collection, not from the observed d after the fact (“post hoc power” computed from an
already-obtained result is not informative and is generally discouraged in reporting).
Related terms
For definitions of the underlying concepts referenced above, see
effect size,
confidence interval, and
p-value. For related interpretation guides, see
correlation coefficient,
chi-square test,
intraclass correlation coefficient (ICC),
Cronbach’s alpha, and
logistic regression.







