Criterion validity asks a narrower, more testable question than construct validity: does this instrument’s score correlate with an independently measured outcome that the field already accepts as meaningful? That outcome is the “criterion.” Criterion validity comes in two forms depending on when the criterion is measured relative to the instrument: concurrent validity (criterion measured at roughly the same time) and predictive validity (criterion measured later). This guide covers both forms in depth, how to choose and evaluate a criterion, the criterion-contamination pitfall, and how unreliability in either measure attenuates the correlation you observe — with a worked numerical example. For how criterion validity fits alongside construct and content validity, see Types of Validity in Research and the dedicated construct validity guide.
Concurrent Validity vs. Predictive Validity
Both are forms of criterion-related validity evidence. The distinction is entirely about timing, not about the underlying logic of the correlation.
| Dimension | Concurrent Validity | Predictive Validity |
|---|---|---|
| When the criterion is measured | At roughly the same time as the instrument | At a defined later time point |
| Typical purpose | Substitute a shorter, cheaper, or less invasive instrument for an established one | Forecast a future state, behaviour, or outcome |
| Typical study design | Cross-sectional: both measures collected in the same session or same short window | Longitudinal or prospective: instrument administered first, criterion collected after a follow-up interval |
| Worked example | A brief self-report depression screener administered alongside a structured clinical interview in the same visit | A graduate-admissions test administered at application, correlated with GPA at degree completion |
| Main practical risk | The “gold standard” comparison instrument may itself be imperfect or costly to justify as a criterion | Attrition, cohort changes, or a long follow-up window can erode the correlation for reasons unrelated to the instrument |
Some methods texts use “concurrent” and “predictive” as if they were separate validity types; the more common current framing, consistent with the Standards for Educational and Psychological Testing (AERA/APA/NCME), treats both as evidence based on relations to a criterion — a single validity argument, gathered at two different points in the measurement timeline.
What Counts as a Valid Criterion
A criterion is only useful if it is itself a defensible, independently obtained measure of the outcome the instrument is meant to predict or track — not simply another self-report version of the same instrument. Common criterion types:
| Criterion type | Example instrument being validated against it | Common limitation |
|---|---|---|
| Structured clinical interview or diagnosis | Brief self-report symptom screener | Interview itself has imperfect inter-rater reliability, which caps the observed correlation |
| Academic or job-performance record (GPA, supervisor rating, output metric) | Admissions test, aptitude test, competency assessment | Performance records are influenced by many factors besides the trait the instrument measures |
| Biomarker or laboratory value | Self-report health-behaviour or symptom scale | Biomarkers can lag or lead the construct of interest, weakening a same-time correlation |
| Documented behavioural outcome (readmission, attrition, incident report) | Risk-screening or readiness instrument | Base rates are often low, which restricts the correlation that is mathematically achievable |
| Licensure, certification, or credentialing outcome | Pre-training aptitude or knowledge test | Pass/fail outcomes are binary, which typically lowers correlation coefficients relative to continuous criteria |
Interpreting the Correlation Coefficient
Criterion validity is usually reported as a single correlation coefficient (Pearson’s r, or a point-biserial/tetrachoric variant when the criterion is binary). General behavioural-science conventions for interpreting r — commonly attributed to Cohen (1988) — treat 0.10 as small, 0.30 as medium, and 0.50 as large. Applied criterion-validity coefficients in real research and assessment settings routinely fall in the 0.20–0.50 range and are still considered practically useful, for two structural reasons this guide covers below: no single criterion fully captures a complex construct (criteria are themselves imperfect), and unreliability in either measure mechanically attenuates the observed correlation regardless of how valid the instrument truly is.
| Observed r | Conventional label | Practical read for criterion validity |
|---|---|---|
| < 0.20 | Small / weak | Investigate whether the criterion is poorly matched, contaminated, or unreliable before concluding the instrument itself lacks validity |
| 0.20 – 0.35 | Small-to-medium | Common and often acceptable for screening tools used alongside other information, not as a sole decision criterion |
| 0.35 – 0.50 | Medium-to-large | Typical target range for an instrument intended to substitute for, or meaningfully predict, an established criterion |
| > 0.50 | Large | Strong evidence, but check for criterion contamination (below) before treating it as confirmatory |
Choosing a Criterion
Before collecting any data, the criterion itself needs the same scrutiny as the instrument under evaluation. A defensible criterion should be:
- Relevant — it should measure the actual outcome of interest, not a convenient proxy that happens to be available.
- Reliable — an unreliable criterion caps the correlation you can ever observe, independent of the instrument’s quality (see Attenuation, below).
- Independent — collected and scored without knowledge of the instrument’s scores, to avoid criterion contamination.
- Free of restricted range — if the sample used to validate the instrument has already been pre-selected on the criterion (e.g., validating an admissions test only on students who were admitted), the correlation will be mechanically suppressed.
- Practically obtainable at the scale the instrument will actually be used — a criterion that requires a resource-intensive gold-standard assessment for every case undermines the rationale for a shorter instrument in the first place.
Methods literature sometimes calls the gap between an ideal criterion and any available real-world measure the “criterion problem” — every available criterion is itself deficient or contaminated to some degree, so criterion-validity evidence should always be read as evidence relative to a specific, named criterion, not as an absolute property of the instrument.
Criterion Contamination
Criterion contamination occurs when the criterion measure is influenced by something other than the outcome it is supposed to represent — most commonly, when whoever is scoring or rating the criterion has knowledge of the instrument’s scores. If a supervisor rating “job performance” (the criterion) already knows a candidate’s test score (the predictor), that knowledge can bias the rating upward or downward independent of actual performance, inflating the observed correlation without reflecting any real predictive relationship. This concern traces back to the classic distinction in Cronbach and Meehl’s foundational 1955 treatment of validity: a criterion can be deficient (missing parts of the true outcome), contaminated (including things it shouldn’t), or both at once.
Practical safeguards against contamination:
- Blind criterion raters to the instrument’s scores wherever feasible.
- Use criteria drawn from records or systems that were generated independently of the study (e.g., administrative data, existing clinical records) rather than purpose-collected ratings from people aware of the study hypothesis.
- Separate the personnel or systems that administer the instrument from those that score the criterion.
Attenuation: Why Observed Correlations Understate True Validity
Every correlation coefficient is capped by the reliability of the two measures being correlated. If either the instrument or the criterion has measurement error — and every real-world measure does, to some degree — the observed correlation will be lower than the “true” relationship between the underlying constructs. This is called attenuation, and it is one of the oldest results in psychometrics, formalised by Charles Spearman in 1904 as the correction for attenuation:
rcorrected = rxy ÷ √(rxx × ryy)
where rxy is the observed correlation between the instrument and the criterion, rxx is the reliability of the instrument (e.g., Cronbach’s alpha — see the Cronbach’s alpha guide), and ryy is the reliability of the criterion (e.g., inter-rater reliability for a rated criterion; see Test-Retest vs. Inter-Rater Reliability).
Worked example
The following is an illustrative calculation using hypothetical numbers, not a reported finding from any specific study. Suppose a new brief screening instrument is validated concurrently against a structured clinical interview. The observed correlation between the two is rxy = 0.40. Internal-consistency reliability (Cronbach’s alpha) for the screening instrument is rxx = 0.85. Inter-rater reliability for the clinical interview is ryy = 0.75.
Applying the formula: rcorrected = 0.40 ÷ √(0.85 × 0.75) = 0.40 ÷ √0.6375 = 0.40 ÷ 0.798 ≈ 0.50.
The disattenuated correlation (0.50) is noticeably higher than the observed correlation (0.40) — not because the instrument changed, but because correcting for measurement error in both variables reveals a stronger underlying relationship than the raw correlation showed. This is why a modest observed criterion-validity coefficient should prompt a check of both measures’ reliability before it is read as a weak validity finding. It is also why the correction should be used to interpret evidence, not to replace the reported observed correlation in a published validity argument — both figures are normally reported together, with the disattenuated value labelled explicitly as a correction.
Criterion Validity vs. Construct and Content Validity
Criterion, construct, and content validity are not competing techniques; they are different kinds of evidence for the same underlying question — does this instrument measure what it claims to? Criterion validity is the most directly testable of the three because it collapses to a single correlation coefficient, but a correlation with one flawed or narrow criterion is not sufficient evidence on its own. The Standards for Educational and Psychological Testing frames validity as a unified concept supported by multiple, converging lines of evidence: content (does the instrument’s coverage match the domain?), internal structure and relations to other variables (does it behave the way the underlying construct should behave, covered in the construct validity guide), and criterion-related evidence (does it correlate with an accepted external outcome, the subject of this guide). See Types of Validity in Research for how all of these fit together, and Reliability vs. Validity and Internal vs. External Validity for two related distinctions that are commonly confused with criterion validity but answer different questions.
Frequently Asked Questions
What is the difference between concurrent and predictive validity?
Both correlate an instrument’s scores with an external criterion. Concurrent validity measures the criterion at roughly the same time as the instrument (often used to justify a shorter or cheaper substitute for an established measure); predictive validity measures the criterion at a later time point (used to justify the instrument as a forecasting tool).
How high should a criterion validity correlation be?
There is no universal cutoff. Correlations of 0.20–0.35 are common and often considered practically useful for screening purposes, while 0.35–0.50 is a more typical target when an instrument is meant to substitute for or meaningfully predict an established criterion. Any interpretation should account for the reliability of both the instrument and the criterion, since unreliability mechanically suppresses the observed correlation (see Attenuation, above).
What is criterion contamination?
It is bias introduced into the criterion measure by something other than the true outcome — most often, a rater’s knowledge of the instrument’s scores influencing how they score the criterion. It inflates the observed correlation and should be guarded against by blinding criterion raters to instrument scores wherever possible.
Can an instrument have criterion validity without construct validity?
In principle a scale can correlate with a criterion for reasons unrelated to the construct it claims to measure — for example, both may share an unrelated confound. This is exactly why the current testing standards treat criterion-related evidence as one part of a broader validity argument rather than sufficient evidence on its own; see the construct validity guide for the fuller picture.
Is criterion validity the same as predictive validity?
No — predictive validity is one of the two forms criterion validity takes. Criterion validity is the umbrella term for any correlation between an instrument and an independently measured criterion; concurrent and predictive validity distinguish the two forms by the timing of that criterion measurement.







