Skip to main content
v2026.11,610 entries · CC-BY 4.0

Criterion Validity: Concurrent and Predictive Evidence Explained

Criterion validity explained: concurrent vs. predictive evidence, how to choose a defensible criterion, criterion contamination, attenuation, and a worked correlation example.

Ask about Criterion Validity: Concurrent and Predictive Evidence Explained

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Criterion validity asks a narrower, more testable question than construct validity: does this instrument’s score correlate with an independently measured outcome that the field already accepts as meaningful? That outcome is the “criterion.” Criterion validity comes in two forms depending on when the criterion is measured relative to the instrument: concurrent validity (criterion measured at roughly the same time) and predictive validity (criterion measured later). This guide covers both forms in depth, how to choose and evaluate a criterion, the criterion-contamination pitfall, and how unreliability in either measure attenuates the correlation you observe — with a worked numerical example. For how criterion validity fits alongside construct and content validity, see Types of Validity in Research and the dedicated construct validity guide.

Concurrent Validity vs. Predictive Validity

Both are forms of criterion-related validity evidence. The distinction is entirely about timing, not about the underlying logic of the correlation.

Dimension Concurrent Validity Predictive Validity
When the criterion is measured At roughly the same time as the instrument At a defined later time point
Typical purpose Substitute a shorter, cheaper, or less invasive instrument for an established one Forecast a future state, behaviour, or outcome
Typical study design Cross-sectional: both measures collected in the same session or same short window Longitudinal or prospective: instrument administered first, criterion collected after a follow-up interval
Worked example A brief self-report depression screener administered alongside a structured clinical interview in the same visit A graduate-admissions test administered at application, correlated with GPA at degree completion
Main practical risk The “gold standard” comparison instrument may itself be imperfect or costly to justify as a criterion Attrition, cohort changes, or a long follow-up window can erode the correlation for reasons unrelated to the instrument

Some methods texts use “concurrent” and “predictive” as if they were separate validity types; the more common current framing, consistent with the Standards for Educational and Psychological Testing (AERA/APA/NCME), treats both as evidence based on relations to a criterion — a single validity argument, gathered at two different points in the measurement timeline.

What Counts as a Valid Criterion

A criterion is only useful if it is itself a defensible, independently obtained measure of the outcome the instrument is meant to predict or track — not simply another self-report version of the same instrument. Common criterion types:

Criterion type Example instrument being validated against it Common limitation
Structured clinical interview or diagnosis Brief self-report symptom screener Interview itself has imperfect inter-rater reliability, which caps the observed correlation
Academic or job-performance record (GPA, supervisor rating, output metric) Admissions test, aptitude test, competency assessment Performance records are influenced by many factors besides the trait the instrument measures
Biomarker or laboratory value Self-report health-behaviour or symptom scale Biomarkers can lag or lead the construct of interest, weakening a same-time correlation
Documented behavioural outcome (readmission, attrition, incident report) Risk-screening or readiness instrument Base rates are often low, which restricts the correlation that is mathematically achievable
Licensure, certification, or credentialing outcome Pre-training aptitude or knowledge test Pass/fail outcomes are binary, which typically lowers correlation coefficients relative to continuous criteria

Interpreting the Correlation Coefficient

Criterion validity is usually reported as a single correlation coefficient (Pearson’s r, or a point-biserial/tetrachoric variant when the criterion is binary). General behavioural-science conventions for interpreting r — commonly attributed to Cohen (1988) — treat 0.10 as small, 0.30 as medium, and 0.50 as large. Applied criterion-validity coefficients in real research and assessment settings routinely fall in the 0.20–0.50 range and are still considered practically useful, for two structural reasons this guide covers below: no single criterion fully captures a complex construct (criteria are themselves imperfect), and unreliability in either measure mechanically attenuates the observed correlation regardless of how valid the instrument truly is.

Observed r Conventional label Practical read for criterion validity
< 0.20 Small / weak Investigate whether the criterion is poorly matched, contaminated, or unreliable before concluding the instrument itself lacks validity
0.20 – 0.35 Small-to-medium Common and often acceptable for screening tools used alongside other information, not as a sole decision criterion
0.35 – 0.50 Medium-to-large Typical target range for an instrument intended to substitute for, or meaningfully predict, an established criterion
> 0.50 Large Strong evidence, but check for criterion contamination (below) before treating it as confirmatory

Choosing a Criterion

Before collecting any data, the criterion itself needs the same scrutiny as the instrument under evaluation. A defensible criterion should be:

  • Relevant — it should measure the actual outcome of interest, not a convenient proxy that happens to be available.
  • Reliable — an unreliable criterion caps the correlation you can ever observe, independent of the instrument’s quality (see Attenuation, below).
  • Independent — collected and scored without knowledge of the instrument’s scores, to avoid criterion contamination.
  • Free of restricted range — if the sample used to validate the instrument has already been pre-selected on the criterion (e.g., validating an admissions test only on students who were admitted), the correlation will be mechanically suppressed.
  • Practically obtainable at the scale the instrument will actually be used — a criterion that requires a resource-intensive gold-standard assessment for every case undermines the rationale for a shorter instrument in the first place.

Methods literature sometimes calls the gap between an ideal criterion and any available real-world measure the “criterion problem” — every available criterion is itself deficient or contaminated to some degree, so criterion-validity evidence should always be read as evidence relative to a specific, named criterion, not as an absolute property of the instrument.

Criterion Contamination

Criterion contamination occurs when the criterion measure is influenced by something other than the outcome it is supposed to represent — most commonly, when whoever is scoring or rating the criterion has knowledge of the instrument’s scores. If a supervisor rating “job performance” (the criterion) already knows a candidate’s test score (the predictor), that knowledge can bias the rating upward or downward independent of actual performance, inflating the observed correlation without reflecting any real predictive relationship. This concern traces back to the classic distinction in Cronbach and Meehl’s foundational 1955 treatment of validity: a criterion can be deficient (missing parts of the true outcome), contaminated (including things it shouldn’t), or both at once.

Practical safeguards against contamination:

  • Blind criterion raters to the instrument’s scores wherever feasible.
  • Use criteria drawn from records or systems that were generated independently of the study (e.g., administrative data, existing clinical records) rather than purpose-collected ratings from people aware of the study hypothesis.
  • Separate the personnel or systems that administer the instrument from those that score the criterion.

Attenuation: Why Observed Correlations Understate True Validity

Every correlation coefficient is capped by the reliability of the two measures being correlated. If either the instrument or the criterion has measurement error — and every real-world measure does, to some degree — the observed correlation will be lower than the “true” relationship between the underlying constructs. This is called attenuation, and it is one of the oldest results in psychometrics, formalised by Charles Spearman in 1904 as the correction for attenuation:

rcorrected = rxy ÷ √(rxx × ryy)

where rxy is the observed correlation between the instrument and the criterion, rxx is the reliability of the instrument (e.g., Cronbach’s alpha — see the Cronbach’s alpha guide), and ryy is the reliability of the criterion (e.g., inter-rater reliability for a rated criterion; see Test-Retest vs. Inter-Rater Reliability).

Worked example

The following is an illustrative calculation using hypothetical numbers, not a reported finding from any specific study. Suppose a new brief screening instrument is validated concurrently against a structured clinical interview. The observed correlation between the two is rxy = 0.40. Internal-consistency reliability (Cronbach’s alpha) for the screening instrument is rxx = 0.85. Inter-rater reliability for the clinical interview is ryy = 0.75.

Applying the formula: rcorrected = 0.40 ÷ √(0.85 × 0.75) = 0.40 ÷ √0.6375 = 0.40 ÷ 0.798 ≈ 0.50.

The disattenuated correlation (0.50) is noticeably higher than the observed correlation (0.40) — not because the instrument changed, but because correcting for measurement error in both variables reveals a stronger underlying relationship than the raw correlation showed. This is why a modest observed criterion-validity coefficient should prompt a check of both measures’ reliability before it is read as a weak validity finding. It is also why the correction should be used to interpret evidence, not to replace the reported observed correlation in a published validity argument — both figures are normally reported together, with the disattenuated value labelled explicitly as a correction.

Criterion Validity vs. Construct and Content Validity

Criterion, construct, and content validity are not competing techniques; they are different kinds of evidence for the same underlying question — does this instrument measure what it claims to? Criterion validity is the most directly testable of the three because it collapses to a single correlation coefficient, but a correlation with one flawed or narrow criterion is not sufficient evidence on its own. The Standards for Educational and Psychological Testing frames validity as a unified concept supported by multiple, converging lines of evidence: content (does the instrument’s coverage match the domain?), internal structure and relations to other variables (does it behave the way the underlying construct should behave, covered in the construct validity guide), and criterion-related evidence (does it correlate with an accepted external outcome, the subject of this guide). See Types of Validity in Research for how all of these fit together, and Reliability vs. Validity and Internal vs. External Validity for two related distinctions that are commonly confused with criterion validity but answer different questions.

Frequently Asked Questions

What is the difference between concurrent and predictive validity?

Both correlate an instrument’s scores with an external criterion. Concurrent validity measures the criterion at roughly the same time as the instrument (often used to justify a shorter or cheaper substitute for an established measure); predictive validity measures the criterion at a later time point (used to justify the instrument as a forecasting tool).

How high should a criterion validity correlation be?

There is no universal cutoff. Correlations of 0.20–0.35 are common and often considered practically useful for screening purposes, while 0.35–0.50 is a more typical target when an instrument is meant to substitute for or meaningfully predict an established criterion. Any interpretation should account for the reliability of both the instrument and the criterion, since unreliability mechanically suppresses the observed correlation (see Attenuation, above).

What is criterion contamination?

It is bias introduced into the criterion measure by something other than the true outcome — most often, a rater’s knowledge of the instrument’s scores influencing how they score the criterion. It inflates the observed correlation and should be guarded against by blinding criterion raters to instrument scores wherever possible.

Can an instrument have criterion validity without construct validity?

In principle a scale can correlate with a criterion for reasons unrelated to the construct it claims to measure — for example, both may share an unrelated confound. This is exactly why the current testing standards treat criterion-related evidence as one part of a broader validity argument rather than sufficient evidence on its own; see the construct validity guide for the fuller picture.

Is criterion validity the same as predictive validity?

No — predictive validity is one of the two forms criterion validity takes. Criterion validity is the umbrella term for any correlation between an instrument and an independently measured criterion; concurrent and predictive validity distinguish the two forms by the timing of that criterion measurement.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →