“Validity” in research methods is not one concept — it is two, and the constant confusion between them is the single most common validity mistake in research writing. One family asks whether an instrument measures what it claims to measure (measurement validity: face, content, construct, criterion). The other asks whether a study design supports the causal or generalizing claim it makes (design validity: internal, external, statistical-conclusion, and construct validity of the cause-and-effect operationalization). Both families share the word “construct validity,” which is itself a frequent source of confusion — they are related but answer different questions. This guide treats each family on its own terms, then explains how they connect to reliability.
Two families, one word
| Family | Core question | Applies to | Types |
|---|---|---|---|
| Measurement validity | Does this instrument actually measure the construct it claims to measure? | A scale, test, questionnaire, or coding scheme | Face, content, construct, criterion (predictive/concurrent) |
| Design validity | Can this study’s design support the causal or generalizing claim being made? | The overall study — its sampling, procedure, and analysis | Internal, external, statistical-conclusion, construct (of the cause-effect operationalization) |
A study can have a well-validated instrument and a weak design (e.g., a validated depression scale used in a study with no control group, so causal claims about what reduced depression scores are unsupported). Equally, a study can have a strong design and a poorly validated instrument (a randomized controlled trial that measures its outcome with a homemade, unvalidated questionnaire). Reporting one without the other is incomplete.
Family 1: Measurement validity — does the instrument measure the construct?
Measurement validity is evaluated for a specific instrument in a specific context of use; a scale is not “valid” in the abstract, it is valid for measuring a particular construct in a particular population for a particular purpose.
Face validity
Face validity is the weakest and most subjective form: whether an instrument appears, on the surface, to measure what it claims to measure — to test-takers, reviewers, or other non-expert observers. A job-satisfaction survey that asks “How satisfied are you with your pay?” has face validity for measuring pay satisfaction; it looks like it’s asking the right thing. Face validity is a useful sanity check and affects respondent buy-in and completion rates, but it is not evidence of psychometric quality — an instrument can look right and still fail to measure the intended construct, or can measure a real construct without looking like it does (some validated personality and psychopathology scales deliberately obscure their intent to reduce social-desirability bias).
Content validity
Content validity asks whether an instrument’s items adequately sample the full domain of the construct being measured, typically judged by subject-matter experts rather than lay reviewers. A statistics-course final exam has content validity if its questions span the topics actually taught (descriptive statistics, hypothesis testing, regression) rather than over-sampling one unit and ignoring others. Content validity is established through systematic expert review — a table of specifications mapping items to domain content, expert rating panels, or a Delphi-style consensus process — not through post-hoc statistical analysis of responses.
Construct validity (measurement sense)
Construct validity — the term Lee Cronbach and Paul Meehl formalized in their 1955 paper “Construct Validity in Psychological Tests” — asks whether an instrument truly measures the theoretical construct it is designed to measure, and is evaluated through a pattern of evidence rather than a single test. Two commonly cited sub-components:
- Convergent validity: scores on the instrument correlate with scores on other measures of the same or a theoretically related construct (a new anxiety scale should correlate with an established, validated anxiety scale).
- Discriminant validity: scores on the instrument do not correlate strongly with measures of theoretically distinct constructs (the same anxiety scale should not correlate highly with a measure of, say, extraversion).
Construct validity is the broadest and most theory-driven form of measurement validity — it is often described as encompassing content and criterion evidence within an overall validity argument, a framing formalized in the Standards for Educational and Psychological Testing (AERA, APA, NCME).
Criterion validity: predictive and concurrent
Criterion validity asks whether scores on an instrument correlate with an external, independently measured outcome (the “criterion”). It splits into two sub-types by timing:
- Predictive validity: the instrument’s scores predict a future outcome. A graduate-admissions test has predictive validity if scores correlate with later graduate GPA.
- Concurrent validity: the instrument’s scores correlate with a criterion measured at roughly the same time. A new, shorter depression screening tool has concurrent validity if its scores correlate with a longer, established clinical interview administered around the same session.
Criterion validity is the most directly testable form because it produces a single correlation coefficient — but it depends entirely on having a defensible, independently measured criterion. Weak criterion validity often traces back to a poorly chosen or poorly measured criterion, not necessarily to the instrument itself.
Family 2: Design validity — can the study support its claim?
The internal/external/statistical-conclusion/construct framework comes from Donald Campbell and Julian Stanley’s original taxonomy of threats to experimental validity, most fully developed by William Shadish, Thomas Cook, and Donald Campbell in Experimental and Quasi-Experimental Designs for Generalized Causal Inference (2002). It evaluates the study as a whole — its design, sample, procedure, and analysis — not any single measurement instrument.
Internal validity
Internal validity asks whether the study design supports the claim that the independent variable, and not some confound, caused the observed change in the dependent variable. Threats include history (an external event coinciding with the study), maturation (participants change over time regardless of treatment), selection (non-equivalent groups from the start), regression to the mean, attrition, and testing effects (repeated measurement changing responses). Random assignment to conditions is the strongest single design feature for protecting internal validity, because it makes confounds equally likely across groups by chance.
External validity
External validity asks whether a causal finding generalizes beyond the specific people, settings, treatments, and times studied — to other populations, real-world settings, or variations in how the treatment is delivered. A drug trial conducted only on healthy young male volunteers may have strong internal validity (a well-controlled RCT) but limited external validity for predicting effects in older patients with comorbidities. Internal and external validity often trade off: the tight control that maximizes internal validity (a lab setting, a narrow sample) can reduce how representative the findings are of real-world conditions.
Statistical-conclusion validity
Statistical-conclusion validity asks whether the statistical inferences about the relationship between variables are warranted — whether the analysis has adequate statistical power, appropriate assumptions, correctly estimated effect sizes, and controlled Type I/Type II error rates. Common threats include underpowered samples (a true effect exists but the study lacks the power to detect it — see CASRAI’s guide to power analysis and sample size calculation), violated statistical assumptions, and fishing/error-rate inflation from testing many comparisons without correction.
Construct validity of the cause-and-effect operationalization
Within the Shadish-Cook-Campbell framework, construct validity has a design-specific meaning distinct from the measurement-validity sense above: it asks whether the specific operationalizations used in the study — the actual manipulation delivered and the actual outcome measured — correctly represent the abstract theoretical constructs the researcher claims to be studying. A study claiming to test “the effect of stress on performance” has weak construct validity if its “stress” manipulation is actually measuring time pressure specifically, or its “performance” outcome is a single easy quiz rather than a construct-relevant task. This is why the same term appears in both families: a design’s causal claim depends on both the manipulation and the outcome measure being valid operationalizations of their intended constructs — the two families connect here rather than being fully separate.
How the two families interact in practice
A concrete way to see the split: imagine a study testing whether a training program improves “critical thinking.”
- Measurement validity question: Does the critical-thinking test used as the outcome actually measure critical thinking (construct validity), cover the relevant sub-skills (content validity), and predict real-world critical-thinking performance (criterion validity)?
- Design validity question: Does the study’s design (random assignment, control group, adequate sample size, generalizable sample) support the claim that the training — rather than a confound — caused any observed change (internal validity), that the result would replicate elsewhere (external validity), and that the statistical test was properly powered and specified (statistical-conclusion validity)?
A methods section that only addresses one family is incomplete. Peer reviewers and methodologists routinely flag studies that validate their instrument thoroughly but never address confounds, or that have an airtight randomized design paired with an unvalidated, ad hoc outcome measure.
Validity and reliability: reliability bounds validity but does not imply it
Reliability — an instrument’s consistency, whether across items (internal consistency, commonly estimated with Cronbach’s alpha), raters (inter-rater reliability), or repeated administrations (test-retest reliability) — is a necessary but not sufficient condition for validity. An instrument that produces inconsistent, noisy scores cannot be a valid measure of anything, because validity requires that scores reliably track something real. But a highly reliable instrument can still measure the wrong thing consistently: a bathroom scale that reads 10 pounds heavy every single time is perfectly reliable (consistent) but not a valid measure of true body weight. The reverse relationship does not hold either — validity places an upper bound on how validity coefficients (such as criterion-validity correlations) can perform given a measure’s reliability, since measurement error attenuates observed correlations. In practice: check reliability first, because a measure that fails reliability cannot pass validity testing, but never stop at reliability alone and report it as if it settled the validity question.
Frequently asked questions
Is construct validity the same thing whether it’s discussed as a measurement property or a design property?
No — this is the exact conflation this guide addresses. As a measurement property (Cronbach & Meehl’s original sense), construct validity asks whether an instrument measures its intended theoretical construct, evaluated via convergent/discriminant evidence. As a design property (Shadish, Cook & Campbell’s sense), it asks whether a study’s specific manipulation and outcome measure correctly operationalize the abstract constructs the causal claim is about. They are related — a design’s construct validity partly depends on its instruments having measurement construct validity — but they are not interchangeable, and a methods section should be specific about which one it means.
Which validity type matters most?
There is no universal ranking; it depends on the study’s purpose. An efficacy trial prioritizes internal validity (isolating the causal effect under controlled conditions); a pragmatic effectiveness trial or a policy-relevant field study prioritizes external validity (generalizability to real-world conditions); an instrument-development study prioritizes construct and content validity; any quantitative study depends on statistical-conclusion validity as a baseline, since a valid design paired with an underpowered or misspecified analysis still produces unreliable conclusions.
Can a study have high internal validity but low external validity, or vice versa?
Yes, and this trade-off is one of the most cited tensions in research design. Tightly controlled laboratory experiments with homogeneous samples typically maximize internal validity at the cost of external validity (results may not generalize to more diverse, real-world settings). Large-scale, real-world field studies typically improve external validity but often sacrifice some internal validity, since fewer confounds can be controlled outside a lab.
How is content validity actually established, since it is not a single statistic?
Common approaches include a formal table of specifications mapping test or survey items against the full content domain, expert-panel review with structured rating scales (e.g., item-level relevance ratings later summarized as a content validity index), and cognitive interviewing with representative respondents to confirm items are interpreted as intended. Unlike criterion validity, content validity is a judgment-based process rather than a single correlation coefficient.
Related CASRAI resources
- Research Methods pillar — the full hub for study design, sampling, analysis, and measurement content.
- Cronbach’s Alpha: What It Measures, How to Interpret It, and When to Use Omega Instead — the internal-consistency reliability estimate most directly related to the validity/reliability discussion above.
- Regression Analysis: Assumptions, Interpretation, and How to Report It — relevant to statistical-conclusion validity threats around assumption violations.
- Power Analysis and Sample Size Calculation — the core defense against underpowered statistical-conclusion validity threats.







