Skip to main content
v2026.11,610 entries · CC-BY 4.0

Psychometrics: How Researchers Measure Things You Can’t Observe

How psychometrics turns unobservable constructs into reliable, valid scores: the construct-to-score pipeline, classical test theory vs. IRT, and how instruments are built and evaluated.

Ask about Psychometrics: How Researchers Measure Things You Can’t Observe

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Psychometrics is the branch of research methodology concerned with measuring things that cannot be observed directly — anxiety, research capacity, patient-reported pain, organisational trust, scientific literacy. None of these can be read off an instrument the way temperature or mass can. Psychometrics is the discipline that turns an unobservable idea (a construct) into a number a researcher can actually analyse, and it supplies the tools for checking whether that number is trustworthy.

This page is the overview for CASRAI’s measurement-and-validity content: it lays out the construct-to-score pipeline, the two properties every instrument is judged on (reliability and validity), and links out to the deeper, single-topic guides on each specific method and statistic.

What Psychometrics Actually Covers

Psychometrics sits at the intersection of measurement theory and statistics. It is not the same as “doing statistics on survey data” — a researcher can run a perfectly valid t-test on numbers that were never checked for whether they actually measure the construct they claim to. Psychometrics is specifically the set of methods for answering two prior questions:

  • Reliability: if I measured this again, under similar conditions, would I get a consistent result?
  • Validity: am I actually measuring the construct I intend to measure, or something else that happens to correlate with it?

Both questions apply to any instrument that stands between a researcher and a latent construct: a survey scale, a coded observation protocol, a diagnostic interview, a standardized test, or an automated sensor reading interpreted as a proxy for behaviour.

The Core Idea: From Construct to Score

Every psychometric instrument follows the same chain, and most measurement problems in a study trace back to a weak link somewhere in it:

  1. Construct — the abstract thing you actually care about (e.g., “research self-efficacy”).
  2. Operational definition — a concrete, observable stand-in for the construct that specifies exactly how it will be measured.
  3. Items — individual questions, tasks, or observation codes intended to sample the construct.
  4. Scale — the rule for combining item responses into a score (summing a Likert set, weighting items, latent-trait scoring).
  5. Score — the number that goes into the analysis, and the point at which reliability and validity evidence has to already be in hand.

A score is only as good as the weakest link in that chain. A perfectly reliable instrument that measures the wrong construct is worthless for its intended purpose; a construct with a strong theoretical definition but items that don’t actually sample it will not produce reliable or valid scores no matter how sophisticated the downstream statistics are.

A Brief History

Psychometrics emerged from late-19th and early-20th century work on individual differences, notably Francis Galton’s studies of human variation and Charles Spearman’s early-20th-century work on correlation and general intelligence, which introduced the statistical foundation later formalized as classical test theory. Classical test theory treats an observed score as the sum of a “true score” and random measurement error, and it underlies most of the reliability statistics still used today, including Cronbach’s alpha (introduced by Lee Cronbach in 1951). In the mid-20th century, item response theory (IRT) developed as an alternative framework that models the probability of a given response as a function of both item properties and the respondent’s underlying trait level, rather than treating all items as interchangeable. Professional standards for test development and use are maintained jointly by the American Educational Research Association (AERA), the American Psychological Association (APA), and the National Council on Measurement in Education (NCME) in the Standards for Educational and Psychological Testing, periodically revised and widely treated as the field’s reference framework.

The Two Pillars: Reliability and Validity

Reliability: Consistency of Measurement

Reliability asks whether a measurement is repeatable. A bathroom scale that reads three different values for the same object in the same minute is unreliable, regardless of whether it’s accurate on average. In psychometrics, reliability is estimated in several distinct ways, and they are not interchangeable — each targets a different source of inconsistency:

For a fuller treatment of what reliability means, how it’s estimated, and how much is “enough,” see Reliability in Research: What It Means and How to Assess It.

Validity: Measuring the Right Thing

Validity asks whether the instrument measures the construct it claims to, not merely whether it produces consistent numbers. An instrument can be highly reliable and still invalid — a scale that consistently measures test-taking speed rather than the intended construct is reliable but not valid for its stated purpose. Modern measurement theory treats validity as a unified concept built from several kinds of evidence rather than several separate “types” of validity, though the classic vocabulary (content, criterion, construct) is still the most common way researchers describe that evidence. See Types of Validity in Research for the full breakdown, and Construct Validity for the specific evidence type most often at stake when an instrument is new or adapted. Whether findings generalize beyond the specific sample and setting studied is a related but distinct question, covered in Internal vs. External Validity.

Classical Test Theory vs. Item Response Theory

Two statistical frameworks dominate instrument development and scoring:

  • Classical test theory (CTT) models an observed score as true score plus error, estimates reliability at the level of the whole scale, and assumes items are roughly interchangeable indicators of the construct. It is simpler, requires smaller samples, and remains the default in most applied research.
  • Item response theory (IRT) models the probability of a specific response as a function of the respondent’s trait level and item-level parameters (difficulty, discrimination, and, in some models, guessing). It supports item-level analysis, computerized adaptive testing, and comparing scores across different item sets, but requires larger samples and more specialized software.

Most research administration, survey, and evaluation contexts use classical test theory because sample sizes and resources rarely justify IRT; large-scale standardized testing programs are where IRT is most commonly used.

How an Instrument Actually Gets Built

Instrument development follows a broadly consistent sequence in the measurement literature, whether the end product is a new survey scale, an adapted questionnaire, or a coded observation protocol:

  1. Define the construct. Write an explicit theoretical definition before drafting a single item — see Research Constructs: Definition and Examples.
  2. Generate an item pool. Draft more items than needed, since weak items will be dropped later.
  3. Expert or content review. Subject-matter experts rate whether items adequately sample the construct — the main source of content-validity evidence.
  4. Pilot test. Administer the draft instrument to a sample similar to the intended population.
  5. Item analysis. Drop items with poor item-total correlations, unclear wording, or floor/ceiling problems.
  6. Estimate reliability. Compute internal consistency and, where relevant, test-retest or inter-rater reliability on the revised item set.
  7. Gather validity evidence. Compare scores against related and unrelated measures, known groups, or external criteria to build a validity case.
  8. Document norms and administration procedures. Record how the instrument should be scored and interpreted so future users apply it consistently.

Common Measurement Pitfalls

  • Construct underrepresentation — the item set is too narrow to capture the full construct as theoretically defined.
  • Construct-irrelevant variance — scores are influenced by something other than the construct, such as reading level, social desirability, or response format.
  • Ceiling and floor effects — most respondents cluster at the top or bottom of the scale, compressing the range and weakening both reliability and discrimination between respondents.
  • Confusing accuracy with precision, or reliability with validity. These are related but distinct properties; see Accuracy vs. Precision in Measurement for the accuracy/precision distinction specifically.
  • Reactivity — the act of measuring changes the behaviour being measured, discussed in The Observer Effect in Research.

Where Psychometrics Shows Up in Research Practice

Researchers encounter psychometric questions well outside psychology departments: patient-reported outcome measures in clinical trials, research-capacity and self-efficacy surveys in institutional evaluation, competency assessments in training programs, and any locally-built questionnaire used to justify a finding. Reviewers and funders increasingly expect a stated reliability estimate and a validity rationale — even a brief one — whenever a study relies on a non-standardized instrument, rather than an assumption that a face-valid-looking questionnaire is automatically adequate.

A Reliability and Validity Map of This Guide Series

These pages, together, form CASRAI’s coverage of measurement, reliability, and validity in research:

Frequently Asked Questions

Is psychometrics only relevant to psychology research?

No. Psychometric principles apply to any research that uses an instrument to measure an unobservable construct — patient-reported outcomes in clinical trials, survey-based program evaluation, education and training assessments, and organizational or institutional research all rely on the same reliability and validity logic, even when the researchers involved don’t identify as psychometricians.

What’s the difference between psychometrics and statistics?

Statistics is the broader toolkit for analysing data once it exists. Psychometrics is specifically about whether the measurement process that produced the data is trustworthy — reliable and valid — before any downstream statistical test is run. A study can use sophisticated statistics on data from a poorly validated instrument and still produce misleading conclusions.

Do I need to run a full psychometric validation for a survey I built for one study?

At minimum, report an internal-consistency estimate (such as Cronbach’s alpha) for any multi-item scale and describe the content-validity rationale for the items — how they were derived from the construct definition. A full validation program (multiple samples, convergent/discriminant evidence, factor analysis) is expected for instruments intended for reuse or publication as a standalone measure, but even a single-study instrument should show basic reliability and a defensible link to its construct definition.

Can an instrument be reliable but not valid?

Yes, and this is one of the most commonly misunderstood points in measurement. Reliability is necessary but not sufficient for validity — an instrument can produce highly consistent scores that nonetheless measure the wrong thing. The reverse is not true: an instrument cannot be valid without also being reasonably reliable, since unreliable scores cannot consistently track a real construct.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →