Measurement, Reliability & Validity
A study is only as good as its measurements. This sub-cluster covers the psychometric concepts that determine whether an instrument actually measures what it claims to: reliability (consistency of measurement, including Cronbach's alpha for internal consistency and test-retest reliability for stability over time) and validity (whether an instrument measures the intended construct, including content, construct, and criterion validity). It also covers instrument development and adaptation — how existing scales are chosen, modified, or built from scratch, and what evidence is needed to justify that choice to reviewers and funders.
Guides
Test-Retest Reliability: Choosing the Retest Interval, and Checking for Drift
The retest interval is the design decision a test-retest coefficient cannot recover from. What the interval trades off, what COSMIN actually asks about it, which ICC form a test-retest design needs, and why only a mean difference or Bland-Altman check reveals systematic drift between occasions.
Visual Analogue Scale (VAS): Construction, Scoring and Minimal Important Change
A visual analogue scale is a 100 mm line with two anchors and no marks in between, scored by measured distance. This guide covers the construction rules that keep it valid, why VAS and NRS scores are not interchangeable, digital-versus-paper equivalence, and the published minimal important change values with the population, baseline severity and derivation method attached to each.
Inter-Rater Reliability: Choosing the Right Coefficient
A selection table from data type and rater design to the right agreement statistic – Cohen’s and weighted kappa, Fleiss’ kappa, ICC, Krippendorff’s alpha and Gwet’s AC1 – with the kappa paradox worked through three 2×2 tables that share identical 95% agreement.
Average Variance Extracted (AVE): Calculation, the 0.50 Threshold, and Discriminant Validity
AVE is the mean proportion of indicator variance a construct explains rather than error. This guide gives the calculation, the origin and conventional status of the 0.50 threshold, and the Fornell-Larcker and HTMT discriminant-validity tests that consume it, including why the methodological literature has moved away from Fornell-Larcker.
Minimal Detectable Change (MDC): SEM-Based Calculation and the MDC-vs-MCID Decision Rule
What MDC is, the three routes to a standard error of measurement and the assumptions each carries, the MDC95 = 2.77 x SEM calculation worked end to end, the individual-versus-group distinction, and the four-cell decision rule for judging a change against both MDC and MCID.
Cronbach’s Alpha: Reliability & Interpretation Guide
What Cronbach’s alpha does and does not tell you about a scale: the formula and its assumptions, why a high alpha is not evidence of unidimensionality, what the benchmark thresholds are worth, and how to report reliability in APA 7th-edition format.
Criterion Validity: Concurrent and Predictive Evidence Explained
Criterion validity explained: concurrent vs. predictive evidence, how to choose a defensible criterion, criterion contamination, attenuation, and a worked correlation example.
Content Validity: Does Your Instrument Cover the Whole Construct?
Content validity is whether an instrument’s items adequately sample the full construct domain, as judged by expert panels. Covers the expert-panel process, the content validity ratio (CVR), the content validity index (CVI), and the distinction from face validity, each worked through a hand-calculated illustrative example.
Response Bias: The Main Types and How to Design Against Them
Response bias is systematic, direction-consistent error in self-report data. This guide covers the five main types — acquiescence, extreme responding, social desirability, recall bias, and order effects — with the specific design fix for each.
The Hawthorne Effect: Why Being Watched Changes the Data
The Hawthorne effect is the tendency for people to change behavior because they know they are being studied. Learn the Western Electric study history, what the Levitt & List and Jones re-analyses actually found, and how to detect and design against reactivity.
Social Desirability Bias: Why Respondents Tell You What You Want to Hear
What social desirability bias is, the self-deception vs. impression-management mechanisms behind it, where it hits hardest, and the countermeasures (indirect questioning, list experiments, randomized response, self-administration) that reduce it.
Psychometrics: How Researchers Measure Things You Can’t Observe
How psychometrics turns unobservable constructs into reliable, valid scores: the construct-to-score pipeline, classical test theory vs. IRT, and how instruments are built and evaluated.
Levels of Measurement: Nominal, Ordinal, Interval and Ratio
A comparison of the four levels of measurement (nominal, ordinal, interval, ratio), the operations and statistics each one permits, and a step-by-step flowchart for classifying any variable.
Construct Validity: Definition, Evidence Types, and Threats
Construct validity is whether an instrument truly measures the theoretical construct it claims to. This guide covers the Messick/AERA-APA-NCME evidence-argument framing, convergent, discriminant, nomological, known-groups and factorial evidence, the multitrait-multimethod matrix, AVE/HTMT, and the two core threats: construct underrepresentation and construct-irrelevant variance.
The Observer Effect in Research: Reactivity, Not Physics
The observer effect (reactivity) is behavior change caused by awareness of being studied. Learn how it differs from the Hawthorne effect, demand characteristics, social desirability bias, and observer bias, where it bites hardest, and how to mitigate it responsibly.
Accuracy vs Precision in Measurement
Accuracy is closeness to the true value (systematic error); precision is closeness of repeated measurements to each other (random error). How to tell them apart, why “precise but inaccurate” is the dangerous case, and how each connects to reliability, validity, calibration, and diagnostic accuracy.
Reliability in Research: What It Means and How to Assess It
What reliability means in research measurement, how it differs from validity, classical test theory, the four types of reliability, what degrades it, attenuation, and how much reliability is enough.
Intraclass Correlation Coefficient (ICC): Forms, Interpretation, and How to Report It
What the ICC measures, the ICC(1,1)/(2,1)/(3,1) forms and when to use each, Koo & Li interpretation benchmarks, ICC vs. kappa/Pearson r/Bland-Altman/Cronbach’s alpha, sample size, and how to report results.
Types of Validity in Research: Measurement Validity vs Design Validity
Measurement validity (face, content, construct, criterion) and design validity (internal, external, statistical-conclusion) are two distinct families constantly conflated under one word. This guide separates them and covers the validity-reliability relationship.
Cronbach’s Alpha: What It Measures, How to Interpret It, and When to Use Omega Instead
Cronbach’s alpha measures internal consistency, not unidimensionality or reliability in general — the most common misreading of the statistic. This guide covers interpretation, why the 0.7 threshold is misapplied, how adding items inflates alpha, when McDonald’s omega is the better choice, and how to report alpha correctly in a methods section.








