Reliability is the consistency of a measurement — whether a scale, test, coding scheme, or instrument produces the same result under the same conditions. It is one of the two properties, alongside validity, that any measurement used in research has to satisfy before its scores can be trusted, and it is routinely confused with validity even though the two answer completely different questions. This page defines reliability, works through classical test theory (the statistical model reliability is built on), covers the four main types researchers report and where CASRAI treats each in depth, explains what degrades reliability and why that costs a study statistical power through attenuation, and covers how reliability should actually be reported.
What Reliability Means in Research
A measurement is reliable to the extent that repeating it — on the same subjects, under the same conditions — produces the same result. Reliability says nothing about whether the measurement is capturing the right thing, only about whether it is capturing something consistently. A bathroom scale that reads two kilograms heavy every single time is highly reliable: step on it repeatedly and it gives the same answer. It is also wrong every time, which is the entire point of separating reliability from validity.
The classic illustration is a dartboard. Darts clustered tightly together, wherever they land, show high reliability — the thrower is consistent. Darts clustered tightly around the bullseye show both reliability and validity — consistent and accurate. Darts clustered tightly in the upper-left of the board, nowhere near the bullseye, show reliability without validity: a consistently biased measurement. Darts scattered randomly across the board show neither. The concrete research version of that upper-left cluster is a survey item that respondents interpret consistently — producing stable, repeatable scores — but that measures something adjacent to the intended construct (social desirability, for instance, instead of the attitude the item was written to capture). The scores are dependable; they are dependably measuring the wrong thing.
Because reliability is a precondition, not a guarantee: an instrument that produces noisy, inconsistent scores cannot be valid, since validity requires that scores reliably track something real. But passing a reliability check only clears the floor. See CASRAI’s guide to types of validity in research for the full validity framework and how the two properties interact in a methods section.
Classical Test Theory: True Score Plus Error
Reliability theory rests on a simple statistical model, classical test theory (CTT), formalized in its modern form by Lord and Novick’s 1968 Statistical Theories of Mental Test Scores: any observed score is modeled as the sum of two unobservable components.
Observed score (X) = True score (T) + Error (E)
The true score is the score a person would get if measurement were perfect — in practice, it is treated as the average score that person would obtain across an infinite number of independent administrations of the same measure. Error is everything else: random noise from item ambiguity, rater inconsistency, momentary attention lapses, or testing conditions, assumed under CTT to be random and uncorrelated with the true score (unlike bias, which is systematic and does not average out — the scale that always reads two kilograms heavy has a bias problem, not a reliability problem).
Reliability, formally, is defined as the ratio of true-score variance to total observed-score variance:
Reliability = Variance(True score) / Variance(Observed score) = Variance(T) / [Variance(T) + Variance(E)]
A reliability coefficient of 0.80 means that, under this model, 80% of the variance in the observed scores reflects genuine differences between subjects (true-score variance), and the remaining 20% is measurement error. Every reliability statistic in common use — Cronbach’s alpha, the intraclass correlation coefficient, Cohen’s or Fleiss’ kappa — is a different way of estimating this same underlying ratio for a different source of error.
The Four Types of Reliability
“Reliability” is not one statistic; it is a family of properties, each isolating a different source of measurement error. A study can report strong reliability on one type and be untested, or weak, on another — reporting one type as if it settled the general question of “is this measure reliable” is a common and avoidable overstatement.
Internal Consistency
Internal consistency asks whether the items within a single multi-item scale, administered once, correlate with one another — whether they appear to be measuring the same underlying construct. It is estimated with Cronbach’s alpha (the most widely reported reliability statistic in the social and health sciences), McDonald’s omega, or split-half correlation. See CASRAI’s guide to Cronbach’s alpha for how it’s calculated, why the conventional 0.70 threshold is routinely over-applied, and when omega is the better-justified choice.
Test-Retest Reliability
Test-retest reliability asks whether the same instrument, administered to the same people at two points in time, produces consistent scores — isolating consistency across time rather than across items or raters. The retest interval is a real design decision, not an afterthought: too short and respondents recall their earlier answers, inflating the estimate; too long and the underlying trait may have genuinely changed, deflating it for reasons that have nothing to do with measurement error. See CASRAI’s comparison of test-retest vs. inter-rater reliability for the statistics used (Pearson’s r, the intraclass correlation coefficient, or Cohen’s kappa for categorical data) and typical interval conventions.
Inter-Rater / Inter-Observer Reliability
Inter-rater (or inter-observer) reliability asks whether two or more independent raters, coders, or observers reach the same result when scoring the same material at essentially the same time — isolating consistency across people rather than across time. It is standard practice wherever a human judgment step sits between raw data and a recorded score: clinical diagnosis, behavioral observation, qualitative coding, and screening decisions in evidence synthesis all require it. Low inter-rater reliability usually points to a problem with the coding scheme or rater training, not the underlying construct. Continuous ratings are typically assessed with the intraclass correlation coefficient (ICC); categorical judgments with Cohen’s kappa (two raters) or Fleiss’ kappa (three or more).
Parallel-Forms (Alternate-Forms) Reliability
Parallel-forms reliability asks whether two different versions of the same instrument — built to the same specification, targeting the same construct at the same difficulty level, but using different specific items — produce equivalent scores when administered to the same people. It is used specifically to avoid the practice effects that complicate test-retest reliability (a respondent can’t simply recall their answers to a different set of items) and is common wherever repeated testing on the same population is expected, such as standardized educational assessment, where the same construct needs to be measured at multiple time points without reusing identical items. The practical cost is that developing and validating two genuinely equivalent forms is considerably more expensive than developing one, which is why parallel-forms reliability is reported far less often than the other three types.
What Degrades Reliability
Reliability is not a fixed property an instrument either has or lacks; it degrades from specific, identifiable sources of error, most of which are addressable in study design rather than inherent to the construct being measured:
- Item ambiguity — vague wording, double-barreled questions, or jargon that different respondents interpret differently introduces noise before any data are even collected. See CASRAI’s guide to questionnaire design for the specific wording failure patterns that most commonly cause this.
- Rater drift — a rater’s application of a coding scheme shifting gradually over the course of a long coding project, so that early and late ratings are no longer produced under equivalent criteria, even without any change in the rater’s training or the coding scheme itself.
- Inconsistent testing conditions — administering an instrument under meaningfully different conditions (noise, time pressure, mode of administration) across subjects or occasions introduces error unrelated to genuine differences between subjects.
- Short scales — a scale built from very few items has less opportunity for item-level noise to average out, which is one reason (though not the only one) that internal-consistency estimates tend to run lower for short scales, all else equal.
- Restricted range — when a sample includes only a narrow slice of the true range on the construct being measured (a homogeneous sample where most people score similarly), reliability estimates — which are correlation-based and therefore sensitive to variance — can appear artificially low even though the instrument functions normally in a more heterogeneous population. A reliability coefficient computed in a restricted-range sample should not be assumed to generalize to a population with a wider range on the construct.
Attenuation: Why Unreliable Measurement Costs You Power
Measurement error does not just add noise in the abstract — it has a specific, predictable statistical consequence: unreliable measures attenuate (weaken) the observed correlations and effect sizes involving them, biasing estimates toward zero. This is why reliability is not merely a reporting formality; it directly affects a study’s ability to detect a real effect.
The relationship is formalized in the correction for attenuation, developed by Charles Spearman in 1904: the observed correlation between two variables is a function of the true correlation between them and the square root of the product of their reliabilities.
robserved = rtrue × √(reliabilityX × reliabilityY)
The practical consequence: if a predictor and an outcome each have a reliability of 0.70, the observed correlation between them can be attenuated to roughly 70% of the true underlying correlation, even when the true relationship is strong. A study using a noisy, low-reliability instrument is therefore not just adding random variability — it is systematically understating the size of real effects and, correspondingly, reducing statistical power to detect them at a given sample size. Improving instrument reliability is, in this sense, a direct and often underused lever for statistical power, alongside the sample-size and effect-size considerations covered in CASRAI’s guide to power analysis and sample size calculation.
How Much Reliability Is Enough?
Conventional thresholds circulate widely — 0.70 as “acceptable,” 0.80 as “good,” 0.90 as “excellent” — and they are useful as shared reference points, but they are conventions, not statistical laws, and treating a single fixed number as a pass/fail bar across every use case is a common overstatement. The threshold that is actually defensible depends heavily on what the score is being used for:
- Group-level research decisions — comparing group means, testing associations between variables, informing policy or theory at an aggregate level — can often tolerate reliability in the 0.70–0.80 range, because random measurement error tends to average out across many observations even as it costs some statistical power per the attenuation relationship above.
- Individual-level decisions — clinical diagnosis, high-stakes assessment, admissions, or any use where a specific person’s score determines a specific consequence for that person — generally demand a substantially higher bar, often 0.90 or above, because measurement error that averages out harmlessly across a group can still meaningfully mislabel or misclassify a specific individual.
This distinction — sometimes described as applying different reliability standards depending on whether a score informs research conclusions about groups versus decisions about individuals — is more consequential than the specific numeric cutoff chosen. A methods section citing “acceptable” reliability without stating which kind of decision the score supports is leaving out the information that actually determines whether the bar was high enough.
How to Report Reliability
A complete, honest reliability report states four things, not just a single coefficient:
- Which type of reliability was assessed — internal consistency, test-retest, inter-rater, or parallel-forms — since each answers a different question and none substitutes for another.
- In which sample — reliability is not a fixed, inherent property of an instrument; it is a property of the scores an instrument produces in a specific sample, under specific conditions. This point is widely misunderstood and worth stating plainly: a scale is not simply “reliable” in the abstract, transportable unchanged to any population. A published reliability coefficient from an instrument’s original validation study describes that study’s sample; a new study using the same instrument in a different population, language, or context should compute and report its own reliability estimate rather than citing only the original figure.
- With a confidence interval — a point estimate alone overstates precision, particularly from a small sample; reporting the interval around it (as recommended, for instance, in Koo and Li’s widely cited guidance on interpreting the intraclass correlation coefficient) lets a reader judge how much the estimate should be trusted.
- What is and is not being claimed — that the scores showed a given level of consistency of the specific type assessed, in this sample, not that the instrument is valid, unidimensional, or reliable in every sense, unless those properties were independently established.
Example phrasing: “Internal consistency for the 8-item scale was acceptable in the current sample (α = 0.81, n = 214). Inter-rater reliability for the coded outcome variable was good (ICC(2,1) = 0.84, 95% CI [0.71, 0.92]) in a subsample of 40 cases double-coded by two raters.” This states the type, the sample, the value with an interval, and keeps each reliability claim scoped to what was actually tested.
Frequently Asked Questions
Is reliability the same thing as validity?
No. Reliability is consistency — whether a measurement gives the same result under the same conditions. Validity is accuracy — whether the measurement actually captures the construct it claims to. A measure can be highly reliable and not valid (a consistently biased instrument); it cannot be valid without also being reliable, since validity requires that scores reliably track something real. See CASRAI’s guide to types of validity in research for the full framework.
What is a good reliability coefficient?
Conventional bands treat roughly 0.70 as acceptable, 0.80 as good, and 0.90 as excellent, but these are widely used conventions rather than fixed statistical standards. The defensible threshold depends on what the score is used for: research decisions about groups can often tolerate somewhat lower reliability than decisions made about specific individuals, which generally demand a substantially higher bar.
Can an instrument be reliable but not valid?
Yes — this is the central point of the reliability/validity distinction. An instrument that produces the same score every time it is administered under the same conditions is reliable regardless of whether that score reflects the intended construct. A scale consistently measuring the wrong thing is a validity failure, not a reliability failure.
Does reliability belong to the instrument or to a specific use of it?
To a specific use. Reliability is a property of the scores an instrument produces in a given sample, population, and context — not an inherent, fixed property of the instrument itself that travels unchanged to every new use. A translated, adapted, or differently-administered version of an instrument, or the same instrument used in a different population, should have its reliability re-assessed rather than inheriting a figure reported elsewhere.
Why does low reliability reduce statistical power?
Because measurement error attenuates (weakens) observed correlations and effect sizes toward zero, per the correction-for-attenuation relationship first described by Spearman in 1904. A true effect measured with a noisy, low-reliability instrument produces a smaller observed effect size than the same true effect measured with a more reliable instrument, which directly reduces the statistical power available to detect it at a given sample size.
Related CASRAI Resources
- Cronbach’s Alpha — internal-consistency reliability: calculation, interpretation, and when to use McDonald’s omega instead.
- Test-Retest vs. Inter-Rater Reliability — direct comparison of the two, with statistics and benchmarks.
- Intraclass Correlation Coefficient (ICC) — the continuous-data reliability statistic behind both test-retest and inter-rater estimates.
- Types of Validity in Research — measurement validity and design validity, and how reliability bounds but does not imply validity.
- Questionnaire Design — wording practices that affect an instrument’s reliability before data collection even begins.
- Correlation Coefficient — the statistic several reliability estimates (test-retest, parallel-forms) are built on.
- Power Analysis and Sample Size Calculation — how attenuation from unreliable measurement affects a study’s statistical power.
- Research Methods hub — the full cluster hub for study design, sampling, analysis, and measurement content.







