Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Direct comparison

Test-Retest vs. Inter-Rater Reliability

Test-retest reliability checks consistency over time; inter-rater reliability checks agreement between raters. Compare statistics, benchmarks, and use cases.

Side-by-side comparison

DimensionTest-Retest ReliabilityInter-Rater Reliability
What it measuresConsistency of the same measure across timeConsistency of the same measure across different raters or observers
What varies between the two measurementsTime (same rater/instrument, two occasions)Rater or observer (same or near-same occasion, two or more raters)
What is held constantRater and instrumentTiming and the material/target being rated
Statistic for continuous dataPearson's r or intraclass correlation coefficient (ICC)Intraclass correlation coefficient (ICC)
Statistic for categorical dataCohen's kappa (same rater, two time points)Cohen's kappa (2 raters) or Fleiss' kappa (3+ raters); raw percent agreement is a weaker, chance-uncorrected alternative
Common interpretation benchmarkICC: <0.50 poor, 0.50-0.75 moderate, 0.75-0.90 good, >0.90 excellent (Koo & Li, 2016)Same ICC bands for continuous ratings; kappa: <0 poor, 0-0.20 slight, 0.21-0.40 fair, 0.41-0.60 moderate, 0.61-0.80 substantial, 0.81-1.00 almost perfect (Landis & Koch, 1977)
Main threat to a good resultPractice/memory effects, true change in the construct over the interval, regression to the meanAmbiguous coding criteria, rater bias or leniency differences, insufficient rater training/calibration
Number of raters/instruments requiredOne (same rater or instrument, two administrations)Two or more
Typical use caseValidating a survey, questionnaire, diagnostic instrument, or physiological measure meant to be stable over timeValidating a coding scheme, clinical diagnosis, qualitative theme coding, or systematic-review screening decisions
Often confused withInternal consistency (Cronbach's alpha) -- that measures item homogeneity within a single administration, not stability over timeIntra-rater reliability -- the same single rater scoring the same material twice, not different raters
Typical interval between comparisonsLong enough that respondents cannot simply recall earlier answers, short enough that the construct hasn't genuinely changed -- often 2-4 weeksNot time-based -- raters typically score the same fixed set of cases independently, close together in time

Common questions

FAQ

What is the basic difference between test-retest and inter-rater reliability?+

Test-retest reliability holds the rater and instrument constant and varies time -- it asks whether the same measurement process gives the same answer twice. Inter-rater reliability holds timing constant and varies the rater -- it asks whether different people scoring the same material agree with each other.

Can a measure have high test-retest reliability but low inter-rater reliability?+

Yes. A single well-trained rater can be highly consistent with themselves over time while a second rater, working from an ambiguous or under-specified coding scheme, produces very different scores. The two statistics isolate different sources of measurement error, so neither one guarantees the other.

What counts as a good ICC value for reliability?+

The commonly cited Koo & Li (2016) guideline treats ICC values below 0.50 as poor, 0.50-0.75 as moderate, 0.75-0.90 as good, and above 0.90 as excellent reliability. The same bands are typically applied to both test-retest and inter-rater ICCs, though acceptable thresholds vary somewhat by field and by how the measure will be used.

Is Cohen's kappa the same thing as inter-rater reliability?+

No -- Cohen's kappa is one specific statistic used to quantify inter-rater reliability for categorical data with two raters. Inter-rater reliability is the broader concept; the appropriate statistic depends on the data type and number of raters (ICC for continuous ratings, Fleiss' kappa for three or more raters on categorical data, Cohen's kappa for exactly two raters on categorical data).

How does test-retest reliability differ from internal consistency (Cronbach's alpha)?+

Test-retest reliability assesses stability across two separate administrations of the same measure over time. Internal consistency (commonly reported as Cronbach's alpha) assesses whether items within a single administration of a multi-item scale correlate with each other, i.e. whether they appear to measure the same underlying construct. A scale can have high internal consistency in one sitting and still show poor test-retest stability, or vice versa.

What is a typical time interval for a test-retest reliability study?+

There is no universal standard, but two to four weeks is a common default in survey and psychometric validation research -- long enough to reduce the chance that respondents simply recall their earlier answers, short enough that the underlying trait or attitude being measured is unlikely to have genuinely changed.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →