Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & Research SupplyReagents, PPE & instruments — chain-of-custody documented.Fast, traceable sourcing built for regulated research environments, from bench consumables to instrumentation.Shop lac.us CodeCASRAIlac.us

Direct comparison

Test-Retest Reliability vs. Inter-Rater

Test-retest reliability = the same measure given to the same sample twice, then correlated (Pearson r or ICC), usually 2-4 weeks apart. Vs. inter-rater.

Ask about Test-Retest Reliability vs. Inter-Rater

Answers are drawn from this comparison and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

How do Test-Retest Reliability, Inter-Rater Reliability compare side by side?

The table below compares Test-Retest Reliability, Inter-Rater Reliability across 11 procurement-relevant dimensions, from what it measures through typical interval between comparisons.

Side-by-side comparison

DimensionTest-Retest ReliabilityInter-Rater Reliability
What it measuresConsistency of the same measure across timeConsistency of the same measure across different raters or observers
What varies between the two measurementsTime (same rater/instrument, two occasions)Rater or observer (same or near-same occasion, two or more raters)
What is held constantRater and instrumentTiming and the material/target being rated
Statistic for continuous dataPearson's r or intraclass correlation coefficient (ICC)Intraclass correlation coefficient (ICC)
Statistic for categorical dataCohen's kappa (same rater, two time points)Cohen's kappa (2 raters) or Fleiss' kappa (3+ raters); raw percent agreement is a weaker, chance-uncorrected alternative
Common interpretation benchmarkICC: <0.50 poor, 0.50-0.75 moderate, 0.75-0.90 good, >0.90 excellent (Koo & Li, 2016)Same ICC bands for continuous ratings; kappa: <0 poor, 0-0.20 slight, 0.21-0.40 fair, 0.41-0.60 moderate, 0.61-0.80 substantial, 0.81-1.00 almost perfect (Landis & Koch, 1977)
Main threat to a good resultPractice/memory effects, true change in the construct over the interval, regression to the meanAmbiguous coding criteria, rater bias or leniency differences, insufficient rater training/calibration
Number of raters/instruments requiredOne (same rater or instrument, two administrations)Two or more
Typical use caseValidating a survey, questionnaire, diagnostic instrument, or physiological measure meant to be stable over timeValidating a coding scheme, clinical diagnosis, qualitative theme coding, or systematic-review screening decisions
Often confused withInternal consistency (Cronbach's alpha) -- that measures item homogeneity within a single administration, not stability over timeIntra-rater reliability -- the same single rater scoring the same material twice, not different raters
Typical interval between comparisonsLong enough that respondents cannot simply recall earlier answers, short enough that the construct hasn't genuinely changed -- often 2-4 weeksNot time-based -- raters typically score the same fixed set of cases independently, close together in time

Common questions

Common questions about Test-Retest Reliability vs Inter-Rater Reliability

What is the basic difference between test-retest and inter-rater reliability?

+

Test-retest reliability holds the rater and instrument constant and varies time -- it asks whether the same measurement process gives the same answer twice. Inter-rater reliability holds timing constant and varies the rater -- it asks whether different people scoring the same material agree with each other.

Can a measure have high test-retest reliability but low inter-rater reliability?

+

Yes. A single well-trained rater can be highly consistent with themselves over time while a second rater, working from an ambiguous or under-specified coding scheme, produces very different scores. The two statistics isolate different sources of measurement error, so neither one guarantees the other.

What counts as a good ICC value for reliability?

+

The commonly cited Koo & Li (2016) guideline treats ICC values below 0.50 as poor, 0.50-0.75 as moderate, 0.75-0.90 as good, and above 0.90 as excellent reliability. The same bands are typically applied to both test-retest and inter-rater ICCs, though acceptable thresholds vary somewhat by field and by how the measure will be used.

Is Cohen's kappa the same thing as inter-rater reliability?

+

No -- Cohen's kappa is one specific statistic used to quantify inter-rater reliability for categorical data with two raters. Inter-rater reliability is the broader concept; the appropriate statistic depends on the data type and number of raters (ICC for continuous ratings, Fleiss' kappa for three or more raters on categorical data, Cohen's kappa for exactly two raters on categorical data).

How does test-retest reliability differ from internal consistency (Cronbach's alpha)?

+

Test-retest reliability assesses stability across two separate administrations of the same measure over time. Internal consistency (commonly reported as Cronbach's alpha) assesses whether items within a single administration of a multi-item scale correlate with each other, i.e. whether they appear to measure the same underlying construct. A scale can have high internal consistency in one sitting and still show poor test-retest stability, or vice versa.

What is a typical time interval for a test-retest reliability study?

+

There is no universal standard, but two to four weeks is a common default in survey and psychometric validation research -- long enough to reduce the chance that respondents simply recall their earlier answers, short enough that the underlying trait or attitude being measured is unlikely to have genuinely changed.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →