Direct comparison
Test-Retest vs. Inter-Rater Reliability
Test-retest reliability checks consistency over time; inter-rater reliability checks agreement between raters. Compare statistics, benchmarks, and use cases.
Side-by-side comparison
| Dimension | Test-Retest Reliability | Inter-Rater Reliability |
|---|---|---|
| What it measures | Consistency of the same measure across time | Consistency of the same measure across different raters or observers |
| What varies between the two measurements | Time (same rater/instrument, two occasions) | Rater or observer (same or near-same occasion, two or more raters) |
| What is held constant | Rater and instrument | Timing and the material/target being rated |
| Statistic for continuous data | Pearson's r or intraclass correlation coefficient (ICC) | Intraclass correlation coefficient (ICC) |
| Statistic for categorical data | Cohen's kappa (same rater, two time points) | Cohen's kappa (2 raters) or Fleiss' kappa (3+ raters); raw percent agreement is a weaker, chance-uncorrected alternative |
| Common interpretation benchmark | ICC: <0.50 poor, 0.50-0.75 moderate, 0.75-0.90 good, >0.90 excellent (Koo & Li, 2016) | Same ICC bands for continuous ratings; kappa: <0 poor, 0-0.20 slight, 0.21-0.40 fair, 0.41-0.60 moderate, 0.61-0.80 substantial, 0.81-1.00 almost perfect (Landis & Koch, 1977) |
| Main threat to a good result | Practice/memory effects, true change in the construct over the interval, regression to the mean | Ambiguous coding criteria, rater bias or leniency differences, insufficient rater training/calibration |
| Number of raters/instruments required | One (same rater or instrument, two administrations) | Two or more |
| Typical use case | Validating a survey, questionnaire, diagnostic instrument, or physiological measure meant to be stable over time | Validating a coding scheme, clinical diagnosis, qualitative theme coding, or systematic-review screening decisions |
| Often confused with | Internal consistency (Cronbach's alpha) -- that measures item homogeneity within a single administration, not stability over time | Intra-rater reliability -- the same single rater scoring the same material twice, not different raters |
| Typical interval between comparisons | Long enough that respondents cannot simply recall earlier answers, short enough that the construct hasn't genuinely changed -- often 2-4 weeks | Not time-based -- raters typically score the same fixed set of cases independently, close together in time |
Common questions
FAQ
What is the basic difference between test-retest and inter-rater reliability?+
Test-retest reliability holds the rater and instrument constant and varies time -- it asks whether the same measurement process gives the same answer twice. Inter-rater reliability holds timing constant and varies the rater -- it asks whether different people scoring the same material agree with each other.
Can a measure have high test-retest reliability but low inter-rater reliability?+
Yes. A single well-trained rater can be highly consistent with themselves over time while a second rater, working from an ambiguous or under-specified coding scheme, produces very different scores. The two statistics isolate different sources of measurement error, so neither one guarantees the other.
What counts as a good ICC value for reliability?+
The commonly cited Koo & Li (2016) guideline treats ICC values below 0.50 as poor, 0.50-0.75 as moderate, 0.75-0.90 as good, and above 0.90 as excellent reliability. The same bands are typically applied to both test-retest and inter-rater ICCs, though acceptable thresholds vary somewhat by field and by how the measure will be used.
Is Cohen's kappa the same thing as inter-rater reliability?+
No -- Cohen's kappa is one specific statistic used to quantify inter-rater reliability for categorical data with two raters. Inter-rater reliability is the broader concept; the appropriate statistic depends on the data type and number of raters (ICC for continuous ratings, Fleiss' kappa for three or more raters on categorical data, Cohen's kappa for exactly two raters on categorical data).
How does test-retest reliability differ from internal consistency (Cronbach's alpha)?+
Test-retest reliability assesses stability across two separate administrations of the same measure over time. Internal consistency (commonly reported as Cronbach's alpha) assesses whether items within a single administration of a multi-item scale correlate with each other, i.e. whether they appear to measure the same underlying construct. A scale can have high internal consistency in one sitting and still show poor test-retest stability, or vice versa.
What is a typical time interval for a test-retest reliability study?+
There is no universal standard, but two to four weeks is a common default in survey and psychometric validation research -- long enough to reduce the chance that respondents simply recall their earlier answers, short enough that the underlying trait or attitude being measured is unlikely to have genuinely changed.







