Written and maintained by CASRAI Editorial Board
Last updated
On this page: why the interval between the two administrations — not the coefficient — is the real methodological problem in a test–retest study; the four distinct threats the interval trades off against each other; what the COSMIN Risk of Bias checklist actually asks about it, verbatim; which intraclass correlation form a test–retest design needs and why absolute agreement is the defensible default; a seeded R demonstration showing that a consistency-form ICC and Pearson’s r are mathematically blind to a systematic shift between occasions; and the Bland–Altman check that catches what a single coefficient cannot.
This page deliberately does not restate ICC form mechanics. The model/type/definition decision, the Shrout & Fleiss and McGraw & Wong notations and the full decision table live on CASRAI’s guide to the intraclass correlation coefficient, and the conversion of a reliability coefficient into a change threshold lives on minimal detectable change. What follows is the part neither of those covers: the interval itself.
What a Test–Retest Coefficient Actually Contains
Test–retest reliability is estimated by measuring the same subjects with the same instrument on two occasions and quantifying the agreement between the two sets of scores. The coefficient that comes out is a variance ratio: between-subject variance over between-subject variance plus within-subject variance. That framing is where the trouble starts, because the within-subject term is not one thing. Between the first and second administration, at least four separable processes can move a subject’s score:
- Random instrument error — the quantity you actually intended to estimate.
- Carryover from the first administration — recall of the earlier answers, practice at the task, or sensitisation to the construct because it was just asked about.
- Genuine change in the construct — recovery, deterioration, learning, seasonal variation, an intervening treatment.
- Regression to the mean — if the first occasion selected or measured subjects at an extreme, the second will sit closer to the centre for purely statistical reasons. See CASRAI’s dictionary entry on regression to the mean.
A test–retest coefficient cannot separate these. It reports one number covering all four, and the interval is the only lever the designer has over their relative sizes. This is the confound that makes the interval the whole methodological problem: a test–retest coefficient measures instrument stability only to the extent that the interval has been chosen so that construct stability holds. Where it does not hold, the same coefficient is reporting something else entirely — and reporting it without any label saying so.
The direction of the error is not symmetric, which is what makes the tradeoff genuinely difficult rather than merely a matter of splitting the difference:
- Too short and carryover inflates the estimate. A respondent reproducing a remembered answer produces agreement that has nothing to do with the instrument’s measurement properties. The instrument looks more reliable than it is, and the resulting standard error of measurement is optimistic.
- Too long and true change deflates it. Real movement in the construct enters the within-subject variance term and is scored as measurement error. The instrument looks less reliable than it is — and this is the direction that quietly contaminates everything downstream, because an inflated error term propagates into an inflated minimal detectable change threshold, which then declares genuine individual improvement to be noise.
The Four Interval-Linked Threats, and What Each One Does to the Number
| Threat | Mechanism | Effect on the coefficient | Which interval direction reduces it |
|---|---|---|---|
| Recall / memory | The respondent reproduces the answer they remember giving, not the answer the item elicits afresh | Inflates | Longer |
| Practice / learning | Performance genuinely improves through exposure to the task, common in cognitive, motor and timed measures | Inflates agreement in ranking but shifts the mean; a consistency coefficient can stay high while every score moves | Longer, or a parallel form |
| Reactivity / sensitisation | Being asked about the construct changes it — attention to a symptom, awareness of a behaviour | Unpredictable direction; typically shifts the mean | Longer |
| True change in the construct | Recovery, deterioration, treatment, maturation, seasonality | Deflates | Shorter |
Because the first three push one way and the fourth pushes the other, there is no interval that eliminates all of them. There is only an interval that is defensible for a specific construct, in a specific population, with a specific instrument — and a reporting obligation to say which one you used and why. Note that practice effects and reactivity are the awkward pair: they do not simply inflate or deflate, they displace. That distinction is what the rest of this page turns on.
What COSMIN Actually Asks About the Interval
The COSMIN Risk of Bias checklist — the standard appraisal instrument for studies of measurement properties, and the one a systematic reviewer will apply to your paper — devotes three of its design-requirement items in the Reliability box (Box 6) to exactly this, and repeats the same three verbatim in the Measurement Error box (Box 7):
- “Were patients stable in the interim period on the construct to be measured?”
- “Was the time interval appropriate?”
- “Were the test conditions similar for the measurements? e.g. type of administration, environment, instructions”
Two features of how those items are scored are worth stating plainly, because they are where papers lose rating points without realising it.
An unstated interval is not neutral — it is a downgrade. The rating options for item 2 place “Doubtful whether time interval was appropriate or time interval was not stated” in the doubtful column. Silence is scored the same as a suspect interval. Similarly, item 1’s very good rating requires “Evidence provided that patients were stable”; merely being “Assumable that patients were stable” drops the study to adequate. Stability is a claim you are expected to evidence, not assert.
A correlation coefficient is only acceptable with separate evidence about drift. Box 6’s statistical-methods item on continuous scores rates a study adequate where a “Pearson or Spearman correlation coefficient [was] calculated with evidence provided that no systematic change has occurred”, and doubtful where the same coefficient was “calculated WITHOUT evidence provided that no systematic change has occurred or WITH evidence that systematic change has occurred”. The checklist is encoding, as a scoring rule, the mathematical fact demonstrated below: a correlation-type statistic cannot itself detect a systematic shift between occasions, so the shift has to be ruled out by a separate analysis.
COSMIN does not name a number of days, and neither should a methods section that is merely copying a convention. Its own position is that what counts as appropriate depends on the construct and the target population. The frequently repeated “two weeks” figure is a working default in patient-reported outcome research, not a standard; it appears on CASRAI’s comparison of test–retest and inter-rater reliability as a typical value for that reason, and it should never be cited as though a standards body had specified it.
Choosing an Interval You Can Defend
The defensible procedure is to derive the interval from two properties of your own study rather than importing a number:
- Establish the memory and practice horizon of the instrument. How long until a respondent could not plausibly reproduce their earlier responses, and until task familiarity has stopped conferring an advantage? A 60-item questionnaire with fine-grained response options has a much shorter recall horizon than a five-item screening tool with binary answers. A timed motor or cognitive task may have a practice horizon far longer than its recall horizon — a person can forget their score and still retain the skill.
- Establish the expected rate of true change in the construct, in this population. A stable trait in a chronic, clinically stable cohort tolerates weeks. An acute post-operative recovery cohort may not tolerate days. If the construct is expected to move measurably inside the shortest interval that clears the recall horizon in step 1, then a clean test–retest estimate is not obtainable in that population and the honest report says so rather than presenting a contaminated coefficient.
- Where the two horizons cannot both be satisfied, change the design rather than the number. The available options are a parallel-forms design (different items measuring the same construct, which removes the recall problem without needing a longer gap — see the parallel-forms section of CASRAI’s guide to reliability in research), restricting the reliability sub-study to a demonstrably stable subgroup, or using an anchor question at the second occasion to identify and exclude subjects who report having changed.
- Hold the conditions constant and record them. COSMIN item 3 above is not a formality. Time of day, administration mode, setting and the exact instructions given are all sources of between-occasion variance that will be attributed to the instrument if they are allowed to vary. Mode is a particularly common lapse in modern studies: paper at occasion one and an app at occasion two turns a test–retest study into an unlabelled method-comparison study.
The anchor-question option deserves a caution: excluding self-reported changers makes the remaining sample more stable but also more homogeneous, and a reliability coefficient is a variance ratio that rises and falls with the heterogeneity of the sample it is computed in. This is the same population-dependence that makes the standard error of measurement the more transferable quantity — de Vet and colleagues note that under classical test theory the SEM has a relatively stable value across populations, while the reliability coefficient does not. Restricting the sample to protect the stability assumption can therefore shrink the very coefficient it was meant to clean up.
Which ICC Form a Test–Retest Design Needs
In a test–retest design the two occasions occupy the position that raters occupy in an inter-rater design, and the same three-part choice applies — model, type, definition. The mechanics of that choice are covered in full on CASRAI’s ICC guide; what is specific to test–retest is how the three decisions resolve:
- Model. The same two occasions apply to every subject, so a two-way model is correct; a one-way model, ICC(1,1), is appropriate only where each subject was measured on a different, randomly chosen pair of occasions. Whether the two-way model should be random or mixed depends on whether you intend the estimate to generalise beyond these particular two occasions, which for a test–retest study is nearly always yes.
- Type. Single measures, the “,1” forms, in almost every case — because in practice one administration on one occasion is what gets used to make a decision about a person. The average-measures form answers the question “how reliable is the mean of two occasions?”, which is a question about a measurement procedure nobody is going to run.
- Definition. Absolute agreement, not consistency, is the defensible default. Scores from occasion one and occasion two are used interchangeably — that is the entire premise of measuring the same person twice — so a systematic offset between them is a real problem, not a nuisance to be partialled out. The next section shows exactly what choosing consistency costs.
Koo and Li’s 2016 guideline remains the standard selection and reporting reference (note that it carries a 2017 erratum in the same journal). Its interpretation bands — below 0.50 poor, 0.50 to 0.75 moderate, 0.75 to 0.90 good, above 0.90 excellent — are the most-quoted part of it, but its more consequential recommendation is to judge reliability from the 95% confidence interval rather than the point estimate. In a test–retest study that recommendation compounds with everything above: a point estimate that is already conditional on an unverified stability assumption, reported without its interval, is two layers of unstated uncertainty deep.
Why a Consistency ICC Cannot See Drift
The claim that a consistency-form coefficient is blind to a systematic between-occasion shift is not an approximation. It is exact, and it is easy to demonstrate. The following computes both ICC forms directly from the two-way ANOVA mean squares in base R, with no packages, on the same 40 subjects measured twice — once with no drift, and once with a constant four-point learning effect added to every subject’s second score.
set.seed(20260826)
n <- 40
true_score <- rnorm(n, mean = 50, sd = 10)
e1 <- rnorm(n, mean = 0, sd = 3)
e2 <- rnorm(n, mean = 0, sd = 3)
occasion1 <- true_score + e1
occasion2_stable <- true_score + e2
occasion2_drifted <- true_score + e2 + 4 # every subject scores 4 points higher
icc_pair <- function(t1, t2) {
n <- length(t1)
d <- data.frame(
score = c(t1, t2),
subject = factor(rep(seq_len(n), 2)),
occasion = factor(rep(c(1, 2), each = n))
)
ms <- summary(aov(score ~ subject + occasion, data = d))[[1]][, "Mean Sq"]
MSR <- ms[1] # between subjects
MSC <- ms[2] # between occasions
MSE <- ms[3] # residual
k <- 2
consistency <- (MSR - MSE) / (MSR + (k - 1) * MSE)
agreement <- (MSR - MSE) / (MSR + (k - 1) * MSE + k * (MSC - MSE) / n)
diffs <- t2 - t1
c(consistency = consistency,
agreement = agreement,
pearson_r = cor(t1, t2),
mean_diff = mean(diffs),
loa_lower = mean(diffs) - 1.96 * sd(diffs),
loa_upper = mean(diffs) + 1.96 * sd(diffs))
}
message("--- No drift between occasions ---")
print(round(icc_pair(occasion1, occasion2_stable), 3))
message("--- Constant +4 point drift between occasions ---")
print(round(icc_pair(occasion1, occasion2_drifted), 3))
Output, from R 4.6.1:
--- No drift between occasions ---
consistency agreement pearson_r mean_diff loa_lower loa_upper
0.917 0.919 0.928 0.201 -7.886 8.288
--- Constant +4 point drift between occasions ---
consistency agreement pearson_r mean_diff loa_lower loa_upper
0.917 0.846 0.928 4.201 -3.886 12.288
Read the two rows against each other. Between a dataset with no drift and a dataset in which every single subject scored four points higher on the second occasion:
- The consistency ICC is identical to three decimal places: 0.917 and 0.917. A constant added to one occasion changes only the between-occasion mean square, and the consistency formula does not contain that term. It is not that consistency is insensitive to drift; it is that drift is algebraically absent from it.
- Pearson’s r is likewise identical: 0.928 and 0.928. Adding a constant to one variable cannot change a correlation. This is why COSMIN scores a bare Pearson correlation as doubtful unless accompanied by separate evidence about systematic change.
- Only the absolute-agreement ICC moves, from 0.919 to 0.846 — and note how modest that drop is. A four-point systematic error on an instrument whose subjects have a standard deviation of about 10 points, affecting every observation, still leaves a coefficient that Koo and Li’s bands label “good”. Even the correct ICC form is a blunt instrument for detecting drift.
- The mean difference is unambiguous: 0.201 against 4.201. It is the one statistic here that reports the problem at full size.
The figures above are a seeded simulation used to demonstrate an algebraic property. They are not measurements of any real instrument and should not be cited as reliability values for anything.
The Bland–Altman Check
Bland and Altman’s 1986 method — plotting the difference between the two measurements against their mean, and drawing the mean difference and the limits of agreement at the mean difference plus or minus 1.96 standard deviations of the differences — is usually introduced as a method-comparison tool for two instruments. Applied to the two occasions of a test–retest study, it is the diagnostic for exactly the failure modes the interval creates, and it reports three things a single coefficient structurally cannot:
- Systematic drift, at full size and in original units. The mean difference is the estimate of the learning, practice or reactivity effect. In the demonstration above it reads 4.201 points — a directly interpretable quantity, not a coefficient that has to be reasoned back into units. If the mean difference is materially non-zero, the stability assumption underpinning the whole study has failed, and the correct response is to report the drift as a finding rather than absorb it into an error term.
- Whether error is constant across the score range. If the scatter of differences fans out with increasing mean, a single reliability coefficient and a single standard error of measurement misstate the error at one end of the scale. This heteroscedasticity is invisible in any single-number summary.
- Where individual disagreement actually lands. The limits of agreement give the interval within which about 95% of between-occasion differences fall, which is the quantity a clinician deciding about one patient needs. Note the relationship recorded on CASRAI’s minimal detectable change page: when the mean difference is zero, the half-width of the limits of agreement and MDC95 are the same quantity expressed two ways. When it is not zero, the asymmetry of the limits is the drift, made visible.
The practical rule that follows: report the mean difference alongside every test–retest coefficient. It costs one line, it is the direct evidence COSMIN’s statistical-methods item asks for, and it is the only routinely available statistic that scales with the drift rather than being algebraically immune to it.
Reporting Checklist
Everything below is either explicitly required by a COSMIN item or needed to make the coefficient interpretable outside the sample it came from:
- The retest interval, as an actual duration and as a distribution if it varied across subjects — not “approximately two weeks”.
- The justification for that interval against both horizons: why it is long enough for recall and practice, and short enough for stability in this population.
- The evidence for stability in the interim — an anchor question, clinical status, absence of intervention — rather than an assumption of it.
- The administration conditions at both occasions, including mode, setting, time of day and instructions, with any differences stated.
- The ICC form in full (model, type, definition), never a bare “ICC = 0.90”, and the software and version that computed it.
- The point estimate with its 95% confidence interval, and a judgment made from the interval.
- The mean difference between occasions, and limits of agreement or an equivalent drift check.
- The sample size and its heterogeneity — the coefficient is a variance ratio and is not transferable to a more or less heterogeneous population, which is why the standard error of measurement travels better than the coefficient does.
For the general reporting frame, GRRAS (Kottner and colleagues, Journal of Clinical Epidemiology, 2011) is the established reliability and agreement reporting checklist, and COSMIN is the broader instrument-evaluation framework it sits inside. This page is part of CASRAI’s research methods cluster; for how test–retest sits alongside the other reliability types see reliability in research measurement and, for the wider measurement frame, psychometrics.
Frequently Asked Questions
What is test-retest reliability?
Test–retest reliability is the degree to which the same instrument, administered to the same subjects on two separate occasions under the same conditions, produces the same scores. It is one of the four reliability types alongside internal consistency, inter-rater and parallel-forms reliability. Strictly, it estimates the stability of the measurement process only under the assumption that the construct itself did not change between occasions.
How long should the interval between test and retest be?
There is no standard duration, and no standards body specifies one. The widely repeated two-week default is a convention in patient-reported outcome research, not a requirement. The interval has to be long enough that respondents cannot reproduce their earlier answers and that practice effects have faded, and short enough that the construct has not genuinely changed in the population being studied. Where those two conditions cannot both be met, the design needs changing — to parallel forms, or to a more stable subgroup — rather than the number being split.
Which ICC should I use for test-retest reliability?
A two-way model, single-measures type, absolute-agreement definition — ICC(2,1) in Shrout and Fleiss notation, ICC(A,1) in McGraw and Wong notation — is the defensible default, because scores from the two occasions are intended to be used interchangeably and so a systematic offset between them counts as a real problem. Use the consistency definition only where relative ranking is genuinely all that matters. The full form-selection decision is set out on CASRAI’s ICC guide.
Can you use Pearson’s r for test-retest reliability?
Not on its own. Pearson’s r is invariant to adding a constant to one occasion, so it cannot detect a systematic shift between administrations — the demonstration above shows r unchanged at 0.928 across datasets differing by a four-point drift on every observation. The COSMIN Risk of Bias checklist reflects this: it rates a study using Pearson or Spearman correlation as doubtful unless separate evidence is provided that no systematic change occurred. See CASRAI’s guide to the correlation coefficient for what r does and does not measure.
What is a carryover effect in a test-retest study?
Carryover is any influence of the first administration on the second: recall of the earlier answers, practice at the task, or sensitisation to the construct because it was just measured. Recall inflates the coefficient by producing agreement that is not attributable to the instrument. Practice and reactivity are more insidious because they shift the mean rather than simply inflating agreement, which is precisely the pattern a consistency-form coefficient cannot see.
Does a high test-retest coefficient mean the instrument is stable?
Not by itself. A high coefficient is consistent with a stable instrument, but also with an interval short enough that respondents remembered their answers, or with a systematic drift affecting every subject equally where a consistency-form coefficient was used. Reading the coefficient as evidence of instrument stability requires the interval justification, the stability evidence, the correct ICC form and the mean difference all to be reported alongside it.
What is the difference between test-retest reliability and internal consistency?
Internal consistency, usually reported as Cronbach’s alpha, is estimated within a single administration and asks whether the items of a scale cohere. Test–retest is estimated across two administrations and asks whether the scale gives the same answer twice. They are not substitutes and neither implies the other: a scale can be internally coherent in one sitting and unstable across weeks. Alpha is also not a defensible source for a standard error of measurement in a test–retest context — COSMIN rates that as inadequate.
Is test-retest reliability the same as inter-rater reliability?
No. Test–retest holds the rater and instrument constant and varies time; inter-rater holds timing constant and varies the rater. See CASRAI’s comparison of the two, and, for choosing among kappa, weighted kappa, Krippendorff’s alpha and the rest, the guide to choosing an inter-rater reliability coefficient.
Sources
- Mokkink LB, de Vet HCW, Prinsen CAC, Patrick DL, Alonso J, Bouter LM, Terwee CB. COSMIN Risk of Bias checklist for systematic reviews of Patient-Reported Outcome Measures. Quality of Life Research 2018;27(5):1171–1179. DOI 10.1007/s11136-017-1765-4. Item wording quoted above is from the July 2018 revision of the checklist, Boxes 6 and 7.
- Koo TK, Li MY. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine 2016;15(2):155–163. DOI 10.1016/j.jcm.2016.02.012. Erratum: J Chiropr Med 2017;16(4):346.
- Bland JM, Altman DG. Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet 1986;1(8476):307–310. DOI 10.1016/s0140-6736(86)90837-8.
- Shrout PE, Fleiss JL. Intraclass correlations: uses in assessing rater reliability. Psychological Bulletin 1979;86(2):420–428. The source of the ICC(model,type) notation used above.
- de Vet HCW, Terwee CB, Ostelo RWJG, Beckerman H, Knol DL, Bouter LM. Minimal changes in health status questionnaires: distinction between minimally detectable change and minimally important change. Health and Quality of Life Outcomes 2006;4:54. DOI 10.1186/1477-7525-4-54. Source of the observation that the standard error of measurement is more population-stable than the reliability coefficient.
- Kottner J, Audige L, Brorson S, Donner A, Gajewski BJ, Hrobjartsson A, Roberts C, Shoukri M, Streiner DL. Guidelines for Reporting Reliability and Agreement Studies (GRRAS) were proposed. Journal of Clinical Epidemiology 2011;64(1):96–106.








