Written and maintained by CASRAI Editorial Board
Last updated
On this page: what minimal detectable change (MDC) is and what it is not, the three routes to a standard error of measurement (SEM) and the assumptions each one carries, the MDC formula worked end to end with real arithmetic, the individual-versus-group distinction, the four-cell decision rule for judging an observed change against both MDC and the minimal clinically important difference (MCID), why an MDC quoted without its source population is not transferable, and a reporting checklist.
Two Different Questions, Routinely Conflated
A patient’s score on an outcome measure moves by 9 points between baseline and follow-up. Two separate questions follow, and answering one does not answer the other:
- Is the change real? That is, is 9 points larger than the noise this instrument produces when nothing has actually changed? This is a measurement error question, and its threshold is the minimal detectable change (MDC).
- Is the change worth anything? That is, is 9 points large enough that a patient or clinician would regard it as a benefit justifying a change in management? This is an importance question, and its threshold is the minimal clinically important difference (MCID).
The most common error in applied outcomes research is treating these as one question, usually by reporting a single “clinically meaningful change” threshold without stating whether it came from a measurement-error calculation or from an anchor-based importance study. They are derived from different data by different methods and they can point in opposite directions. De Vet and colleagues made exactly this distinction the subject of a dedicated methodological paper, arguing that some distribution-based methods promoted as ways of estimating minimally important change had in fact only estimated minimally detectable change — a different quantity that says nothing about importance (de Vet HC, Terwee CB, Ostelo RW, Beckerman H, Knol DL, Bouter LM. Minimal changes in health status questionnaires: distinction between minimally detectable change and minimally important change. Health Qual Life Outcomes. 2006;4:54; DOI 10.1186/1477-7525-4-54).
This page treats them separately, calculates MDC properly, and then gives the decision rule for using both together.
What MDC Is
MDC is the smallest change in an individual’s score that exceeds the measurement error of the instrument, at a stated level of confidence. Below the MDC, an observed change is statistically indistinguishable from the instrument measuring the same unchanged person twice and getting two different numbers.
Terminology varies by literature and the variants are not always identical in construction:
| Term | Where it is used | Note |
|---|---|---|
| MDC (minimal detectable change) | Rehabilitation, physiotherapy, sports medicine | Most common in clinical practice literature; usually written MDC95 or MDC90 with the confidence level attached. |
| SDC (smallest detectable change) | COSMIN and health-measurement methodology | COSMIN’s preferred term. Same construction. Usually distinguished as SDCind and SDCgroup. |
| SRD (smallest real difference) | Clinimetrics, some biomechanics literature | Synonym for the individual-level quantity. |
| MDD (minimal detectable difference) | Occasionally, interchangeably | Check the formula before assuming equivalence; some authors use it for a group-level quantity. |
Under the COSMIN taxonomy, this quantity sits in the reliability domain — specifically under the measurement property measurement error, which the consensus defines as the systematic and random error of a patient’s score that is not attributed to true changes in the construct to be measured. The reliability domain contains three measurement properties: internal consistency, reliability, and measurement error (COSMIN taxonomy of measurement properties; Mokkink LB, Terwee CB, Patrick DL, et al. J Clin Epidemiol. 2010;63(7):737–745, DOI 10.1016/j.jclinepi.2010.02.006).
What MDC is not
- MDC is not responsiveness. COSMIN treats responsiveness as its own domain, defined as the ability of an outcome measure to detect change over time in the construct to be measured — it concerns the validity of a change score. MDC concerns the error in a change score. A measure can have a small MDC and still be unresponsive, because it registers change that is not the change you care about.
- MDC is not an effect size. It is expressed in the raw units of the instrument, not standardised. Two instruments’ MDCs are not comparable to each other, and an MDC cannot be interpreted with Cohen’s benchmarks. See CASRAI’s guides to choosing and reporting effect size and interpreting Cohen’s d for that separate question.
- MDC is not a p-value or a significance test. Comparing an individual’s change to the MDC is a confidence-interval statement about one person’s score, not a hypothesis test on a sample. The distinction between statistical and clinical verdicts is covered in CASRAI’s comparison of statistical significance vs. clinical significance.
- MDC is not a property of the instrument alone. It is a property of the instrument as measured in a particular sample. See why an MDC without its population is not transferable below.
Step 1: Obtain the Standard Error of Measurement
Every MDC calculation starts from the SEM, which expresses the typical size of the discrepancy between an observed score and the underlying true score, in the instrument’s own units. There are three routes to it, and they are not interchangeable in what they assume.
Route 1: from a reliability coefficient and a standard deviation
SEM = SD × √(1 − R)
where SD is the standard deviation of the scores in the sample and R is the reliability coefficient estimated in that same sample. This is the formula given in de Vet et al. (2006) and it is the route used in the large majority of published MDC values, because reliability studies routinely report both quantities.
The assumption that is most often violated here: R must be an agreement-type reliability coefficient. For continuous data that normally means an absolute-agreement intraclass correlation — ICC(2,1) in the Shrout and Fleiss notation, ICC(A,1) in McGraw and Wong’s. A consistency ICC (ICC(3,1) / ICC(C,1)) deliberately removes systematic between-occasion or between-rater bias from the error term. Feeding a consistency ICC into this formula therefore produces a smaller SEM and a smaller, optimistic MDC that ignores a source of error a real repeat measurement is exposed to. CASRAI’s guide to the intraclass correlation coefficient and its forms sets out which form to choose and why the choice is the single most common error in published reliability research.
The same caution applies to the source of the coefficient: Cronbach’s alpha is an internal-consistency coefficient describing agreement among items at a single administration, not agreement across occasions. Substituting alpha for a test–retest coefficient in this formula answers a different question and generally understates measurement error for change scores. CASRAI’s overview of reliability in research and the comparison of test-retest vs. inter-rater reliability cover which coefficient answers which question.
Route 2: from the standard deviation of the difference scores
If you have the raw test–retest data rather than a published coefficient, compute the difference for each subject (retest minus test) and take the standard deviation of those differences, SDdiff:
SEM = SDdiff / √2
This route makes no assumption about the heterogeneity of the sample entering through an SD of raw scores, and it is generally the more defensible calculation when you hold the data. It is also the route that makes the connection to Bland–Altman analysis explicit — see the identity below.
Route 3: from a repeated-measures ANOVA
SEM = √(MSerror)
The square root of the residual (error) mean square from a repeated-measures analysis of variance on the test–retest data. This is the most flexible route when the design has more than two occasions or more than one rater, because it lets you decide explicitly which variance components count as error. Whether the occasion (or rater) main effect is pooled into the error term is precisely the agreement-versus-consistency decision described in Route 1, made visible as a modelling choice. CASRAI’s guides to standard deviation and choosing random vs. fixed effects cover the underlying variance decomposition.
The assumptions the test–retest design itself must satisfy
All three routes inherit the validity of the repeated-measurement design that generated the data. State these explicitly when you report an MDC, because they are the assumptions most likely to be silently false:
- Stability. The interval between measurements must be long enough that recall and practice effects have faded, and short enough that no true change has occurred. In a stable chronic condition this is achievable; in acute recovery or early rehabilitation it frequently is not, and true change contaminates the error estimate upward.
- Homoscedastic error. A single MDC assumes measurement error is roughly constant across the score range. Plot the test–retest differences against their means; if the scatter fans out, one MDC misstates the error at one end of the scale. Report the pattern rather than a single number.
- Approximately normal difference scores. The 1.96 multiplier in the next step is a normal-distribution quantile. Skewed differences make it approximate.
- No floor or ceiling effects. Scores compressed against a boundary cannot show their full error, which deflates the apparent SEM for that subgroup.
- No systematic bias between occasions. Check the mean difference. A non-zero mean difference is a learning effect or drift, not random error, and it needs reporting separately rather than being absorbed into the MDC.
Step 2: The MDC Formula
MDC95 = 1.96 × √2 × SEM ≈ 2.77 × SEM
Two constants, each doing a specific job:
- 1.96 is the standard normal quantile for a two-sided 95% interval. Change it to change the confidence level.
- √2 (≈ 1.414) is there because a change score is the difference between two independently measured values, each carrying its own error. The variance of a difference between two independent quantities is the sum of their variances, so the standard error of a change score is SEM × √2, not SEM.
| Confidence level | Quantile | Full multiplier (quantile × √2) | Formula |
|---|---|---|---|
| 90% | 1.645 | 2.33 | MDC90 = 2.33 × SEM |
| 95% | 1.96 | 2.77 | MDC95 = 2.77 × SEM |
| 99% | 2.576 | 3.64 | MDC99 = 3.64 × SEM |
Always report the subscript. An unlabelled “MDC” is ambiguous between values that differ by roughly 19% (MDC90 vs MDC95), which is more than enough to flip a responder classification.
Substituting Route 1’s SEM into the formula gives the fully expanded version that some papers quote directly:
MDC95 = 1.96 × √2 × SD × √(1 − R)
The Bland–Altman identity worth knowing
Substitute Route 2’s SEM into the MDC95 formula and the √2 cancels:
MDC95 = 1.96 × √2 × (SDdiff / √2) = 1.96 × SDdiff
That is exactly the half-width of the 95% limits of agreement in a Bland–Altman analysis. When the mean difference between occasions is zero, MDC95 and the Bland–Altman limits of agreement are the same quantity expressed two ways. If a paper reports limits of agreement, you already have the MDC without needing the ICC at all — and if the mean difference is not zero, the asymmetry of the limits is telling you about a systematic bias the MDC formula would have hidden.
Worked Example, End to End
The figures below are an illustrative construction using a hypothetical 0–100 point outcome scale. They are not the published values of any specific real instrument — use them to follow the arithmetic, not to cite an MDC for any measure.
Inputs (both from the same test–retest reliability study, one week apart, in a stable outpatient sample of 60 people):
- Standard deviation of baseline scores: SD = 12.0 points
- Absolute-agreement test–retest reliability: ICC(2,1) = 0.90
Step 1 — SEM:
SEM = 12.0 × √(1 − 0.90) = 12.0 × √0.10 = 12.0 × 0.3162 = 3.79 points
Step 2 — MDC:
MDC95 = 1.96 × 1.4142 × 3.79 = 2.77 × 3.79 = 10.5 points
MDC90 = 1.645 × 1.4142 × 3.79 = 2.33 × 3.79 = 8.8 points
Step 3 — interpret. For an individual patient on this scale, an observed change smaller than 10.5 points cannot be distinguished from measurement error with 95% confidence. The 9-point change from the opening of this page would not clear MDC95 (10.5) but would clear MDC90 (8.8) — which is precisely why the unlabelled figure is useless and the subscript is not optional.
Individual versus group: the √n correction
The MDC above applies to one person measured twice. When the question is whether a group mean has changed, the relevant error is the standard error of a mean change, which shrinks with sample size. The group-level threshold divides by the square root of the number of subjects:
MDCgroup = MDCindividual / √n
For the worked example with n = 40 participants:
MDCgroup = 10.5 / √40 = 10.5 / 6.32 = 1.66 points
The gap between 10.5 and 1.66 is the whole reason a trial can report a statistically solid mean improvement while no individual participant in it demonstrably improved. Both statements can be simultaneously true and correctly calculated. Report which one you are making. Using a group-level threshold to classify individual responders understates the error by a factor of √n and manufactures responders out of noise; using an individual-level threshold to judge a group mean is needlessly conservative and understates a real treatment effect. See CASRAI’s guides to reporting confidence intervals and power analysis and sample size for the group-level machinery.
The MDC-vs-MCID Decision Rule
MCID — called minimal important change (MIC) in COSMIN terminology, and the term this page uses interchangeably — is the smallest change a patient would perceive as beneficial and that would justify a change in management. The concept was introduced by Jaeschke, Singer and Guyatt in 1989 (Control Clin Trials. 1989;10(4):407–415, DOI 10.1016/0197-2456(89)90005-6), originating in respiratory quality-of-life outcomes research.
Critically, an MCID cannot be derived from measurement error at all. It requires an external anchor — typically a patient-reported global rating of change — against which score changes are calibrated. De Vet et al. argue directly for anchor-based methods on exactly this ground: they include a definition of what is minimally important, whereas distribution-based methods, however sophisticated, do not in themselves provide any indication of the importance of an observed change. Any paper deriving a “clinically important difference” from 0.5 SD, or from one SEM, or from any other purely distributional rule, has calculated a detectability quantity and labelled it an importance quantity.
The four cells
With both thresholds in hand, an observed individual change falls into one of four cells:
| Change ≥ MCID | Change < MCID | |
|---|---|---|
| Change ≥ MDC | Real and meaningful. The change exceeds measurement error and clears the importance threshold. This is a defensible responder classification. | Real but not meaningful. The instrument genuinely detected a change; it is too small to matter to the patient. Report it as detected, not as benefit. |
| Change < MDC | The trap cell. The observed value clears the importance threshold but sits inside measurement error. You cannot claim a real change in this individual. This cell exists only when MCID < MDC — see below. | Neither. No evidence of real or important change. |
The criterion: is the instrument fit for individual decisions?
The bottom-left cell is not an edge case; it is a diagnostic. Its existence tells you something about the instrument, and de Vet et al. frame the comparison as the practical payoff of keeping the two concepts distinct — appreciating the distinction makes it possible to judge whether the minimally detectable change of an instrument is sufficiently small to detect minimally important changes.
Stated as a rule:
- If MCID > MDC, the instrument can resolve the smallest change that matters. An individual change clearing the MCID also clears measurement error. The bottom-left cell is empty and individual-level responder analysis is defensible.
- If MCID < MDC, the instrument is too noisy for the decision you are asking it to make at the individual level. There is a band of changes that would matter to a patient but that this instrument cannot distinguish from its own error. You have three honest options: report at group level only (where the √n correction shrinks the threshold), use MDC as the operative responder threshold and state plainly that you are using a detectability rather than an importance criterion, or use a less noisy instrument.
Continuing the worked example, suppose an anchor-based study on the same 0–100 scale had established an MCID of 8 points:
| Scenario | SD | ICC | SEM | MDC95 | MCID | Verdict |
|---|---|---|---|---|---|---|
| A — as calculated above | 12.0 | 0.90 | 3.79 | 10.5 | 8 | MCID < MDC. Not fit for individual responder classification. |
| B — better reliability, same sample | 12.0 | 0.95 | 2.68 | 7.4 | 8 | MCID > MDC. Individual decisions defensible. |
The difference between an instrument that supports individual clinical decisions and one that does not is, in this construction, the difference between a reliability coefficient of 0.90 and one of 0.95 — two values a reader would casually describe with the same word, “excellent.”
Why an MDC Without Its Source Population Is Not Transferable
This is the part most often omitted, and it is the reason MDC values quoted from a table in a review article are frequently wrong for the population they are being applied to.
De Vet et al. note that in classical test theory the SEM has a rather stable value across different populations — that stability is one of the standard arguments for reporting SEM rather than a reliability coefficient. But stable is not invariant, and it does not license lifting a number out of one study and into another. Three specific mechanisms break transferability:
- The reliability coefficient is strongly sample-dependent. An ICC is a variance ratio: between-subject variance over total variance. Estimate it in a more heterogeneous sample and the between-subject variance rises while the error variance does not, so the ICC rises — without the instrument having become any more precise. This is why an ICC on its own tells you nothing about detectable change.
- Route 1 multiplies a sample-dependent SD by a sample-dependent coefficient. The two dependencies partially offset, which is exactly why the product is more stable than either factor, but they do not cancel exactly. Consider the same instrument studied in a broader sample: SD = 18.0 with ICC = 0.95 (higher, because the sample is more heterogeneous). Then SEM = 18.0 × √0.05 = 4.02, and MDC95 = 11.2 points. The reliability coefficient improved from 0.90 to 0.95 and the detectable change got worse, 10.5 to 11.2. A reader who compared only the ICCs would have drawn the opposite conclusion.
- The stability assumption behaves differently in different populations. An MDC estimated in a stable chronic cohort and an MDC estimated in an acute-recovery cohort are not measuring the same thing, because in the second the retest interval admits true change that inflates the apparent error.
The practical consequence: an MDC is only interpretable alongside the sample it came from. When you report one, report the SEM, the reliability coefficient and its ICC form, the SD, the retest interval, the sample size, and the clinical characteristics of the sample. When you use a published one, check that its source sample resembles yours in severity range and stability — and if it does not, recompute rather than borrow. An MDC quoted as a bare number, with no population attached, cannot be checked and should not be relied on for an individual clinical decision.
Common Errors
| Error | Why it is wrong | Effect |
|---|---|---|
| Using 1.96 × SEM without the √2 | That is the confidence interval for a single score, not for a change between two scores | Understates MDC by about 29%; inflates apparent responders |
| Using a consistency ICC (ICC(3,1)) rather than an agreement ICC | Systematic between-occasion bias is excluded from the error term | Optimistic SEM and MDC |
| Using Cronbach’s alpha as R | Internal consistency describes agreement among items at one sitting, not across occasions | Answers a different question; generally understates error for change scores |
| Reporting “MDC” with no confidence subscript | MDC90 and MDC95 differ by roughly 19% | Ambiguous, non-reproducible threshold |
| Applying MDCgroup to individual patients | The √n correction only applies to a mean | Understates individual error by √n; manufactures responders |
| Calling a distribution-based value (0.5 SD, 1 SEM) an MCID | No anchor for importance was consulted | A detectability quantity mislabelled as an importance quantity |
| Quoting an MDC without its source sample | The value is conditional on the sample’s SD, reliability and stability | Threshold may be badly wrong for the target population |
| Treating a single MDC as valid across the whole score range | Assumes homoscedastic error | Misstates error at the scale extremes |
Reporting Checklist
When an MDC appears in a methods or results section, these items make it reproducible and checkable. Everything below is derivable from data you already collected in the reliability study:
- The SEM, with which of the three routes produced it
- The reliability coefficient, named by form (e.g. “ICC(2,1), absolute agreement”) — never a bare “ICC = 0.90”
- The SD used, and whether it is the baseline SD or the SD of difference scores
- The confidence level, as a subscript on the MDC
- Whether the value is individual-level or group-level, and if group-level, the n
- The retest interval and the justification for it against the stability assumption
- The mean difference between occasions, as evidence there was no systematic drift
- A statement on homoscedasticity (a Bland–Altman plot is the usual evidence)
- The sample: size, condition, severity range, setting
- If an MCID is also used, its anchor and source, and an explicit comparison of MCID to MDC
This page sits in CASRAI’s research methods cluster, alongside related coverage of psychometrics, types of validity, accuracy vs. precision in measurement and clinical outcome assessment validation.
Frequently Asked Questions
How do you calculate minimal detectable change?
Two steps. First obtain the standard error of measurement, most commonly as SEM = SD × √(1 − R), where R is an absolute-agreement test–retest reliability coefficient from the same sample. Then multiply: MDC95 = 1.96 × √2 × SEM, which is approximately 2.77 × SEM.
What is the difference between MDC and MCID?
MDC answers whether a change is real — larger than the instrument’s measurement error. MCID answers whether a change is meaningful — large enough that a patient or clinician would regard it as a benefit. MDC is computed from reliability data; MCID requires an external anchor such as a patient global rating of change. Neither substitutes for the other, and a change can clear one without clearing the other.
Why is there a √2 in the MDC formula?
Because a change score is the difference between two separately measured values, each carrying its own measurement error. The variance of a difference between two independent quantities is the sum of their variances, so the standard error of the change is SEM × √2. Omitting it understates MDC by about 29%.
What is the difference between MDC90 and MDC95?
Only the confidence level, which changes the quantile: MDC90 = 2.33 × SEM (from 1.645), MDC95 = 2.77 × SEM (from 1.96). MDC95 is roughly 19% larger and therefore more conservative. Always report the subscript — an unlabelled MDC is ambiguous between the two.
Is minimal detectable change the same as smallest detectable change?
Yes, in construction. SDC is the term COSMIN and the health-measurement methodology literature prefer; MDC is more common in rehabilitation and clinical practice literature. Smallest real difference (SRD) is a third synonym. Check the formula in any given paper rather than assuming, particularly with MDD, which some authors use for a group-level quantity.
What counts as a good MDC value?
There is no universal benchmark, because MDC is in the instrument’s own raw units and is not standardised. The only meaningful comparison is against the minimal clinically important difference for the same instrument in a comparable population: an MDC smaller than the MCID means the instrument can resolve the smallest change that matters, and an MDC larger than the MCID means it cannot — at least not for individual-level decisions.
Can I use a published MDC for my own study population?
Only after checking that its source sample resembles yours in severity range, setting and stability, and only if the paper reports the sample it came from. MDC depends on the standard deviation and the reliability coefficient of the sample it was estimated in, and a higher reported ICC does not imply a smaller MDC — a more heterogeneous sample raises the ICC without making the instrument more precise. If the source population differs materially, recompute from your own test–retest data.
Is MDC the same as responsiveness?
No. Under the COSMIN taxonomy, responsiveness is a separate domain concerning the validity of a change score — whether the instrument detects change in the construct it is supposed to measure. MDC sits under measurement error in the reliability domain and concerns how much of an observed change could be noise. An instrument can have a small MDC and poor responsiveness.
Sources
- de Vet HC, Terwee CB, Ostelo RW, Beckerman H, Knol DL, Bouter LM. Minimal changes in health status questionnaires: distinction between minimally detectable change and minimally important change. Health and Quality of Life Outcomes. 2006;4:54. DOI 10.1186/1477-7525-4-54. Full text (PMC1560110).
- Mokkink LB, Terwee CB, Patrick DL, Alonso J, Stratford PW, Knol DL, Bouter LM, de Vet HC. The COSMIN study reached international consensus on taxonomy, terminology, and definitions of measurement properties for health-related patient-reported outcomes. Journal of Clinical Epidemiology. 2010;63(7):737–745. DOI 10.1016/j.jclinepi.2010.02.006.
- COSMIN taxonomy of measurement properties — the reliability domain (internal consistency, reliability, measurement error) and the responsiveness domain.
- Jaeschke R, Singer J, Guyatt GH. Measurement of health status. Ascertaining the minimal clinically important difference. Controlled Clinical Trials. 1989;10(4):407–415. DOI 10.1016/0197-2456(89)90005-6.








