Skip to main content
v2026.11,772 entries · CC-BY 4.0

The Donabedian Model: Separating Structure, Process and Outcome Measures in a QI Project

A working guide to Donabedian’s structure-process-outcome triad for QI teams: how to classify a candidate measure, why process measures produce signal long before outcomes do, how strong the causal chain is in published validation studies, and where balancing measures belong.

Ask CASRAI · included with Regulatory Radar

Ask about The Donabedian Model: Separating Structure, Process and Outcome Measures in a QI Project

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Written and maintained by CASRAI Editorial Board

Last updated

Most improvement teams can recite the triad. Far fewer can say which of their own measures is which — and the ones that get it wrong tend to fail the same way. They build the aim statement around an outcome that cannot plausibly move inside the project’s window, collect it faithfully for six months, and end with a number that never leaves the noise band. The team then reports “no significant improvement” for a change package that may have worked perfectly well.

Classifying a measure as structure, process or outcome is not a labelling exercise for the methods section. Donabedian’s triad is a claim about a causal chain, and the position of a measure on that chain predicts two operationally decisive things: how quickly the measure can respond to your intervention, and how confidently you can attribute a change in it to your work rather than to case mix, seasonality or chance. Get that right and your measurement plan does useful work. Get it wrong and no amount of statistical care will rescue it.

What Donabedian actually proposed

Avedis Donabedian set out the framework in “Evaluating the Quality of Medical Care” in the Milbank Memorial Fund Quarterly in 1966 — a paper important enough that Milbank reprinted it in full in 2005. He returned to it in JAMA in 1988 with the operational statement that most people are actually quoting when they cite “the Donabedian model”. In that paper he argued that quality assessment requires “detailed information about the causal linkages among the structural attributes of the settings in which care occurs, the processes of care, and the outcomes of care.”

The three domains, as the European Observatory on Health Systems and Policies restates them:

  • Structure — the attributes of the setting in which care occurs: material resources (facilities, equipment, drug supply), intellectual resources (knowledge, information systems, protocols) and human resources (staffing levels, qualifications, skill mix).
  • Process — the components of care actually delivered. This covers both clinical activity (what was prescribed, what was referred, what was assessed) and organisational activity (how supply is managed, how waiting lists are run).
  • Outcome — the effects of care on the health status of patients and populations, split between final outcomes (mortality, morbidity, disability) and intermediate outcomes (blood pressure, functional ability, knowledge gained).

The load-bearing proposition is directional: good structure increases the likelihood of good process, and good process increases the likelihood of good outcome. Note the hedge. Donabedian claimed probabilistic linkage, not determinism — which is exactly why a project can improve process reliably and still see a flat outcome chart, without either result being wrong.

The classification test that actually works

Practitioners misclassify measures because they sort on how a measure feels rather than on what the number counts. Three questions settle almost every case.

  1. Would this number still exist if no patient were treated today? If yes, it is a structure measure. Nurse-to-patient ratio, proportion of staff holding a given qualification, whether a written protocol exists, whether an order set is live in the EHR, number of functioning infusion pumps — these are standing attributes of the setting. They exist on a quiet Sunday.
  2. Does the number count an action performed by or on the care team, expressed as a proportion of eligible cases? If yes, it is a process measure. The defining shape is a numerator of eligible cases in which the action happened over a denominator of all eligible cases.
  3. Does the number describe a change in the patient’s state, experience or resource consumption? If yes, it is an outcome measure. Mortality, complication rate, readmission, length of stay, functional score, patient-reported experience.

The three misclassifications that recur

Training completion is structure, not process. “Percentage of ward nurses who completed the new sepsis-bundle education” is a property of the workforce, not an account of care delivered to a patient. It is a legitimate and useful structure measure — it just cannot substitute for the process measure, which is the proportion of eligible patients who actually received the bundle element within the window your local protocol or the applicable national measure specification defines. Teams that track training completion and call it process routinely report 100% “compliance” alongside unchanged patient-level care.

Timeliness is process, not outcome. “Time to antibiotic” or “door-to-needle time” looks like a result because it produces a number at the end of an episode. It is not. It measures the timeliness of an action the team performed; the patient’s state has not been described. It behaves like a process measure statistically too — it responds within weeks and needs no risk adjustment beyond eligibility.

Patient satisfaction is outcome, not process. Because satisfaction is generated during the interaction, teams file it under process. Donabedian’s logic puts it under outcome: it is an effect of care on the patient, not a description of what was done. A scoping review of the model in outpatient settings classified patient satisfaction, adverse events such as falls and medication errors, follow-up care and complaint-handling rates all as outcome indicators. If your instrument is a formal patient-reported tool, the design questions that follow are those covered in patient-reported outcome measures in hospital quality reporting.

Why your process measure moves and your outcome measure does not

This is the single most useful thing the triad tells a QI team, and it comes from a specific piece of evidence. Mant and Hicks (BMJ, 1995) took the interventions with proven mortality benefit in acute myocardial infarction, used meta-analysis and large randomised trials to estimate how much optimal use of them would shift mortality in a typical district general hospital, then ran the sample-size arithmetic on how many years of data a hospital would need to detect that difference. Their conclusion: process measures derived from randomised trial evidence could detect relevant differences between hospitals that comparison of hospital-specific mortality would not identify. Mortality, they concluded, is an insensitive indicator of the quality of care.

The mechanism is arithmetic, not ideology. An outcome measure is diluted by everything the intervention does not touch. If a bundle element is relevant to a fraction of patients, and the outcome has a low base rate, and the outcome is driven mostly by patient factors you did not change, then the effect size reaching your chart is a small fraction of the effect size in your process. Meanwhile the process measure counts only eligible patients and only the step you changed — the signal arrives undiluted and unadjusted.

The practical consequence for a project running one or two quarters:

  • Process measures are your primary evidence of change. They are what your run chart is for. Weekly or biweekly points, an explicit eligible-population denominator, and the run-chart rules for distinguishing real signal from noise will tell you within weeks whether the change package is being delivered.
  • Outcome measures are your honesty check, not your proof. Track them, chart them, and state in advance that the project is not powered to move them. A flat outcome chart alongside a shifted process chart is an interpretable, publishable result.
  • Structure measures are usually step functions. A protocol goes live once; a staffing establishment changes once a year. Charting a structure measure over time yields almost no information. Record structure as project context — the state of the system when the intervention ran — rather than as a time series.

Beware the denominator trap on process measures. If the eligible population is not specified with explicit exclusions before you start, apparent improvement is often denominator drift: cases quietly stop being counted as eligible. Fix the denominator definition at the same moment you fix the numerator.

How strong is the causal chain, really?

Two studies are worth knowing precisely, because they are what the framework rests on empirically and because both are more nuanced than the textbook diagram.

Structure to process to outcome, measured end to end. Moore and colleagues (Journal of Trauma and Acute Care Surgery, 2015) evaluated a Canadian provincial trauma system — 57 centres, 63,971 patients, 2005–2010. Structural performance came from on-site accreditation reports scored against American College of Surgeons criteria; process was a composite of conformity to 15 validated process indicators; outcome was risk-adjusted mortality, complications, readmission and length of stay. They found a statistically significant correlation between structure and process of r = 0.33, and between process and outcome of r = −0.33 for readmission and r = −0.27 for length of stay. The authors concluded the model is valid for evaluating trauma care.

Read those coefficients carefully. They are real but modest — structure explains roughly a tenth of the variance in process. And the published correlations between process and outcome are reported for readmission and length of stay; a significant process–mortality correlation is not among the reported findings. The chain holds in aggregate across dozens of centres over five years. That is a very different evidential situation from a single ward over ten weeks.

Structure straight to outcome. The RN4CAST study (Aiken and colleagues, The Lancet, 2014) linked discharge data for 422,730 surgical patients aged 50 and over in 300 hospitals across nine European countries to surveys of 26,516 nurses. Each additional patient in a nurse’s workload was associated with a 7% increase in the odds of a patient dying within 30 days of admission (OR 1.068, 95% CI 1.031–1.106), and every 10% increase in the proportion of nurses holding a bachelor’s degree with a 7% decrease (OR 0.929, 95% CI 0.886–0.973). Two staffing and education variables — pure structure — tracked hospital mortality across nine health systems. That is the strongest available warrant for taking structure measures seriously, and it is also a reminder that detecting the effect took 300 hospitals. The same measures on one unit would show nothing. Staffing-linked structure and outcome pairs of this kind are the substance of nurse-sensitive indicators.

Where balancing measures fit

Teams trained on the Model for Improvement arrive with a different vocabulary: outcome, process and balancing measures. This is not a rival taxonomy and the two do not need reconciling, because they classify on different axes.

Donabedian classifies by position in the causal chain. The Institute for Healthcare Improvement’s family of measures classifies by the role a measure plays in your project. IHI defines outcome measures as showing how the system is working from the patient’s perspective, process measures as showing whether the parts of the system are performing as planned, and balancing measures as answering whether changes designed to improve one part of the system are causing new problems elsewhere — reintubation rates when you shorten ventilator time, readmissions when you shorten length of stay. IHI recommends a set of typically four to ten measures, usually with a single outcome measure tied to the aim.

A balancing measure is therefore not a fourth Donabedian category. It is a structure, process or outcome measure that you selected because of what it might reveal about unintended harm. Reintubation rate is an outcome measure serving a balancing role. Overtime hours are a structure measure serving a balancing role. When you write your methods, say both things: what the measure is, and what job it does. Building the aim, the drivers and the measures together is what a driver diagram exists to force.

The literature’s own imbalance — and what it means for you

A scoping review of the model applied to outpatient care quality (2025) examined 18 studies from Canada, Germany and China and found that all of them addressed all three dimensions — but that the indicator sets were heavily lopsided. In one analysis of paediatric care, 89% of indicators captured process quality, 9% outcome and 2% structure; in another, 78% process, 20% outcome and 2% structure. The reviewers also noted that most of this work stopped at constructing index systems rather than applying them in practice, and called for multicentre validation.

Two implications. First, the near-absence of structure indicators is a genuine gap, not just a fashion — and given the RN4CAST findings, an odd one. If your project changed staffing, skill mix, equipment availability or protocol existence, record it explicitly rather than assuming it is context. Second, the process-heavy pattern is not a defect. It reflects the same signal-detection reality Mant and Hicks quantified. Process dominance is what a measurement plan looks like when it is built to detect change on a realistic timescale.

Reporting the classification

If the project is heading for publication or an internal QI committee, SQUIRE 2.0 — the 18-item Standards for QUality Improvement Reporting Excellence, revised in 2015 and indexed by the EQUATOR Network — is the checklist reviewers will apply. Its Measures item asks for the measures chosen for studying processes and outcomes of the intervention, including the rationale for choosing them, their operational definitions, and their validity and reliability, alongside methods for assessing data completeness and accuracy.

“Rationale” is where the Donabedian classification earns its place. A methods paragraph that says we selected X as our primary process measure because the project could not be powered to detect a change in Y, our outcome measure, within the study period; Y was tracked to confirm no deterioration is stronger than one that simply lists measures under three headings. It shows the team understood what its data could and could not demonstrate.

Where the model runs out

Three limitations matter in practice.

It is drawn as a linear chain and care is not linear. The three domains interact and feed back — outcome data changes protocols, which is outcome acting on structure. Critics have argued the linearity limits the model’s ability to represent how the domains influence each other.

Attribution stays hard. Establishing the relationship between structure, process and outcome in a specific setting is the model’s known weak point. The 1988 paper flags this as the missing information the field needed; the 2015 trauma validation supplies correlations around 0.3. The model tells you what kind of measure you have. It does not license a causal claim you have not earned by design.

The frame predates a lot of modern care. Sociopolitical and socioeconomic determinants, cross-organisation collaboration, and digital health delivery sit awkwardly in a 1966 taxonomy of settings, activities and effects. Adapted versions exist — a systematic review in JMIR addressed applying the model to eHealth — but the base framework does not accommodate them natively.

None of this makes the triad less useful. It makes it a classification scheme with real predictive value about measurement behaviour, rather than a theory of causation you can lean on for a single project.

Frequently asked questions

Is patient satisfaction a process or an outcome measure?

Outcome. It describes an effect of care on the patient, not what the team did. Published applications of the model classify patient satisfaction alongside adverse events and complaint handling as outcome indicators. The interaction that produced the satisfaction is the process; the patient’s reported experience of it is the outcome.

Is staff training completion a structure or a process measure?

Structure. It describes an attribute of the workforce that exists independently of any patient encounter. The corresponding process measure is the proportion of eligible patients who actually received the trained-for action. Reporting training completion as though it were process is one of the most common ways a QI report overstates delivery.

Where do balancing measures fit in the Donabedian model?

They are not a fourth category. Donabedian sorts by position in the causal chain; the IHI family of measures sorts by the role a measure plays in the project. Every balancing measure is also a structure, process or outcome measure. Reintubation rate used to guard against harm from earlier extubation is an outcome measure doing a balancing job.

Do I need a measure in all three categories?

Not necessarily, and forcing one into each box produces filler. What you need is at least one process measure that can actually move within your timeline, an outcome measure tied to your aim that you track honestly even if it will not shift, and an explicit record of the structural conditions the intervention ran under. IHI’s guidance points at a total set of roughly four to ten measures rather than a quota per domain.

My process measure improved but the outcome did not. Did the project fail?

Not necessarily, and that combination is expected rather than anomalous. Mant and Hicks showed that outcome measures such as hospital mortality are insensitive to quality differences that process measures detect readily. The alternative explanations to rule out are that the process you improved is not in fact on the causal path to that outcome, or that fidelity was high but reach was low. Both are answerable from your own data; neither is answered by declaring failure.

Is the Donabedian model the same as a logic model?

No, though they are often drawn similarly. A logic model maps a specific programme’s inputs, activities and intended results, and is built fresh for each programme. Donabedian’s triad is a general classification of quality-of-care measures that applies across programmes. Donabedian himself adapted the industrial input–process–output idea, which is why the two look alike on a whiteboard.

Which came first, the 1966 paper or the 1988 one?

The 1966 Milbank Memorial Fund Quarterly paper introduced the framework; the 1988 JAMA paper gave the operational definitions and the explicit statement about causal linkages that most citations are actually drawing on. Milbank reprinted the 1966 paper in full in 2005, which is why you will see both dates attached to the same work.

References

  • Donabedian A. Evaluating the quality of medical care. Milbank Memorial Fund Quarterly 1966;44(3 Suppl):166–206. Reprinted in Milbank Quarterly 2005;83(4):691–729. PubMed 16279964
  • Donabedian A. The quality of care. How can it be assessed? JAMA 1988;260(12):1743–8. PubMed 3045356
  • Mant J, Hicks N. Detecting differences in quality of care: the sensitivity of measures of process and outcome in treating acute myocardial infarction. BMJ 1995;311(7008):793–6. PMC2550793
  • Moore L, Lavoie A, Bourgeois G, Lapointe J. Donabedian’s structure-process-outcome quality of care model: validation in an integrated trauma system. Journal of Trauma and Acute Care Surgery 2015;78(6):1168–75. PubMed 26151519
  • Aiken LH, Sloane DM, Bruyneel L, et al. Nurse staffing and education and hospital mortality in nine European countries: a retrospective observational study. The Lancet 2014;383(9931):1824–30. PMC4035380
  • Application of the Donabedian three-dimensional model in outpatient care quality: a scoping review. PMC12045680
  • Institute for Healthcare Improvement. Model for Improvement: Establishing Measures. ihi.org
  • Ogrinc G, Davies L, Goodman D, et al. SQUIRE 2.0 (Standards for QUality Improvement Reporting Excellence): revised publication guidelines from a detailed consensus process. EQUATOR Network · PMC5411027
  • Busse R, Klazinga N, Panteli D, Quentin W (eds). Improving healthcare quality in Europe: characteristics, effectiveness and implementation of different strategies. European Observatory on Health Systems and Policies. NCBI Bookshelf NBK549261

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.