Written and maintained by CASRAI Editorial Board
Last updated
Realist evaluation is an approach to evaluating complex interventions — programmes, policies, clinical services, organisational change efforts — that asks a different question than most evaluation designs start with. Instead of does it work?, it asks what works, for whom, in what circumstances, and through what mechanism? The method was developed by Ray Pawson and Nick Tilley and set out in their 1997 book Realistic Evaluation, and its organising device is the context-mechanism-outcome (CMO) configuration: a claim that a given mechanism fires (or doesn’t) in a given context, producing a given outcome. This guide works through what each element of a CMO configuration actually is, why that logic is the thing that separates realist evaluation from outcome-only evaluation designs, and a fully worked, explicitly illustrative example.
What realist evaluation is — and what it isn’t
Realist evaluation belongs to a family of theory-driven evaluation approaches: rather than treating an intervention as a black box and simply comparing outcomes between an intervention group and a control group, it tries to open the box and explain why the intervention produced the pattern of outcomes it did. That explanatory ambition is what it shares with determinant frameworks like the Consolidated Framework for Implementation Research (CFIR) — but realist evaluation is older, methodologically distinct, and organised specifically around testing CMO configurations rather than coding a fixed list of constructs.
Two mix-ups are worth clearing up before anything else, because both are common:
- Realist evaluation is not realist synthesis (realist review). Realist synthesis applies the same CMO logic to a body of existing literature, building and refining programme theory by synthesising published studies of similar interventions across different settings — that is secondary research. Realist evaluation applies the identical logic to primary data collected about one specific intervention, in one specific setting, as it actually ran. See Realist Review: What It Is and the RAMESES Reporting Standards for the synthesis side of the family; this guide is about the primary-evaluation side. They share an intellectual ancestor (Pawson) and a reporting-standards project (RAMESES), but a realist evaluation is not “a realist review of one study” — it’s a different research act with its own data-collection logic.
- Realist evaluation is not a synonym for “qualitative evaluation.” RAMESES II, the reporting-standards project for realist evaluations, explicitly frames them as typically mixed-methods studies: quantitative outcome data is often exactly what tells you a CMO configuration held or didn’t, while qualitative data (interviews, observation, documents) is usually what lets you identify the mechanism and characterise the context in the first place. Realist evaluation is a logic of inquiry, not a ban on numbers.
The CMO configuration, precisely
A CMO configuration is a single proposition with three linked parts:
- Context (C) — the pre-existing conditions that determine whether a mechanism can fire: who the participants are, what resources and constraints they already have, the institutional setting, the social and cultural conditions surrounding the intervention. Context is not “background information” in realist evaluation — it is a causally active ingredient, on equal footing with the intervention itself.
- Mechanism (M) — not the intervention, and this is the distinction people new to the method most often get wrong. The intervention is a resource or opportunity offered into a context; the mechanism is the reasoning, decision, or capacity change that happens inside the people (or organisation) exposed to that resource, in response to it. Pawson and Tilley call this generative causation: a programme doesn’t cause an outcome directly the way a physical force does — it works by giving people reason, opportunity, or capacity to act differently, and it’s that internal response, not the programme itself, that generates the outcome.
- Outcome (O) — the result, which realist evaluation treats as plural and pattern-shaped from the start: intended outcomes, unintended outcomes, and outcomes that vary systematically by subgroup, because different subgroups sit in different contexts.
The configuration itself is written as a single testable proposition: in context C, mechanism M is triggered (or blocked), producing outcome O. A realist evaluation of one programme typically produces several CMO configurations, not one — because the same intervention, offered into different contexts, is expected to trigger different mechanisms and therefore produce different outcomes for different people. That expectation of legitimate, explainable variation is the method’s central commitment, and it’s also what most directly separates it from the evaluation designs covered next.
Why this differs from outcome-only evaluation frameworks
Several evaluation frameworks already covered on this site are built to measure whether and how much an outcome changed. They are not wrong or inferior instruments — they answer a different question than CMO logic does, and realist evaluation is often used alongside them rather than instead of them:
- RE-AIM scores an intervention along five outcome-and-implementation dimensions (Reach, Effectiveness, Adoption, Implementation, Maintenance), each with its own denominator. It tells you the size and spread of what happened. It does not, by itself, explain why Effectiveness was high in one clinic and flat in another — that’s a CMO-shaped question.
- A programme logic model maps inputs through activities and outputs to outcomes and impact in a single assumed causal chain. It documents what the programme intends to do and in what order; it doesn’t specify why the chain would hold for some participants and break for others, because a single chain has no room for context-dependent variation by design.
- A conventional RCT or pre/post outcome evaluation reports an average treatment effect (or a null result) across the whole sample. Realist evaluation treats that single average as potentially hiding the more useful finding: the intervention may have worked well for one subgroup, done nothing for a second, and backfired for a third — and those three results, averaged together, can look identical to “no effect” or “a modest effect” everywhere.
CMO logic’s distinguishing move is refusing to ask “did it work” as a single yes/no question and instead building an explicit theory of which contexts activate which mechanisms to produce which outcomes — then testing that theory against data, configuration by configuration, rather than testing a single averaged hypothesis.
A worked CMO configuration example
The following is an illustrative composite, constructed to demonstrate CMO configuration logic. It does not describe a real institution, a real programme, or real study data — no metrics below are drawn from any actual evaluation.
Consider a hospital outpatient clinic that introduces a text-message medication-reminder service for patients on a chronic-disease regimen. An outcome-only evaluation would ask: did adherence go up? A realist evaluation instead builds and tests separate CMO configurations for the subgroups the initial programme theory (drawn from staff interviews and prior literature) suggested would respond differently:
| Context (C) | Mechanism (M) | Outcome (O) |
|---|---|---|
| Patient already intends to take medication, has a stable phone number and reliable signal, but is prone to simply forgetting amid a busy schedule. | The text arrives as a well-timed cue that closes an existing gap between intention and action — it doesn’t create motivation, it prompts an already-motivated person at the right moment. | Adherence improves. The mechanism (cue-triggered recall) fires because the context (existing intention + reliable phone access) allows it to. |
| Patient has unstable housing and a prepaid phone with irregular credit, so texts arrive late, arrive in a batch, or don’t arrive at all. | The cueing mechanism never has the chance to fire — not because reminders don’t work in principle, but because the delivery channel is unreliable in this context. | No measurable change in adherence for this subgroup — a null result that is not evidence the intervention “doesn’t work,” but evidence that this context blocks the mechanism from being triggered at all. |
| Patient is managing a condition they experience as stigmatising and has not disclosed it to people who share their phone or household. | An unsolicited, visible text about the condition triggers a privacy/exposure concern rather than a helpful reminder — a threat response instead of a cueing response. | Adherence for this subgroup may worsen, or the patient may opt out of the service entirely, even though the same message triggered a helpful mechanism for the first subgroup. |
Averaged across the whole clinic population, these three configurations could easily net out to “a small, statistically marginal improvement in adherence” — the kind of result an outcome-only evaluation reports and then struggles to act on. The CMO breakdown gives the clinic something an average can’t: a testable explanation of who the service actually helps, who it’s neutral for, and who it may be actively harming, plus a concrete lever for each (fix the delivery channel for the second group; make the service opt-in and use less identifiable message wording for the third).
Building and testing CMO configurations
Realist evaluation proceeds iteratively rather than in a single measurement pass:
- Articulate an initial, rough programme theory. Drawn from existing literature, programme documents, and interviews with staff and participants about how they believe the intervention is supposed to work. This produces a first, provisional set of candidate CMO configurations — deliberately treated as hypotheses, not findings.
- Collect data capable of testing context, mechanism, and outcome separately. This typically means mixed methods: realist interviews (which probe participants’ own reasoning about why they did or didn’t respond to the intervention) alongside quantitative outcome data, documentary analysis, and observation, chosen specifically because a single data source rarely evidences all three elements of a configuration at once.
- Use retroduction to explain the pattern in the data. Retroductive reasoning works backward from an observed outcome pattern to the underlying mechanism that plausibly generated it, given the context in which it occurred — the inferential move that separates realist analysis from simply describing what happened.
- Refine, retain, or reject each configuration. Some initial configurations survive testing largely intact; others need their context or mechanism component revised to fit the data; some are abandoned because no evidence supports the proposed mechanism actually firing. The result is a refined, evidenced set of CMO configurations, not a single verdict on the programme.
- State the programme theory at a middle range. Realist evaluation deliberately avoids two failure modes: a theory so abstract it explains nothing about this specific programme, and a theory so tied to this one setting’s specifics that it offers no transferable lesson. The aim is a middle-range theory — specific enough to be genuinely testable, general enough to inform a similar intervention run somewhere else.
Reporting a realist evaluation
Realist evaluations have their own dedicated reporting standard, distinct from the standard used for realist syntheses. RAMESES II (Wong G, Westhorp G, Manzano A, Greenhalgh J, Jagosh J, Greenhalgh T. “RAMESES II reporting standards for realist evaluations.” BMC Medicine 2016;14:96) sets out what a realist-evaluation write-up needs to include — among other things, the initial programme theory the evaluation set out to test, the data sources used to test each element of a configuration, and the refined CMO configurations themselves, stated explicitly rather than left implicit in the discussion section. It is a companion project to the original RAMESES standards for realist syntheses, built by an overlapping author team, but the two checklists are not interchangeable: using the synthesis checklist to report a primary evaluation (or vice versa) will leave reviewers looking for items that don’t apply to what was actually done.
Common pitfalls
- Treating the intervention itself as the mechanism. “The text message” is not a mechanism; “a timely cue closing an intention-action gap” is. If a CMO configuration’s mechanism column could be answered by just re-describing the programme activity, it hasn’t been specified yet.
- Writing context as a demographic checklist. Age, sex, and diagnosis can be part of context, but context in realist terms is whatever conditions determine if the mechanism can fire — that can be organisational (staff turnover, referral pathways), relational (trust in the referring clinician), or resource-based (phone reliability in the worked example above), and a CMO configuration that never specifies this is doing outcome evaluation with extra vocabulary, not realist evaluation.
- Collapsing multiple configurations into one average finding. The entire value of the method is lost if the write-up reverts to reporting a single pooled result once the analysis is done — the configurations, stated explicitly and separately, are the deliverable.
- Skipping the initial programme theory. A realist evaluation that starts data collection without an explicit, falsifiable starting theory has nothing for the data to test against, and typically produces post-hoc storytelling rather than genuinely tested configurations.
Frequently asked questions
Is realist evaluation qualitative or quantitative?
Neither exclusively — it’s most often mixed-methods. Quantitative outcome data commonly establishes whether a proposed configuration’s outcome actually occurred; qualitative data commonly establishes why, by surfacing the mechanism and characterising the context. RAMESES II frames realist evaluations as typically mixed-methods studies rather than a purely qualitative approach.
How is a mechanism different from the intervention?
The intervention is the resource or opportunity offered (a text message, a training session, a new referral pathway). The mechanism is the change in reasoning, decision-making, or capacity that happens inside the people exposed to that resource, which is what actually generates the outcome. The same intervention can trigger different mechanisms — or none at all — depending on context, which is exactly why one intervention produces different outcomes for different people.
How many CMO configurations should an evaluation produce?
There’s no fixed target. The number should be driven by how many genuinely distinct context-mechanism-outcome patterns the data actually supports, not by a template. A small, well-evidenced set of configurations that explains real variation in outcomes is more useful than a large set padded out to look thorough.
Is realist evaluation the same as a realist review or realist synthesis?
No. Realist synthesis (realist review) applies CMO logic to published literature across multiple studies as a form of evidence synthesis; realist evaluation applies it to primary data collected about one specific, real-world implementation. See Realist Review: What It Is and the RAMESES Reporting Standards for the synthesis side of the method.
Can realist evaluation be combined with frameworks like CFIR or RE-AIM?
Yes, and it commonly is. RE-AIM or a logic model can establish that an outcome pattern varied; realist evaluation’s CMO logic is then used to explain why it varied, in terms of the contexts and mechanisms at work. CFIR and realist evaluation both aim to explain rather than just measure, but CFIR codes against a fixed construct list, while realist evaluation builds and tests configurations specific to the programme theory at hand.








