Skip to main content
v2026.11,610 entries · CC-BY 4.0

Building a Programme Logic Model That Survives Evaluation

Most logic models fail at evaluation time not because the inputs-activities-outputs-outcomes-impact chain is wrong, but because each stage got filled with a description instead of a countable indicator. This guide covers the vagueness trap, the four tests a measurable indicator has to pass, and a full worked example.

Ask about Building a Programme Logic Model That Survives Evaluation

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

A programme logic model is a one-page diagram of a causal chain: the inputs a project commits, the activities it runs, the outputs those activities produce, the outcomes that follow for the people or systems involved, and the longer-term impact those outcomes are meant to add up to. Most logic models look fine in the planning meeting where they get drawn and then fail quietly at evaluation time — not because the five-box chain is the wrong shape, but because most of the boxes were filled in with a description of what the project intends to do rather than a stated, countable indicator of it. “Increase awareness” and “build capacity” are activities a reader can picture; they are not things an evaluator can measure, and a logic model whose outcome row can’t be measured is a logic model that can’t be evaluated, however coherent the diagram looks in a grant narrative. This guide covers what belongs in each stage of the chain, the specific vagueness trap that breaks evaluability, and a full worked example carrying a genuinely measurable indicator end to end.

The five-stage chain, and what each box actually has to contain

Every stage in a logic model answers a different question, and each one needs a different kind of evidence to be evaluable rather than just plausible.

Inputs

Question it answers: what is committed to the project before anything happens.

What makes it measurable: a quantified resource commitment, not a description — staff time in FTE or hours, a dollar budget, licensed materials, or access to an existing dataset or system. “Adequate staffing” is not an input in the evaluable sense; “2.0 FTE for 12 months” is.

Activities

Question it answers: what the project actually does with those inputs.

What makes it measurable: a countable, scheduled unit of delivery — sessions run, sites visited, materials distributed — against a stated denominator of what was planned. “Provide training and support” is not an activity indicator; “deliver four 90-minute sessions per site within 8 weeks” is.

Outputs

Question it answers: the direct, immediate product of the activities — what came out the other end.

What makes it measurable: a count with an explicit eligible-population denominator, so the number can be read as a rate, not a raw total. “People were trained” is not an output indicator; “108 of 120 eligible staff completed training and passed the competency check” is — the denominator is what turns a count into a number anyone else can judge.

Outcomes

Question it answers: what changed for the people, practices, or systems the project touched, usually split into short-term (knowledge, skills, immediate behaviour), intermediate (sustained practice or policy change), and sometimes long-term outcome tiers.

What makes it measurable: a baseline value, a target value, and a timeframe, all attached to the same indicator. A number with no baseline can’t show change; a target with no timeframe can’t be checked; an indicator collected only once can’t distinguish the project’s effect from normal variation.

Impact

Question it answers: the longer-term, higher-level result the outcomes are meant to add up to — usually measured at a population, system, or organisational level, over a longer horizon and often with other contributing factors in play.

What makes it measurable: the same baseline-target-timeframe discipline as outcomes, plus an honest acknowledgement of attribution limits — a single project rarely gets to claim sole credit for a population-level number, and a defensible logic model says so rather than implying a stronger causal claim than the evaluation design can support.

Outputs vs. outcomes: the mistake that breaks the most logic models

The single most common way a logic model fails at evaluation time is not a missing box — it is an output dressed up as an outcome. “120 staff were trained” is an output: it counts what the project delivered. It says nothing about whether those staff changed what they actually do. “95% of eligible admissions received a fall-risk screen within 4 hours” is an outcome: it measures a change in practice, not a count of training delivered. A project can hit every output target — full attendance, every session delivered on schedule, every material distributed — and still produce no outcome change at all, if the activities didn’t actually move behaviour. Reviewers who have seen this pattern before will look specifically for at least one outcome-level indicator that is not simply a restatement of an output count with a percentage sign attached.

The vagueness trap: what “genuinely measurable” actually requires

A logic-model indicator survives evaluation only if it clears four tests, all at once:

  • A countable unit and, where relevant, a denominator — a rate or proportion, not a bare count that has no population to be judged against.
  • A data source that already exists or can realistically be built before the indicator is due, not one the project hopes to invent later.
  • An explicit target value — a number, not a direction. “Increase” is a direction; “increase from 71% to at least 90%” is a target.
  • An explicit timeframe tied to when the indicator will actually be measured, not left implicit until the final report is due.

The table below applies those four tests to one indicator per stage — the vague version is the kind of language that survives a planning meeting, and the fixed version is what an evaluator can actually check against real data.

Stage Vague version (unevaluable) Measurable version
Inputs “Adequate staffing and funding secured” “2.5 FTE funded for 12 months; $40,000 materials budget confirmed at launch”
Activities “Deliver training and ongoing support” “12 scheduled sessions delivered across 3 sites within 8 weeks (12/12 = 100%)”
Outputs “Staff were trained” “108 of 120 eligible staff (90.0%) trained and passed the competency check”
Outcomes “Practice improved” “Screening within 4 hours of admission rises from a 70.9% baseline to ≥90% within 6 months”
Impact “Fewer adverse events” “Fall rate per 1,000 patient-days falls by ≥25% relative to baseline within 12 months”

Assumptions and external factors: the columns most templates leave out

A logic model that only lists inputs through impact is telling half the causal story. Two things sit underneath the chain and rarely get their own column, and their absence is exactly what makes a null result impossible to interpret afterward:

  • Assumptions — the reasons the project believes activities will actually produce outputs, and outputs will actually produce outcomes. “Staff who pass the competency check will apply it at the bedside” is an assumption, not a certainty, and it is worth stating and, where possible, checking directly rather than only inferring it from the outcome number.
  • External factors — conditions outside the project’s control that could move the same indicators regardless of what the project does. A concurrent staffing shortage, an unrelated policy change, or a parallel initiative touching the same population can all shift an outcome or impact number independently of the project being evaluated.

Recording both explicitly does real evaluation work: if an outcome indicator doesn’t move, the assumptions and external-factors rows are the first place to look for why — a failed assumption, a confounding external event, or an activity that genuinely didn’t work are three different findings that call for three different responses, and a logic model without those columns can’t tell them apart.

Worked example: a fall-prevention training programme logic model

Illustrative composite — not a real institution. The scenario, staffing figures, dates, and results below are a constructed example built to show a complete, internally consistent logic model. They are not attributed to any specific hospital, health system, or published study.

A hospital nursing service runs a 12-month pilot to reduce inpatient falls on three medical-surgical units by training staff on a standardized fall-risk screening and care-planning protocol.

Inputs (committed at launch): 2.5 FTE clinical educator time for 12 months; a $40,000 curriculum-licensing and printing budget; an existing EHR fall-risk screening module configured for the three pilot units.

Stage Indicator Baseline Target Result
Activities Scheduled training sessions delivered ÷ sessions planned 100% delivered within 8 weeks 12 of 12 delivered (100%)
Output Staff trained and passing competency check ÷ eligible staff 0% ≥85% within 8 weeks 108 / 120 = 90.0%
Outcome (short-term) Admissions screened within 4 hours ÷ total admissions 781 / 1,102 = 70.9% (pre-pilot 6-month period) ≥90% within 6 months of rollout 1,056 / 1,109 = 95.2% (post-rollout 6-month period)
Outcome (intermediate) Standardized care plan documented ÷ patients flagged high-risk Not tracked pre-pilot — the standardized flag did not exist before rollout ≥95% within 6 months 402 / 412 = 97.6%
Impact Falls ÷ patient-days × 1,000 (unit-reported) 55 falls / 30,400 patient-days = 1.81 per 1,000 (pre-pilot 6-month period) ≥25% relative reduction within 12 months 37 falls / 31,200 patient-days = 1.19 per 1,000 — a 34.5% relative reduction

Assumption: staff who pass the competency check will actually apply the screening and care-planning steps at the bedside, not just on the test — the intermediate care-plan-documentation indicator exists specifically to check that assumption directly, rather than inferring it from the fall-rate number alone. External factor: unit staffing ratios and any concurrent hospital-wide initiative touching the same units over the same 12 months were tracked alongside the pilot, since either could move the fall rate independently of the training programme.

Notice what the table does that a narrative description would not: every row has a denominator, a baseline (or an honest note that none existed), a target set before the result was known, and a timeframe. That is what lets a reader check the claim against the number, rather than take the project’s word for “falls went down.” This is also why the fall-rate-per-1,000-patient-days framing matters more than a raw fall count — see NDNQI nursing quality indicators for how unit-level nurse-sensitive indicators like this one get defined and benchmarked, and nurse-sensitive indicators for the broader measure set this kind of outcome sits inside.

Common ways a logic model becomes unevaluable after it’s built

  • Output/outcome conflation — reporting an activity count with a percentage sign on it and calling it a result.
  • An indicator with no denominator — “100 people reached” means nothing without stating the eligible population it’s a share of.
  • No baseline captured before the project starts — without a pre-project number, there is no way to show the project moved anything, only that a number existed afterward.
  • A target with no timeframe — “will improve” can never be checked; “will reach X by month 6” can.
  • Skipping assumptions and external factors — so a flat or negative outcome result can’t be attributed to a failed assumption, an external shock, or a genuinely ineffective activity, because nothing was recorded that could distinguish the three.
  • Impact claimed at a level the evaluation design can’t support — a single-site pilot claiming population-level or systemic impact is a design mismatch, not a measurement problem, and it needs a smaller, honestly-scoped impact claim or a larger evaluation design, not a more confident sentence.

From logic model to evaluation plan

Once every stage carries a real indicator, the logic model turns directly into the skeleton of an evaluation plan: each indicator becomes one row specifying its data source, collection method, collection frequency, and who is responsible for it. The logic model itself doesn’t choose an evaluation design or explain why an intervention did or didn’t work — it states what changed and by how much. For the determinants side of that question — the barriers and facilitators that explain an implementation result — see the Consolidated Framework for Implementation Research (CFIR). For a five-dimension outcomes framework built specifically around real-world reach and adoption, see the RE-AIM framework. Planning-stage frameworks such as PRECEDE-PROCEED and evidence-appraisal pathways such as the Iowa Model and the Johns Hopkins Nursing EBP Model sit alongside a logic model rather than replacing it — a logic model states the causal chain and its indicators; these frameworks structure how you plan, appraise evidence for, or explain the implementation around that chain. Where a project also needs to demonstrate feasibility before a full rollout, see pilot study. Logic models are also a standard attachment in funded-project reporting — see grant narrative / project narrative for how the same inputs-to-outcomes chain gets written into a funder-facing document.

Frequently asked questions

What’s the difference between an output and an outcome in a logic model?

An output is a direct count of what the project delivered — sessions run, people trained, materials distributed. An outcome is a measured change in behaviour, practice, knowledge, or condition among the people or systems the project touched. A project can produce every planned output and still show no outcome change; the two are evaluated separately, never collapsed into one number.

What’s the difference between a logic model and a theory of change?

They’re related but not the same tool. A theory of change is the explicit causal narrative — the reasoning for why a set of activities is expected to lead to a given outcome, including the assumptions that reasoning depends on. A logic model is typically the more compact, structured diagram of the resulting inputs-activities-outputs-outcomes-impact chain, often built after the theory of change has worked out the underlying causal logic. Many funders ask for both; a logic model without an underlying theory of change is a diagram with no stated reason to believe it will work.

How many indicators does each stage need?

One well-specified indicator per stage is enough to make a logic model evaluable; more than two or three per stage usually signals the model is trying to measure everything rather than the things that actually matter for deciding whether the project worked. Depth should go into making each indicator genuinely measurable — denominator, data source, target, timeframe — rather than into adding more of them.

Do I need a logic model before I can write an evaluation plan?

Effectively, yes. An evaluation plan specifies how each indicator will be measured, by whom, and how often — and none of that can be written until the logic model has already stated what those indicators are. Building the two in parallel usually means the evaluation plan gets rewritten once the logic model’s vague spots get caught.

What if I don’t have baseline data before the project starts?

Say so explicitly in the logic model rather than skipping the baseline row silently. A genuinely unavailable baseline (a new measure with no pre-project data, as in the worked example’s care-plan-documentation indicator above) is a real, common evaluability limitation, not a flaw to hide — and it changes what the eventual result can claim: a post-only number can describe where the project ended up, but it can’t, on its own, demonstrate that the project caused a change.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.