Skip to main content
v2026.11,772 entries · CC-BY 4.0

Proctor’s Implementation Outcomes: All Eight, and How to Measure Each

Proctor’s eight implementation outcomes — acceptability, adoption, appropriateness, feasibility, fidelity, implementation cost, penetration, and sustainability — with the validated instrument or measurement method for each.

Written and maintained by CASRAI Editorial Board

Last updated

When a reviewer asks “which outcome are you measuring,” a research-methods section that answers with a symptom score or a satisfaction rating has answered the wrong question. Proctor and colleagues’ 2011 taxonomy names eight outcomes that belong to implementation itself, not to the treatment being implemented or the service system delivering it — and gives each one a distinct definition so “it worked” stops doing all the work in a methods section.

Why implementation outcomes are a separate category

Proctor’s framework sits between two other outcome levels that grant and journal reviewers already expect to see, and confusing the three is one of the more common reasons an implementation-science methods section gets sent back:

  • Implementation outcomes — the eight below. They answer “did the implementation effort succeed,” independent of whether the underlying treatment works.
  • Service outcomes — the Institute of Medicine’s STEEEP dimensions: efficiency, safety, effectiveness, equity, patient-centeredness, timeliness. These describe the service system’s performance once the innovation is in place.
  • Client/patient outcomes — satisfaction, function, symptomatology, and similar clinical or consumer endpoints. These describe whether the treatment itself worked for the people who received it.

The point of separating them: a program can be implemented with high fidelity and full adoption (implementation success) while the underlying treatment produces no measurable clinical benefit (a service/client-outcome question) — or the reverse, a genuinely effective treatment can fail to reach anyone because it was never actually adopted. Conflating the two means a study can’t tell which failure occurred.

The eight outcomes, and how each is actually measured

Three of the eight now have short, psychometrically validated self-report instruments; the other five are measured structurally — by counting, checklisting, or costing — because they describe organizational or process facts rather than perceptions.

Acceptability

Definition: the perception among stakeholders that a given innovation is agreeable, palatable, or satisfactory.

How it’s measured: the Acceptability of Intervention Measure (AIM) — a 4-item self-report scale developed and psychometrically tested by Weiner et al. (2017) in Implementation Science. Respondents (providers, staff, or clients) rate agreement with items like “this intervention meets my approval.” AIM, its companion measures below, and feasibility are described in the implementation-science literature as “leading indicators” — they can be collected early, before adoption or fidelity data exist, giving a warning sign well before a program stalls.

Appropriateness

Definition: the perceived fit, relevance, or compatibility of the innovation for a given practice setting, provider, or consumer — distinct from acceptability, which is about whether people like it, not whether it fits.

How it’s measured: the Intervention Appropriateness Measure (IAM), also a 4-item scale from the same Weiner et al. (2017) validation study, developed alongside AIM and FIM specifically because appropriateness, acceptability, and feasibility are conceptually distinct enough that a single combined score obscures which one is failing.

Feasibility

Definition: the extent to which an innovation can actually be used or carried out within a given agency or setting, given its real staffing, workflow, and resource constraints.

How it’s measured: the Feasibility of Intervention Measure (FIM), the third 4-item scale from Weiner et al. (2017). AIM, IAM and FIM were built and tested together — 36 implementation scientists and 27 mental-health professionals independently sorted a 31-item pool back to its intended construct, and all three measures showed acceptable-to-strong internal consistency and discriminant validity in the original study.

Adoption

Definition: the intention, initial decision, or action to try or employ an innovation — the uptake decision itself, at the individual-provider or organizational level.

How it’s measured: not by a self-report attitude scale, but structurally — administrative or EHR-based uptake counts (how many eligible providers or units started using the practice), or by tracking progression through the Stages of Implementation Completion (SIC), a phased milestone framework (Saldana et al.) that treats adoption as a documented event a site passes through rather than an opinion a respondent reports.

Fidelity

Definition: the degree to which an intervention was delivered as prescribed in the original protocol or design.

How it’s measured: against an intervention-specific fidelity checklist, but the widely cited structural framework for building one is the NIH Behavior Change Consortium’s treatment-fidelity model (Bellg et al., 2004), which breaks fidelity into five components a study should address separately: study design, provider training, treatment delivery, treatment receipt, and enactment of skills by the recipient. A single global “was it done right” rating collapses distinctions the BCC framework treats as separately actionable — training fidelity and delivery fidelity fail for different reasons and need different fixes.

Implementation cost

Definition: the cost impact of an implementation effort, distinct from the cost of the treatment itself — what it actually costs an organization to get a practice adopted and running, not to deliver it once running.

How it’s measured: the Cost of Implementing New Strategies (COINS) method (Saldana et al., 2014, Children and Youth Services Review), which maps costs and staff time onto the same SIC milestone stages used for adoption — pre-implementation activities (engagement calls, partner meetings, feasibility and billing determinations), active implementation, and sustainment each get their own cost line, rather than one lump “implementation cost” figure that hides where the money actually goes.

Penetration

Definition: the extent to which an innovation is integrated within a service setting — operationally, the proportion of the eligible population that actually receives it, once the practice is in place.

How it’s measured: as a calculated ratio (recipients ÷ eligible population), not a survey instrument. This is the outcome most often confused with RE-AIM’s Reach dimension; the two overlap but aren’t identical — Reach in RE-AIM is typically scoped to a specific program’s participant pool, while penetration in Proctor’s framework is scoped to the full eligible population within a service setting, including people the program never contacted at all. See the RE-AIM framework guide for how Reach is defined and reported there.

Sustainability

Definition: the extent to which a newly implemented treatment is maintained or institutionalized within a service setting’s ongoing, stable operations — implementation outcomes don’t end at go-live.

How it’s measured: the Program Sustainability Assessment Tool (PSAT) (Luke et al., 2014), a 40-item instrument spanning 8 sustainability-capacity domains (5 items each) — organizational capacity, program adaptation, program evaluation, communications, strategic planning, funding stability, partnerships, and environmental support. It has been used to rate more than 2,500 public health, clinical, and social-service programs since publication, and a validated shorter version exists for lower-burden re-administration.

Reviewers most often flag two things: the wrong outcome, and the wrong instrument

Two mistakes recur in implementation-science manuscripts and grant proposals:

  • Reporting a service or client outcome and calling it an implementation outcome. “Patients reported high satisfaction” is a client outcome; it says nothing about acceptability (did providers find the intervention agreeable to deliver), fidelity, or penetration. State explicitly which of the eight is being reported, and which tier (implementation / service / client) it belongs to.
  • Using a global “success” rating in place of the specific outcome that actually failed. A program can score well on acceptability and appropriateness (people like it and think it fits) while failing on feasibility (they can’t actually staff it) — a combined score would average that out and hide the real barrier. Report each outcome that was measured separately, not a composite.

This taxonomy is not the same project as Implementation Mapping‘s five-task planning process (which designs strategies before rollout) or the CDC Program Evaluation Framework‘s six-step evaluation cycle (which evaluates a program broadly, not implementation specifically) — Proctor’s eight outcomes are what those and other planning frameworks are ultimately trying to produce or improve, and are the vocabulary a methods or evaluation section should report results against regardless of which planning framework was used to design the rollout.

Reporting checklist for a methods section or grant aim

  • Name which of the eight outcomes are being measured in this study — not all eight are relevant to every study, and claiming all eight without measuring them invites a reviewer to ask for the missing data.
  • For acceptability, appropriateness, and feasibility, report the instrument (AIM/IAM/FIM), respondent group (provider, staff, client, or organizational leader — the same instrument can be administered to different stakeholders with different results), and timing (early leading-indicator administration vs. post-implementation).
  • For fidelity, specify which of the five BCC components (design, training, delivery, receipt, enactment) the study’s fidelity checklist actually covers — most studies cover delivery and skip enactment, which is worth stating rather than implying full coverage.
  • For penetration, report both the numerator and denominator used (recipients and eligible population), since the same penetration percentage means something different depending on how narrowly “eligible” was defined.
  • For sustainability, report the assessment timepoint relative to when active implementation support ended — a PSAT score collected while the original implementation team is still on-site measures something different from one collected a year after they’ve left.

Frequently asked questions

Is penetration the same as RE-AIM’s Reach?

Related but not identical. Both describe how much of an eligible population is actually receiving an innovation, but they’re typically scoped differently — RE-AIM’s Reach is usually reported against a program’s own participant pool, while Proctor’s penetration is scoped to the full eligible population within the service setting. State explicitly which denominator was used; the two numbers are not interchangeable even when they look similar.

Do all eight outcomes need to be measured in every implementation study?

No. Proctor’s own framing treats these as conceptually distinct constructs to choose among, not a mandatory checklist. A pilot feasibility study typically reports acceptability, appropriateness, and feasibility (the three “leading indicator” self-report measures); a longer-running effectiveness-implementation hybrid trial is more likely to add fidelity, adoption, and eventually sustainability once there’s something to sustain.

What’s the difference between fidelity and adherence?

In most implementation-science usage they’re treated as closely related, with fidelity as the broader umbrella term and adherence (along with competence) as components of it under frameworks like the NIH BCC model — adherence asks whether the prescribed steps were delivered; competence asks how skillfully. A study reporting “high fidelity” without specifying whether that means adherence, competence, or both is under-specifying its own claim.

See also: RE-AIM Framework, Implementation Mapping, Hybrid Effectiveness-Implementation Trial Designs, Knowledge-to-Action Framework, Translational Research, and the Research Methods pillar.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about Proctor’s Implementation Outcomes: All Eight, and How to Measure Each

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.