Skip to main content
v2026.11,610 entries · CC-BY 4.0

The CDC Framework for Program Evaluation in Public Health: The 6 Steps and 4 Standards

CDC’s six-step framework treats stakeholder engagement as gating, not a formality, then sequences describing the program, focusing the design, gathering evidence, justifying conclusions, and ensuring use — filtered throughout by four standards: utility, feasibility, propriety, and accuracy.

Ask about The CDC Framework for Program Evaluation in Public Health: The 6 Steps and 4 Standards

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

The CDC Framework for Program Evaluation in Public Health is a six-step process for planning and conducting an evaluation of a public health program, published by the Centers for Disease Control and Prevention in 1999 (MMWR 48, No. RR-11) and still the process CDC recommends to grantees, state and local health departments, and program managers today. It is deliberately generic — it does not prescribe a specific study design, statistical method, or outcome measure — because it is meant to fit programs as different as a tobacco-cessation campaign, an immunization outreach effort, and a chronic-disease surveillance system. What it gives instead is a planning sequence: six steps, applied recursively rather than once, filtered at every step through four standards the evaluation has to satisfy.

This guide walks through the six steps in the order CDC presents them, the four standards that run underneath all six, and how the sequence differs from adjacent frameworks — a logic model, RE-AIM, and PRECEDE-PROCEED — that evaluators often reach for at the same time and sometimes conflate with this one.

What the framework is, and what it isn’t

The framework is a process model, not an outcomes model: it tells you what to do and in what order, not what to measure. That distinction matters because the three frameworks most often confused with it each solve a different problem:

  • A logic model is a single planning artifact — typically inputs, activities, outputs, outcomes, and impact laid out in one diagram — that gets produced inside Step 2 (Describe the Program) below, not a competing framework.
  • RE-AIM is an outcomes-measurement framework: it specifies five dimensions (Reach, Effectiveness, Adoption, Implementation, Maintenance) to score once you already know what you’re measuring. It slots into Step 4 (Gather Credible Evidence) as one way to structure the evidence you collect.
  • PRECEDE-PROCEED is a program-planning model, used before a program is designed, to work backward from a health outcome to the behavioral, environmental, and educational factors a program should target. The CDC framework picks up once a program already exists (or is far enough along in design) to be described and evaluated.

In practice, an evaluator commonly uses PRECEDE-PROCEED to design the program, builds a logic model as part of describing it, uses the CDC framework to sequence the evaluation itself, and uses RE-AIM (or another outcomes framework) inside that sequence’s evidence-gathering step. They are not competitors; they operate at different points in the same project.

The six steps, as a planning sequence

CDC presents the six steps as a cycle, not a line: later steps routinely send you back to re-engage stakeholders or re-describe the program as the evaluation’s real shape becomes clearer. Treat the order below as where to start, not a rule against revisiting an earlier step.

1. Engage stakeholders

This is the step evaluators most often shortcut, and CDC treats it as gating rather than procedural — the framework’s own guidance is that an evaluation designed without the people who will act on its findings is unlikely to change anything, however methodologically sound it is. Three stakeholder groups matter for a public health program specifically: those involved in operating the program (staff, partner organizations, funders), those served or otherwise affected by it (participants, their communities), and primary intended users of the evaluation findings — a distinct group from either of the first two, since the person who has to decide whether to continue funding the program is not necessarily the person running it or the person receiving services. Naming intended users explicitly, before designing the evaluation, is what makes Step 5 (Justify Conclusions) and Step 6 (Ensure Use) achievable later rather than an afterthought.

2. Describe the program

A shared, written description of the program — its need, expected effects, activities, resources, stage of development, context, and logic model — before evaluation design starts. The point isn’t documentation for its own sake: stakeholders routinely hold different implicit models of what the program is supposed to do, and an evaluation built on an undocumented, unshared assumption tends to produce findings that different stakeholders interpret (or dispute) differently after the fact. The program’s stage of development matters directly to what follows — a program still being piloted needs a formative, process-focused evaluation; a mature, stable program can support a summative, outcome-focused one. Describing the program wrongly at this step is the single most common cause of an evaluation that answers a question nobody was actually asking.

3. Focus the evaluation design

Not every program aspect can be evaluated at once, and CDC’s framework treats this as a deliberate scoping decision rather than a resource constraint to apologize for. Focusing means settling, in writing, the evaluation’s purpose (to gain insight, change practice, assess effects, or affect participants directly — these call for different designs), its specific users and uses, the evaluation questions themselves, and the methods and agreed-on standards for the work. A design that tries to answer every stakeholder’s question with equal depth usually ends up answering none of them well; focusing is where the stakeholder list from Step 1 gets translated into a short, prioritized set of evaluation questions the remaining steps can actually answer.

4. Gather credible evidence

Credibility, in CDC’s framing, is a property of the evidence relative to the specific question and specific users from Step 3 — not an absolute methodological ranking where a randomized trial always outranks a chart review. Gathering evidence means specifying indicators, sources, quality, quantity, and logistics: what counts as evidence of success, where it comes from, how much of it is enough to support a conclusion, and how it will actually be collected within the program’s real operating constraints. This is the step where an outcomes framework like RE-AIM, or a design choice like a pre-post comparison, a regression discontinuity, or a qualitative case study, gets selected — the choice should follow from the questions focused in Step 3, not the other way around.

5. Justify conclusions

Evidence alone doesn’t produce a conclusion; it has to be linked to a set of values through explicit standards, compared against some frame of reference, and interpreted before it’s justified as a conclusion anyone should act on. CDC names four elements here: standards (the values used to judge performance — e.g., is a 20% enrollment increase good, adequate, or poor, and against what benchmark), analysis and synthesis (discovering and summarizing what the data show), interpretation (what the findings actually mean), and judgment (the conclusion itself, stated in relation to the standards). Skipping the explicit-standards step is a common failure mode: a program can report an accurate number and still leave stakeholders to argue about whether that number means the program worked, because nobody agreed in advance on what “worked” would look like.

6. Ensure use and share lessons learned

An evaluation with excellent evidence and a defensible conclusion still fails if nobody acts on it. CDC’s framework treats dissemination and use as a step to design for, not a hoped-for byproduct: it names design (planning for use from the start, not bolting it on at the end), preparation (building stakeholders’ capacity to actually use findings), feedback (an ongoing exchange with stakeholders throughout, not only at the end), follow-up (staying available to help interpret and apply findings after the report is delivered), and facilitation (removing barriers between “here is what we found” and someone actually changing what the program does). This is also where the earlier stakeholder-engagement step pays off directly — users identified and involved from Step 1 are the ones most likely to act on findings delivered in Step 6.

The four standards that run through every step

CDC’s framework adapts a widely used set of program-evaluation standards — originally developed by the Joint Committee on Standards for Educational Evaluation — organized into four groups. These aren’t a seventh step performed at the end; they’re a lens applied continuously, at every step above, to check the evaluation is being done well rather than merely being done.

  • Utility — will the evaluation serve the actual information needs of its intended users? An evaluation can be accurate and still fail this standard if it answers a question nobody with decision-making power actually had.
  • Feasibility — is the evaluation realistic, practical, diplomatically sound, and frugal given the program’s real operating context, staff time, and budget?
  • Propriety — is the evaluation conducted legally, ethically, and with due regard for the welfare of those involved and affected, including participants whose data are being used?
  • Accuracy — does the evaluation produce and convey technically adequate, valid information about the program’s merit, worth, or effectiveness?

In practice, these standards create real tension the four steps above don’t resolve on their own — the most accurate possible design (a large randomized trial, extensive primary data collection) is routinely infeasible for a modestly resourced local program, and an evaluator has to negotiate that tradeoff explicitly rather than silently defaulting to whichever standard is easiest to satisfy.

Applying the sequence: an illustrative walkthrough

Illustrative example, not a real program or a real health department. The scenario below is a composite used only to show what each step produces in practice — it does not describe any specific, identifiable public health program.

Consider a hypothetical county health department piloting a community-based fall-prevention program for adults 65 and older, run through senior centers.

Step What it produces, concretely
1. Engage stakeholders A short list naming senior-center staff, the county’s aging-services funder, and program participants as the three stakeholder groups, with the funder named explicitly as the primary intended user of the evaluation’s findings (their renewal decision is what the evaluation needs to inform).
2. Describe the program A one-page logic model: inputs (trainer hours, screening tools), activities (balance classes, home-hazard checklists), outputs (classes delivered, participants screened), outcomes (self-reported falls, confidence in balance), noted explicitly as a first-year pilot, not a mature program.
3. Focus the design Given the pilot stage from Step 2, the evaluation is scoped as formative and process-focused (is the program reaching and retaining participants as designed?), not an effectiveness trial — deliberately deferring the effectiveness question to a later year.
4. Gather evidence Attendance logs, a pre/post self-reported balance-confidence scale administered at senior centers, and staff-completed fidelity checklists for each class session.
5. Justify conclusions Attendance and confidence-scale results are compared against the standard agreed with the funder in Step 1 (60% session-retention as the pilot’s threshold for “on track”), not an arbitrary post-hoc benchmark.
6. Ensure use Findings are delivered to the funder ahead of the renewal decision, with a short follow-up session to help senior-center staff interpret the fidelity-checklist results and adjust class scheduling before year two.

Does program evaluation need IRB review?

Often not, but “it’s program evaluation” isn’t itself the determination. Under the Common Rule (45 CFR 46), what makes an activity “research” requiring IRB oversight is a specific two-part regulatory test, and institutional review offices commonly treat program evaluation, quality improvement, and accreditation self-studies as falling outside that definition on the generalizable-knowledge element specifically — the evaluation is designed, at the outset, to inform decisions about this program in this setting, not to produce findings intended to apply beyond it. That intent can shift mid-project (a pilot evaluation later written up for publication with generalizable claims, for example), which is why institutions require a documented screening determination from an authorized reviewer rather than a program manager’s own judgment call. See quality improvement vs. human subjects research for the full determination walkthrough.

How the framework relates to implementation science

The CDC framework is close kin to, but not the same discipline as, implementation science, which studies the methods that promote uptake of evidence-based practices into routine settings. A program evaluation asks whether this program worked; implementation science asks why an intervention that works in trials does or doesn’t get adopted, delivered with fidelity, and sustained in real-world settings generally. The two overlap heavily in practice — an impact evaluation of a program’s outcomes and an implementation-science study of the same program’s adoption barriers are often run side by side — but they answer different questions and are frequently confused because both use terms like “fidelity” and “sustainability.” Also distinct: research evaluation frameworks like the UK’s REF or CoARA assess the quality and impact of research outputs and researchers themselves — a different object of evaluation entirely from evaluating a public health program.

Frequently asked questions

What are the four standards of the CDC evaluation framework?

Utility, feasibility, propriety, and accuracy. They’re adapted from the Joint Committee on Standards for Educational Evaluation and are applied continuously across all six steps, not as a separate seventh step.

Is the CDC framework the same thing as a logic model?

No. A logic model is a single planning diagram (inputs, activities, outputs, outcomes, impact) typically produced while completing Step 2, Describe the Program. The CDC framework is the six-step process the logic model is embedded inside.

Does a program evaluation always need IRB approval?

Not automatically. Institutions commonly exclude program evaluation from “research” requiring IRB review on the generalizable-knowledge element of the Common Rule’s definition, but this requires a documented determination from an authorized reviewer, not a self-certification by whoever is running the evaluation.

How is this different from RE-AIM?

RE-AIM is an outcomes-measurement framework (five specific dimensions to score); the CDC framework is a six-step process for planning and conducting the evaluation as a whole. RE-AIM commonly gets used inside Step 4, Gather Credible Evidence, of the CDC framework rather than replacing it.

Who published the CDC framework and when?

CDC, in the Morbidity and Mortality Weekly Report Recommendations and Reports series, 1999 (volume 48, No. RR-11), titled “Framework for Program Evaluation in Public Health.”

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.