Skip to main content
v2026.11,610 entries · CC-BY 4.0

Healthcare Failure Mode and Effect Analysis (HFMEA): Process Mapping, Hazard Scoring, and the Decision Tree

HFMEA is the proactive counterpart to root cause analysis: map a high-risk process, score every failure mode on the VA’s Hazard Scoring Matrix, and use the decision tree to decide which ones actually need a corrective action before a patient is harmed.

Ask about Healthcare Failure Mode and Effect Analysis (HFMEA): Process Mapping, Hazard Scoring, and the Decision Tree

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

A hospital root cause analysis starts after something has already gone wrong. Healthcare Failure Mode and Effect Analysis (HFMEA) is the tool built to run before that — a structured way to walk through a process step by step, find the points where it is most likely to fail, and fix them while the failure is still hypothetical.

HFMEA was developed by the Veterans Health Administration’s National Center for Patient Safety (NCPS), which describes it as combining elements of traditional industrial Failure Mode and Effect Analysis (FMEA) with root cause analysis and other risk-assessment methods, using both quantitative risk scoring and a qualitative decision tree to identify which hazards actually need to be addressed. NCPS still maintains the methodology today, most recently in its Guidebook for Proactive Risk Assessment (November 2025), alongside the original HFMEA process flipbook and worksheet.

This guide is written for hospital patient-safety officers, quality directors, risk managers, and infection preventionists who are asked to run — or sit on the team for — a proactive risk assessment of a high-risk process. It walks through HFMEA’s five steps, the Hazard Scoring Matrix, and the decision tree that separates a failure mode worth acting on from one worth documenting and monitoring, using a medication-administration process as a worked illustration throughout.

HFMEA as the Proactive Counterpart to RCA

Root cause analysis and HFMEA answer two different questions about the same underlying goal — a process that doesn’t hurt patients.

  • RCA is reactive. It starts from a specific event that already happened — most often a sentinel event — and works backward through the timeline to find what actually caused it. See our companion guide on root cause analysis in healthcare for team composition and timeline reconstruction, and on the RCA2 action hierarchy for turning findings into corrective actions that hold.
  • HFMEA is proactive. It starts from a process that hasn’t failed (yet), maps every step, and asks which steps are most likely to fail and how bad it would be if they did — before any patient is harmed.

The two are complementary, not competing: an HFMEA on a high-risk process routinely turns up the same failure modes an RCA team would otherwise have to reconstruct after the fact. Many hospitals run an HFMEA specifically on a process where a near miss, a related RCA finding elsewhere in the organization, or an incoming regulatory or equipment change has raised the process’s risk profile.

The Five Steps of HFMEA

NCPS structures HFMEA into five sequential steps. The first three are preparation; the real analysis happens in step four.

Step 1: Define the Topic

Name the specific process to be studied, and scope it deliberately — “medication administration on a medical-surgical unit” is analyzable; “medication safety” is not. A topic that’s too broad produces a process map with dozens of steps and a hazard analysis no team can finish; a topic scoped too narrowly misses the handoffs where risk actually concentrates.

Step 2: Assemble the Team

HFMEA is a multidisciplinary exercise by design. NCPS guidance calls for a team that includes the people who actually perform each step of the process, a facilitator experienced in the method, and — deliberately — at least one person unfamiliar with the process, whose outside perspective tends to catch workaround steps the regular staff no longer notice because they’ve become routine.

Step 3: Graphically Describe the Process

Before any hazard can be scored, the team maps the process as a flow diagram, breaking it into its major subprocesses in the order they actually happen — not the order the policy manual describes them. For a medication-administration process, that typically separates into subprocesses such as prescribing, transcribing/order verification, dispensing, and administration at the bedside, each of which gets broken down further into its individual steps in step 4. Process mapping done well here is what makes the hazard analysis exhaustive rather than a brainstorm limited to whatever failure modes the team happens to remember.

Step 4: Conduct a Hazard Analysis

This is the core of HFMEA, and it runs in three parts for every step identified in the process map:

  1. List potential failure modes. For each process step, the team brainstorms every plausible way that step could fail — not just the ones that have happened before. A single step in a medication-administration process (e.g., “verify patient identity before administration”) might carry several distinct failure modes: identity not checked at all, checked against the wrong source, or checked using a workaround (room number instead of two identifiers) that fails when a patient has been moved.
  2. Score each failure mode on the Hazard Scoring Matrix (severity × probability — see below).
  3. Run every failure mode that scores high enough through the Decision Tree to determine whether it proceeds to action planning or is documented and set aside.

Step 5: Identify Actions and Outcome Measures

For every failure mode the decision tree routes to action, the team decides whether to eliminate the failure mode, control it (reduce its probability or improve detectability), or — for a low-severity, well-controlled failure mode a team consciously chooses not to act on — accept it, with the rationale documented. Each planned action gets an owner, a target date, and an outcome measure: a way to confirm afterward that the redesigned step actually reduced the hazard rather than just moving it somewhere else in the process.

The Hazard Scoring Matrix

Every failure mode identified in step 4 gets scored on two independent dimensions, following NCPS’s Hazard Scoring Matrix:

  • Severity — how bad the outcome would be if the failure mode occurred and reached a patient, typically rated across four bands from minor/no harm through catastrophic (death or major permanent loss of function).
  • Probability — how likely the failure mode is to occur, typically rated across four bands from remote through frequent.

The two ratings combine into a single hazard score for that failure mode. A failure mode that is both severe and likely scores at the top of the matrix and proceeds automatically to the decision tree; one that is minor and remote is documented and generally set aside without further analysis. The exact severity and probability category definitions your team should use are in NCPS’s published matrix and worksheet — reproduce them for your team’s calibration session rather than improvising wording, since a hazard score is only comparable across failure modes if every team member is rating against the same anchors.

One category overrides the numeric score entirely: a single point weakness — a step where, if it fails, nothing downstream catches the failure before it reaches the patient — gets routed to the decision tree regardless of how low its probability score comes out, because a rare failure with no backstop is exactly the pattern that produces a sentinel event.

Hazard score band What it typically means What happens next
Low severity × low probability Minor consequence, rarely occurs Document; no further analysis required
Mixed (one dimension high, one low) Either infrequent-but-serious or frequent-but-minor Team judgment on whether to route to the decision tree
High severity × high probability, or any single point weakness Likely, and/or catastrophic if it occurs Always proceeds to the decision tree

The Decision Tree: Does This Failure Mode Need Action?

A high hazard score doesn’t automatically mean “build a corrective action.” HFMEA’s decision tree is what turns a scored failure mode into an actual go/no-go decision, by asking three questions in sequence for each failure mode that reached this stage:

  1. Is this a single point weakness? If the step fails with no downstream safeguard to catch it, the team generally proceeds straight to action planning regardless of the other two questions.
  2. Is there an existing, effective control measure for this failure mode? A control measure that’s documented but not actually followed in practice doesn’t count — this is where the team’s own front-line members matter, since they know which controls are real and which exist only in the policy manual.
  3. Is the failure mode easily detectable if it occurs, before it reaches the patient? A failure that’s likely to be caught by a downstream check (a pharmacist verification, a second-nurse check) carries lower net risk than an identical failure mode with no realistic chance of being caught.

Only a failure mode that is not a single point weakness, does have an effective existing control, and is readily detectable can reasonably be set aside with documentation rather than action. Any other combination routes to step 5 — identify and assign a corrective action.

A Worked Illustration: Medication Administration

The following is an illustrative walkthrough of how the method applies to one subprocess, not a real hospital’s HFMEA — use it to see the mechanics, not as a template to copy without your own team’s process map.

Subprocess: bedside medication administration. One process step: “nurse scans patient wristband before administering.” A plausible failure mode: the wristband is unreadable (smudged, wrong side of the bed rail) and the nurse manually enters the medical record number from memory of the chart instead of rescanning or fetching a new wristband.

  • Severity: potentially catastrophic — a manual entry error can result in a medication reaching the wrong patient.
  • Probability: occasional — wristband degradation is a known, recurring issue on units with limited relabeling supplies.
  • Single point weakness check: yes, if no second identifier check exists downstream of the scan step.
  • Decision tree: single point weakness alone routes this to action planning, independent of the other two questions.
  • Candidate action (step 5): hard-stop the administration workflow when a scan fails, rather than allowing manual entry, paired with a unit-level wristband-quality outcome measure tracked alongside the existing medication use evaluation program.

Where HFMEA Fits Against Related Tools

Patient-safety and quality teams work with several risk-assessment frameworks that sound similar but answer different questions. Getting the right one in front of a surveyor — or a new team member — matters:

Tool When it runs What it produces
HFMEA Proactively, before an incident, on a chosen high-risk process A scored, prioritized list of failure modes with corrective actions
Root cause analysis Reactively, after a sentinel event or serious near miss A causal statement and, via RCA2’s action hierarchy, ranked corrective actions
Hazard vulnerability analysis Annually, for emergency-preparedness planning A prioritized list of external/internal hazards (storms, power loss, mass casualty) for the emergency operations plan — not a process-level tool
ISO 14971 During medical device design and post-market surveillance A device manufacturer’s risk management file — a regulatory deliverable, not a hospital operations tool

HFMEA and hazard vulnerability analysis are especially easy to conflate because both produce a scored hazard list — but HVA scores external and facility-wide threats to continuity of operations, while HFMEA scores failure points inside one specific clinical or administrative process. A hospital typically needs both, run by different teams, for different purposes.

Why Hospitals Run HFMEA: The Accreditation Angle

Joint Commission accreditation standards are widely reported to require accredited hospitals to proactively assess at least one high-risk process on a recurring cycle, and FMEA/HFMEA is the tool most commonly cited in accreditation guidance as satisfying that requirement — verify the exact current expectation and cadence against your own accreditor’s standards manual before citing a specific interval in a survey response, since accreditation requirements are revised periodically. In practice, this is also where an HFMEA typically lives administratively: as a documented Performance Improvement Project inside a hospital’s QAPI plan and PIP write-up, with the hazard analysis and resulting actions as its evidence of the “systematic, data-guided” quality-improvement activity QAPI requires.

Quality directors preparing for CPHQ certification and patient-safety officers pursuing CPPS certification will both encounter HFMEA as core exam content — it sits squarely inside the proactive-risk-assessment competency both credentials test.

Frequently Asked Questions

What does HFMEA stand for?

Healthcare Failure Mode and Effect Analysis — a proactive risk-assessment method developed by the VA National Center for Patient Safety, adapted from industrial FMEA for clinical and administrative healthcare processes.

How is HFMEA different from RCA?

RCA is reactive — it starts after a specific event and works backward to find the cause. HFMEA is proactive — it starts from a process that hasn’t yet failed and works forward to find and fix its weakest points before they cause harm. See our root cause analysis guide for the reactive side of this pair.

What is a single point weakness in HFMEA?

A process step where, if that one step fails, nothing downstream catches the failure before it reaches the patient. HFMEA’s decision tree routes any single point weakness to action planning regardless of how the failure mode scored on the Hazard Scoring Matrix, because a rare failure with no backstop is a recurring pattern behind sentinel events.

What is the VA Hazard Scoring Matrix?

A two-dimensional scoring tool that rates each identified failure mode on severity (how bad the outcome would be) and probability (how likely the failure is to occur), following the categories in NCPS’s published matrix. The combined score determines whether a failure mode automatically proceeds to the decision tree.

Can HFMEA replace RCA?

No — they cover different situations. A hospital still needs RCA (or, for a lower-harm cluster of events, an apparent cause analysis) to investigate events that have already happened. HFMEA reduces how often that becomes necessary by fixing high-risk processes proactively, but it doesn’t replace the reactive investigation a sentinel event still requires.

How long does a full HFMEA take?

It varies with process complexity and team availability rather than following a fixed timeline — a narrowly scoped process (a single medication-administration subprocess, for example) with a committed team can complete process mapping through action planning in a small number of structured working sessions; a broader process with many subprocesses and stakeholders takes proportionally longer. Scoping the topic tightly in step 1 is the single biggest lever a team has over how long the analysis takes.

Related Reading

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.