Skip to main content
v2026.11,610 entries · CC-BY 4.0

Root Cause Analysis in Healthcare: Team, Timeline, and Causal Factors

A hospital root cause analysis is only as good as the process behind it. This guide covers RCA team composition, timeline reconstruction from records and interviews, and how to identify and write a defensible causal statement — the fundamentals underneath any specific action-hierarchy framework.

Ask about Root Cause Analysis in Healthcare: Team, Timeline, and Causal Factors

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

A root cause analysis (RCA) is only as good as the process that produced it — before a hospital gets to any corrective action, it has to build the right team, reconstruct an accurate timeline, and separate what actually caused the event from what merely happened to be present. Most of the RCA literature and training material hospitals encounter jumps straight to techniques (five whys, fishbone diagrams) or to the follow-on question of what to do with the findings. This guide covers the ground underneath that: who should be in the room, how to build a defensible timeline from record review and interviews, and how to write a causal statement that survives scrutiny.

Written for hospital patient-safety officers, quality directors, risk managers, and infection preventionists who convene or facilitate RCA teams. If you already know this part and want the framework for turning findings into corrective actions that actually hold, see our companion guide on the RCA2 action hierarchy.

What Triggers a Root Cause Analysis

Hospitals most commonly convene a formal RCA after a sentinel event — a patient safety event that reaches a patient and results in death, permanent harm, or severe temporary harm, and is therefore not primarily related to the natural course of the patient’s illness. The Joint Commission’s Sentinel Event Policy is reported to call for a comprehensive systematic analysis and corrective action plan within 45 business days of the event or its discovery (confirm the current figure against the Sentinel Event chapter of your organization’s applicable accreditation manual, since the policy is revised periodically). Many hospitals also run a full RCA, or a scaled-down version, for serious near misses and other high-severity events that didn’t reach the patient — waiting for actual harm before investigating a process failure discards the free information a near miss provides.

Not every reportable event warrants a full RCA. Lower-severity, higher-frequency events are often better served by an apparent cause analysis (a shorter, single-facilitator review) or by aggregating several related low-harm events and looking for a shared cause across them, rather than running a full multidisciplinary RCA on each one individually. The triage decision — full RCA vs. a lighter-weight review — is usually made by the patient safety or risk management office at intake, based on actual or potential severity, not on how much attention the event has already drawn.

Building the RCA Team

The single most consistent piece of guidance across RCA methodology — reflected in both AHRQ’s patient safety literature and the VA National Center for Patient Safety’s own RCA guidance — is that the analysis has to be done by a multidisciplinary team, not by one investigator working alone or by the department where the event occurred reviewing itself. A team assembled correctly includes:

  • Frontline staff who are close to the process, but not necessarily the specific individuals involved in the event. The VA’s own patient-safety guidance frames this as needing people who are “the most familiar with the situation” — that’s a process-familiarity requirement, not a requirement to interview or seat the involved staff as team members. Direct participants are almost always interviewed as part of evidence-gathering; whether they sit on the analysis team itself is a judgment call that depends on your organization’s just-culture posture and on avoiding both hindsight bias and a chilling effect on future reporting.
  • A facilitator trained in RCA methodology who has no personal stake in the outcome and can keep the team from converging on the first plausible explanation or on an individual to blame.
  • Subject-matter expertise from every department or role the event actually touched — nursing, pharmacy, a specific service line, biomedical engineering for a device-related event, environmental services for an environmental-controls failure. A team that’s all nursing when the event involved a medication-dispensing system will miss the pharmacy-side and engineering-side contributing factors entirely.
  • Leadership participation, both to give the process organizational credibility and to make sure a finding that requires resources or authority outside the team’s own scope has a route to actually get acted on.

What the team should generally avoid: anyone whose direct reporting line creates a conflict of interest in the findings, and — per a well-established principle in the patient-safety literature — a team composition or process that treats the RCA as an exercise in assigning individual blame rather than understanding the system that allowed the event to occur. That distinction between systems thinking and blame is also the organizing idea behind a formal just culture framework, and teams that have adopted one tend to run cleaner RCAs because the blame question has already been separated out of the analysis.

Reconstructing the Timeline

Before a team can identify a cause, it needs an accurate, agreed-upon sequence of what actually happened — not what the policy says should have happened. AHRQ’s patient-safety guidance describes the starting point as data collection and reconstruction of the event through record review and participant interviews. In practice that means pulling every contemporaneous record that touches the event (the medical record, medication administration record, device logs, staffing schedules, any relevant alarm or monitoring data) and interviewing everyone who was present or involved, as close to the event as practical while memory is still fresh.

A few practices make the resulting timeline more defensible:

  • Build the timeline before assigning any interpretation to it. The first pass should be a neutral, time-stamped sequence of what occurred — this happened, then this happened — with interpretation (why something happened, whether it should have happened) deliberately held for a separate step. Collapsing sequencing and interpretation into one pass is one of the most common ways teams anchor on an early, incomplete explanation.
  • Interview separately before comparing accounts. Interviewing participants individually, before they’ve had a chance to align their recollections with each other, surfaces genuine discrepancies in what different people believed was happening at the time — discrepancies that are themselves often a contributing factor (a misunderstanding about who was responsible for a step, or about what a handoff communicated).
  • Distinguish what a record shows from what a person recalls, and note where the two conflict rather than silently picking one. A conflict between the documented time of a step and a participant’s recollection of when it happened is data, not noise.
  • Capture system state, not just actions — staffing levels at the time, whether a device was functioning normally, what else was competing for the same staff member’s attention. A timeline of actions alone tends to reproduce an individual-blame framing; a timeline that also captures conditions is what makes a systems-level cause visible.

From Timeline to Cause: Active Errors, Latent Conditions, and Causal Factors

Once the sequence is established, the team’s job shifts to identifying not just how the event occurred but why. Patient-safety literature (reflected in AHRQ’s own RCA guidance) frames this as separating active errors — the specific action or omission at the point of care that immediately preceded the event — from latent conditions: the underlying system weaknesses (staffing, design, training, process, culture) that made an active error more likely, or made it more likely to reach the patient once it occurred. This is the same logic behind the “Swiss cheese model” of accident causation: a single active error rarely causes harm on its own; harm occurs when it passes through multiple layers of an organization’s defenses that all happened to have a gap in the same place at the same time.

A root cause, in this framework, is a latent condition — a system-level factor that, if corrected, would prevent recurrence not just of this specific event but of the broader failure mode. A properly identified causal factor generally has a few characteristics:

  • It states a clear cause-and-effect relationship rather than simply describing what happened. “The infusion pump’s dose-limit alert was disabled” is an observation; “the dose-limit alert was disabled because the unit’s standard configuration didn’t enforce it, which allowed an order outside the safe range to be administered without a system check” is a causal statement.
  • It uses specific, neutral language rather than vague or blame-laden descriptors. “Staff was careless” is neither specific nor actionable; “the medication label for the two look-alike vials was not distinguishable at the point of selection” identifies something a corrective action can actually target.
  • A human action described as a cause needs a preceding system cause behind it. “The nurse administered the wrong dose” is rarely a complete causal statement on its own — the team’s job is to keep asking why that action was possible or likely given the system the person was working in, until it reaches a factor that isn’t itself just another individual’s action.
  • A procedure violation is not, by itself, a root cause. If staff routinely deviate from a written procedure, the root cause is usually why the deviation is routine (an unworkable procedure, a workaround the organization tacitly accepted, competing priorities that make compliance impractical) rather than the deviation itself.

Teams commonly use a structured technique to get from the timeline to these causal factors — repeatedly asking why a condition existed until the answer reaches a system-level factor, or a fishbone/cause-and-effect diagram to organize contributing factors by category (people, process, equipment, environment). Either is a facilitation tool for organizing the team’s thinking, not a substitute for the discipline above: a “why” chain that stops at a human action, or a fishbone diagram populated with vague entries, produces the same weak result the structured technique was supposed to prevent.

A Worked Example

A patient received a medication dose ten times higher than ordered because a decimal point was misread during manual transcription from a paper order to the electronic system. The active error is the transcription itself. Working backward through latent conditions, the team’s timeline and interviews establish: the unit was short-staffed on the shift in question; the transcribing nurse was covering an additional patient assignment outside their usual unit; the electronic system had no independent dose-range check that would have flagged the resulting order as implausible; and a verbal double-check step existed in policy but was not consistently performed during high-volume shifts. None of those latent conditions alone caused the event — the causal factors are the absence of a system-level dose check and a verification step that depended on staff finding time for it during exactly the conditions (high volume, short staffing) when it was least likely to happen. “The nurse misread a decimal point” explains what happened; it is not the root cause, and an action plan that stops at retraining that one nurse leaves every other latent condition in place for the next short-staffed, high-volume shift. This is deliberately a composite, illustrative scenario, not a report of an actual event at any specific institution.

Documenting Findings for Sign-Off

A completed RCA record generally needs to show, in a form a reviewer or surveyor can follow without having sat in the room: the event description and timeline as reconstructed, the evidence base (records reviewed, interviews conducted), the causal factors identified with the reasoning connecting each one to the timeline, and the resulting action plan. Teams that build the timeline and causal statements carefully during the analysis produce this documentation almost as a byproduct; teams that skip straight to a discussion of “what should we do differently” without first agreeing on a defensible timeline and causal chain tend to produce documentation that reads as a conclusion in search of a justification. Once the causal factors are in hand, grading each proposed action’s strength — rather than defaulting to training and policy memos — is exactly what our RCA2 action hierarchy guide covers in detail; and once an action is in place, treating “completed” and “effective” as two separate questions, the same discipline behind a PDSA cycle, is what confirms the underlying risk actually went down.

Frequently Asked Questions

What is a root cause analysis in a hospital setting?

A root cause analysis (RCA) is a structured, multidisciplinary process for investigating a patient safety event by reconstructing an accurate timeline of what occurred and identifying the underlying system-level factors — not just the immediate human action — that allowed the event to happen, so that corrective action addresses the actual cause rather than a single individual’s action.

Who should be on a hospital RCA team?

A multidisciplinary, interdisciplinary team including frontline staff familiar with the process (not necessarily the specific individuals directly involved in the event), a trained facilitator with no stake in the outcome, subject-matter expertise from every department the event touched, and leadership participation for organizational credibility and follow-through authority.

What is the difference between an active error and a latent condition?

An active error is the specific action or omission at the point of care that immediately preceded the event. A latent condition is an underlying system weakness — staffing, design, training, process, or culture — that made the active error more likely to occur or more likely to reach the patient. A root cause is typically a latent condition, not the active error itself.

How long does a hospital have to complete an RCA after a sentinel event?

Under The Joint Commission’s Sentinel Event Policy, a comprehensive systematic analysis and corrective action plan is reported to be due within 45 business days of the event or its discovery — confirm the current figure against your accreditation manual’s Sentinel Event chapter, since the policy is revised periodically.

Does every patient safety event require a full root cause analysis?

No. Sentinel events and other high-severity events generally warrant a full RCA; lower-severity or higher-frequency events are often better handled through a shorter apparent cause analysis or by aggregating similar low-harm events to look for a shared cause, reserving the full multidisciplinary RCA process for the events where its cost is justified.

What is the “Swiss cheese model” and how does it relate to root cause analysis?

It’s a widely used conceptual model of accident causation in which an organization’s safety defenses are pictured as slices of Swiss cheese: harm occurs not because of one single failure, but because the holes (weaknesses) in multiple layers of defense happen to line up at the same time. It underlies why RCA looks for the several latent conditions that combined to let an active error reach the patient, rather than stopping at the first error found.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.