Written and maintained by CASRAI Editorial Board
Last updated
A run chart plots a quality measure over time against its median, and four probability-based rules — shift, trend, too many or too few runs, and the astronomical data point — tell you whether a change in the pattern is a real signal or just the ordinary noise every process produces. For an infection preventionist tracking a monthly CLABSI rate, a quality director defending a measure to a board, or a risk manager trying to tell whether a new protocol actually moved a number, that distinction is the entire point of collecting the data in the first place. This guide works through each rule with its exact counting criteria, a full worked example on a healthcare metric, and where a run chart’s limits mean the right next step is a Shewhart control chart instead.
What a Run Chart Actually Tests
A run chart is a line graph of a measure plotted in time order, with a single reference line drawn at the median of the plotted values — not the mean. The median is used specifically because it is not distorted by the one or two extreme points that quality-improvement data routinely produces (a single bad month during a staffing crisis, for example), which keeps the reference line representative of the typical process performance rather than pulled toward an outlier.
Every point on the chart falls into one of three positions relative to that median line: above it, below it, or on it. The four rules below are built entirely out of that above/below/on classification, applied in time order. None of them require a control limit, a standard deviation, or any distributional assumption — that simplicity is exactly what makes a run chart usable by a unit-level improvement team without a statistician on hand, and it is also exactly what a PDSA cycle needs for its own measurement step: the run chart is the standard way to display PDSA data over time, annotated with when each cycle started.
The Four Rules, With the Exact Criteria
These four rules come from the run-chart methodology formalized by Perla, Provost, and Murray (“The run chart: a simple analytical tool for learning from variation in healthcare processes,” BMJ Quality & Safety, 2011) and adopted by the Institute for Healthcare Improvement as its standard run-chart interpretation guidance. Each rule has a specific counting method — approximate application is where most misreads happen.
Rule 1 — Shift
A shift is six or more consecutive data points, all on the same side of the median. Points that fall exactly on the median line are not counted toward a shift and do not break one either — skip them and keep counting the next point on either side. Six was chosen because it is the smallest run length that has roughly a 1-in-64 (2⁶) probability of occurring by chance alone if the process were actually unchanged, which is small enough to treat as a genuine signal rather than routine variation.
Rule 2 — Trend
A trend is five or more consecutive data points that all increase, or all decrease. Two consecutive points of exactly equal value break a trend (there is no direction between them); a single repeated value does not automatically end a trend on its own if the surrounding points still form a consistent direction — apply judgment on ties the same way you would on an astronomical point, and if it’s ambiguous, don’t call the rule. Five points was set as the threshold for the same reason as the shift rule: a consistent five-point run in one direction is well outside what random fluctuation typically produces.
Rule 3 — Too Many or Too Few Runs
A run is one or more consecutive points on the same side of the median; the chart’s total run count is the number of times the data crosses from one side of the median to the other, plus one. This rule is the least intuitive of the four because it doesn’t use a fixed number — it compares your observed run count against a table of upper and lower limits keyed to the number of useful observations (points not sitting exactly on the median). Too few runs means the data is clumping into shifts longer than chance would produce (a signal, even if no single run reaches six points). Too many runs means the data is crossing the median unusually often — a sawtooth pattern typically caused by alternating measurement conditions (e.g., two different data sources or shifts feeding the same chart) rather than a real process change. Both directions are flagged from the same published limit table; a chart with, for example, 20 useful observations is expected to fall between roughly 6 and 15 runs, and a count outside that range in either direction is a signal worth investigating.
Rule 4 — Astronomical Data Point
An astronomical point is a single value so obviously different from everything around it that essentially any reasonable observer, without doing arithmetic, would agree it doesn’t belong with the rest of the data — a month with zero CLABSIs sitting among a run of 4–6/month, or one catastrophic spike. This is deliberately the one qualitative rule among the four: there is no fixed multiple-of-range or standard-deviation cutoff in the original methodology, because a run chart carries no calculated control limits to test against (that calculation is what a control chart adds — see below). Use it sparingly and require real consensus before calling a point astronomical; treating ordinary variation as astronomical is the most common way teams over-read noise as signal on a run chart.
How Many Data Points Before the Rules Are Meaningful
The shift, trend, and runs rules all lose reliability on short baselines. The generally accepted floor is 10–12 data points before applying shift and trend rules with any confidence, and the runs-table comparison needs enough useful observations for its limits to mean anything — a five-point chart cannot statistically support a shift call even if all five points happen to fall on one side of the median, because five-in-a-row is not improbable enough on its own. Practically: start plotting and annotating the chart with cycle markers before your first improvement test, not after, so the chart already carries a real baseline period by the time you need to interpret a change.
Worked Example: Monthly CLABSI Rate
A hospital’s infection prevention team tracks its CLABSI rate per 1,000 central-line days, plotted monthly. Over a 20-month period the rate runs roughly flat around a median of 1.8 for the first 11 months, then a central-line insertion bundle and daily necessity review go live at month 12.
- Months 1–11 (baseline): values oscillate above and below 1.8 with no run longer than 3–4 points — no shift, no trend, run count consistent with the table limit for 11 useful observations. This is the expected picture of a stable process: noise, not signal.
- Months 12–18 (post-intervention): seven consecutive months fall below the original 1.8 median. That is a shift (six or more consecutive points on one side) — a real, statistically improbable pattern change, not a lucky month. The team can reasonably attribute the shift to the bundle rather than chance.
- Months 14–18 within that same run: the rate additionally declines month over month for five straight months (1.4, 1.2, 1.0, 0.9, 0.7) — that also satisfies the trend rule independently of the shift already called. Two rules firing on overlapping data isn’t double-counting; it’s two independent lines of evidence for the same real improvement.
- Month 9, in the baseline period: a single month spikes to 4.6 against neighbors all under 2.0. The team’s clinical judgment (a documented cluster tied to a single unit’s temporary agency-staffing gap) supports calling this an astronomical point — investigated and explained, but correctly not treated as evidence the whole process shifted, since it doesn’t participate in a run.
Recalculating the median for the post-shift period (months 12–20) rather than leaving the original baseline median on the chart is standard practice once a shift is confirmed — otherwise every future point gets compared against a median that no longer represents current performance.
When a Run Chart Isn’t Enough — Moving to a Control Chart
A run chart is a screening tool, not a full statistical process control (SPC) method. It has no calculated control limits, so it can miss a real signal that doesn’t happen to trip one of the four pattern rules, and conversely it offers nothing beyond those four heuristics if a team wants a formal answer to “is this process in statistical control.” A Shewhart control chart (the general term for the XmR, p, u, c, and other chart types built for different data types) adds calculated upper and lower control limits, typically at three standard deviations from the centerline, which let you distinguish common-cause variation (the routine noise every process has) from special-cause variation (a signal) using a formal, calculated boundary rather than the run chart’s simpler pattern rules.
The practical sequence most quality-improvement programs use: start with a run chart because it requires less data and no chart-type selection to begin displaying and interpreting a new measure immediately; once a process is established and the team needs to distinguish common- from special-cause variation with statistical rigor — particularly before declaring a process “in control” to a board, a regulator, or a payer — move to the appropriate control chart type for the data (a p-chart for a proportion like a compliance rate, a u-chart for a rate like infections per line-days, an XmR chart for a continuous or low-count individual measure). A root-cause investigation triggered by a run-chart signal, in turn, typically feeds into root cause analysis to identify what actually changed.
Common Misapplication Errors
- Using the mean instead of the median. A mean shifts with every new point and is pulled by outliers; the whole point of the median reference line is that it stays stable enough for the shift/trend/runs rules to mean something consistent over time.
- Calling a shift or trend on too few points. Six consecutive points out of only seven total plotted is not the same evidence as six consecutive points out of thirty — the rules assume enough baseline data exists for the pattern to be improbable by chance.
- Forgetting to skip points on the median. A point sitting exactly on the median line doesn’t count toward, and doesn’t break, a shift or trend run — simply skip it and continue counting the next point that falls clearly above or below.
- Leaving a stale median in place after a confirmed shift. Once a shift rule fires and is investigated and accepted as real, recalculate the median for the new period; comparing new data against an outdated baseline manufactures false signals going forward.
- Treating irregular spacing as time order. The rules assume evenly spaced, sequentially collected observations (one point per week, per month, per unit of output). Mixing reporting cadences on one chart — or backfilling missing months out of order — breaks the probability basis the six- and five-point thresholds rely on.
Frequently Asked Questions
What’s the difference between a run chart and a control chart?
A run chart uses only a median line and four pattern rules to flag non-random signals; a control chart adds calculated upper and lower control limits (typically ±3 standard deviations) that formally separate common-cause from special-cause variation. A run chart needs less data and no chart-type decision to start; a control chart gives a statistically rigorous answer once the process and data type are established.
How many data points does a run chart need?
A commonly used floor is 10–12 points before the shift and trend rules carry real statistical weight, and enough useful (non-median) observations for the runs-table comparison to be meaningful. Fewer points than that can still be plotted, but pattern-rule calls made on them should be treated as provisional.
Why does a run chart use the median instead of the mean?
The median resists distortion from the extreme high or low points quality data routinely produces, keeping the reference line representative of typical performance rather than pulled toward an outlier month.
Can more than one rule fire on the same data?
Yes — a genuine, sustained improvement often satisfies both the shift and trend rules at once on overlapping points, as in the CLABSI example above. That’s two independent pieces of evidence for the same real change, not a duplicate finding.








