Written and maintained by CASRAI Editorial Board
Last updated
An observed-to-expected (O/E) ratio answers one specific question: given the patients this hospital or unit actually treated, how many adverse events would a risk-adjustment model have predicted, and how does that compare to what actually happened? For infection preventionists, patient-safety officers, quality directors, and risk managers, that distinction matters because a raw event count — eight CLABSIs, three readmissions, twelve pressure injuries — says almost nothing on its own. A tertiary referral center running a busy oncology and transplant service will accumulate more device-days, more comorbidity, and more baseline risk than a community hospital doing routine surgery, and will generate more adverse events as a simple consequence of who it treats, not necessarily how well it treats them. The O/E ratio is the mechanism built to strip that case-mix difference out of the comparison.
What “Expected” Actually Means
The “expected” value in an O/E ratio is not a target, a national average applied uniformly, or a benchmark a hospital is being asked to hit. It is a model-predicted count specific to that hospital’s own patients, produced through a process statisticians call indirect standardization:
- A risk-adjustment model is built (typically by a national surveillance or quality program, using a large reference population) that relates known risk factors — device type, procedure category, patient acuity, facility characteristics, and similar variables — to the rate at which an adverse event occurs in that reference population.
- The hospital’s own patients — their actual device-days, actual procedure mix, actual risk-factor distribution — are then run through that model’s coefficients.
- The output is the expected (or “predicted”) count: how many events the reference model says should occur, if this specific hospital’s case mix behaved exactly like the reference population’s did.
That is why two hospitals with identical observed event counts can have very different expected counts, and therefore very different O/E ratios — the ratio is scored against each hospital’s own risk profile, not against a shared flat number. A hospital that took on a sicker, higher-risk population and still produced the model-predicted number of events is performing as expected, even if its raw count of events is higher than a lower-acuity hospital down the road.
How the Ratio Is Actually Calculated
Mechanically, the calculation itself is simple — the work is almost entirely in getting the expected value right:
- Define the population and time period. This is usually a defined set of device-days, procedure-days, or patient-days over a reporting period (a quarter, a year), scoped to a specific unit, procedure category, or facility.
- Score that population through the risk-adjustment model to produce the expected (predicted) count of events, using the hospital’s actual mix of risk factors as inputs.
- Count the observed events that actually occurred in that same population and period, using the surveillance definition’s case-finding rules.
- Divide: O/E = observed ÷ expected.
An O/E of 1.0 means the hospital produced exactly as many events as its own risk-adjusted model predicted. Above 1.0 means more events occurred than predicted; below 1.0 means fewer. National Healthcare Safety Network (NHSN) surveillance — the source of the closely related standardized infection ratio (SIR) used for CLABSI, CAUTI, and other healthcare-associated infections — calculates the expected count using negative binomial regression against its national baseline model, and, notably, will not calculate or report a ratio at all when the predicted count falls below a set minimum (CMS’s own reporting requirements set a 1.000-predicted-infections floor for the NHSN measures used in hospital pay-for-reporting programs). That floor exists for a reason covered below: an O/E ratio built on a very small expected count is statistically unstable, no matter how the arithmetic comes out.
A Worked Example
The numbers below are an illustrative, hypothetical worked example built to show the mechanics clearly — they are not data from any real facility.
Say a hospital’s central line-associated bloodstream infection (CLABSI) surveillance covers one ICU for a calendar quarter. The risk-adjustment model, scored against that ICU’s actual central-line-days and the unit’s location-type coefficients in the reference model, predicts an expected count of 5.2 infections. Actual surveillance for the quarter identifies 8 infections meeting the NHSN case definition.
O/E = 8 ÷ 5.2 = 1.54.
Read correctly, that does not mean “54% worse than the national average.” It means this specific ICU, given its own device-days and risk profile, produced 54% more infections than the reference model predicted for a unit exactly like this one. That is a materially different and more useful statement — it isolates a genuine deviation from the unit’s own expected performance rather than penalizing the unit for having sicker patients or more device-days than a lower-acuity comparator.
Now compare a second, smaller step-down unit in the same illustrative hospital. Its lower device-day volume produces an expected count of only 1.1 infections for the quarter; two infections actually occurred. O/E = 2 ÷ 1.1 ≈ 1.82 — numerically a larger ratio than the ICU’s 1.54. But an expected count of 1.1 is built on a very thin denominator: one or two additional or fewer infections would swing the ratio dramatically, purely from chance. This is precisely the situation the confidence-interval discussion below, and NHSN’s predicted-count reporting floor, exist to guard against — a small expected count can produce a dramatic-looking O/E ratio that is really just statistical noise.
Why O/E Is Not Comparable Across Differently Risk-Adjusted Models
The single most common misreading of an O/E-style ratio is treating the number itself as portable — comparing a 1.5 from one measure or one time period directly against a 1.5 from another, as though “1.5” means the same thing everywhere. It usually does not, for at least three reasons:
- Different measures use different models entirely. NHSN’s SIR is built on its own national baseline model for a specific device/procedure type. CMS’s Hospital Readmissions Reduction Program (HRRP) calculates its excess readmission ratio from a completely separate hierarchical logistic regression fit to Medicare claims data, and by statutory design floors that ratio at 1.0 — a hospital that outperforms its expected readmission rate gets no credit below 1.0 in the payment calculation, an asymmetry that has nothing to do with the O/E arithmetic itself and everything to do with a policy choice layered on top of it. AHRQ’s Patient Safety Indicators use a third, claims-based comorbidity adjustment. A 1.5 on one of these is not the same statistical object as a 1.5 on another, even though all three are “observed over expected” at their core.
- Reference models get periodically rebased. National surveillance programs update their baseline models to a new reference period as practice patterns and case mix shift nationally. When that happens, the “expected” value for the exact same unit, with the exact same patients, can change from one baseline era to the next — not because the unit’s performance changed, but because the yardstick did. An O/E ratio calculated under an older baseline is not directly comparable to one calculated under a newer one without checking which baseline period each was scored against.
- Denominator and scope differences hide inside the ratio. Two hospitals reporting O/E ratios for “the same” measure can still differ in unit-type mapping, case-finding practices, or reporting period length in ways that shift the expected count without showing up anywhere in the headline number.
The practical rule: an O/E ratio is only meaningful as a comparison against 1.0 (this hospital’s own expected performance) or as a trend for the same measure, same model version, same unit, over time — not as a league-table number to rank hospitals against each other unless every one of those elements is held constant.
Reading a Confidence Interval That Crosses 1.0
Most credible O/E-style reporting attaches a confidence interval to the ratio, and reading that interval correctly is at least as important as reading the point estimate.
The observed count that drives the numerator is a count of discrete events, and counts of rare events carry real statistical variability — the same underlying risk could plausibly have produced one fewer or one more event by chance alone, especially at low volumes. The confidence interval expresses that uncertainty as a range: if the interval spans, say, 0.9 to 2.3, the data cannot statistically distinguish this hospital’s true underlying performance from “exactly as expected” (an O/E of 1.0 sits inside that range), even though the point estimate of 1.54 looks elevated on its face.
Two things follow from that, both easy to get backwards:
- A CI that crosses 1.0 is not proof there is no problem — it means the current data volume cannot confirm one statistically. A genuine, real elevation in risk can still be sitting inside a wide interval simply because too few events have accumulated yet to narrow it. This is exactly why NHSN declines to calculate a SIR at all below its predicted-count floor: at very low expected counts, the interval is so wide that the point estimate alone is close to meaningless, so the program does not report a number that would invite over-interpretation.
- Conversely, a narrow interval entirely above 1.0 on a high-volume unit is a much stronger statistical signal than the same point estimate on a low-volume unit with a wide interval — the width of the interval, not just the ratio itself, is what tells a quality director how much confidence the number deserves.
The takeaway for a quality committee reviewing these numbers: read the interval before reacting to the point estimate, and treat “1.0 is inside the interval” as “not yet statistically distinguishable from expected” rather than “confirmed fine.”
Where O/E-Style Ratios Show Up in Hospital Quality Measurement
The observed-to-expected mechanic is not unique to one measure — it is the underlying logic behind several distinct programs a quality or patient-safety office tracks:
- The standardized infection ratio (SIR) used across NHSN’s healthcare-associated infection measures (CLABSI, CAUTI, SSI, MRSA bacteremia, and C. difficile).
- The excess readmission ratio in CMS’s Hospital Readmissions Reduction Program, which risk-adjusts and floors at 1.0 by statute.
- The risk-adjusted components feeding the HAC Reduction Program‘s composite score.
- The individual, claims-adjusted indicators explained in AHRQ’s Patient Safety Indicators, several of which are themselves O/E-style ratios before being combined into the PSI-90 composite.
Worth a specific note of disambiguation: “risk adjustment” also appears in an entirely different context — Medicare Advantage payment, where HCC risk-adjustment coding uses documented diagnoses to set a capitated payment amount for an individual patient. That is a payment-methodology use of “risk adjustment” aimed at pricing an individual’s expected cost of care. The O/E ratio covered here is a quality-measurement use of the same general statistical family, aimed at comparing a population’s observed outcomes to a model-predicted count — related mathematics, different purpose, and not interchangeable in conversation with finance or coding teams.
Common Misreadings to Avoid
- Treating O/E as a raw incidence rate. It is a ratio against a model prediction, not a rate per device-day or per admission — a hospital with a low raw infection rate can still have an elevated O/E if its risk profile predicted an even lower rate.
- Comparing ratios calculated under different baseline eras without checking which reference period each was scored against.
- Publicizing a striking O/E ratio built on a very small expected count without disclosing how thin that denominator is — a 2.0 built on an expected count of 0.6 is not the same claim as a 2.0 built on an expected count of 40.
- Reading “the interval crosses 1.0” as a clean bill of health rather than as “not yet statistically resolved” — particularly relevant when volume is climbing toward, but has not yet crossed, a reporting threshold.
Frequently Asked Questions
What counts as a “good” O/E ratio?
An O/E ratio at or below 1.0, with a confidence interval that does not extend meaningfully above 1.0, indicates performance at or better than what the risk-adjustment model predicted for that unit’s own case mix. There is no universal numeric cutoff for “good” beyond that — the interpretation depends on the specific program’s own reporting thresholds.
Is an O/E ratio of exactly 1.0 good or bad?
Neither, on its own — it means the hospital produced exactly the number of events the model predicted for its actual patient population. It is the reference point the ratio is built around, not a pass/fail line by itself.
Can O/E ratios be compared directly between two different hospitals?
Only if both are scored under the same risk-adjustment model, the same baseline period, and the same measure definition. Otherwise the “expected” side of each ratio reflects a different reference model, and the two numbers are not measuring the same thing even though they look like comparable decimals.
Why does a smaller unit end up with a wider confidence interval on the same point estimate?
Because the interval’s width is driven by how many events the calculation is built on, not by the ratio’s value. A small expected count means a small observed-event denominator, and small counts of rare events carry proportionally more statistical noise — the same point estimate on a high-volume unit rests on far more data and produces a correspondingly narrower interval.
What is the difference between an O/E ratio and a standardized infection ratio (SIR)?
The SIR is a specific, named application of the O/E mechanic to NHSN’s healthcare-associated infection surveillance, using NHSN’s own national baseline models to generate the expected count. “O/E ratio” is the general statistical concept; “SIR” is one program’s specific implementation of it.








