Written and maintained by CASRAI Editorial Board
Last updated
Three carbapenem-resistant Klebsiella isolates land on one ward inside a fortnight. Before anyone can ask where, someone has to draw the curve. The epidemic curve is the first analytic artefact of an outbreak investigation and the one that decides what the rest of the investigation asks about: it converts a list of cases into a shape, and the shape tells you whether you are looking at one contaminated event, a persistent reservoir, or staff-mediated spread — and roughly when the exposure happened.
This page is about the construction decisions, not the concept. Choosing the interval. Deciding which date to plot. Reading the pattern. And the step most generic epidemiology explainers leave out: using the finished curve to bound the exposure window, so your interviews and chart reviews target a specific range of days instead of the whole month.
It is written for a hospital infection preventionist working a unit-level cluster. It sits inside CASRAI’s patient safety cluster.
What an epidemic curve is, precisely
CDC’s Principles of Epidemiology in Public Health Practice (3rd edition) defines it exactly: “An epidemic curve is a histogram that displays the number of cases of disease during an outbreak or epidemic by times of onset. The y-axis represents the number of cases; the x-axis represents date and/or time of onset of illness.”
Two things in that definition are load-bearing and routinely violated in hospital practice:
- It is a histogram, not a line chart or a run chart. The bars adjoin because time is continuous; the area of each bar is proportional to the case count in that interval. Plotting cases as a connected line implies interpolation between time points that does not exist, and it destroys the shape information you are drawing the graph to get.
- The x-axis is time of onset. Not specimen collection date, not result date, not admission date. Those are different graphs that answer different questions — see below, because in a hospital they are usually all you have.
An epidemic curve is descriptive epidemiology. It generates hypotheses; it does not test them. Testing comes afterwards, with a case-control study or a retrospective cohort — see cohort versus case-control for which design fits a closed unit population versus an open one.
Step 1: freeze the case definition before you plot anything
Every bar on the curve is a decision about who counts. Change the case definition halfway through and the curve changes shape for reasons that have nothing to do with transmission.
CDC’s rule: a case definition is clinical criteria plus restrictions by time, place and person, applied consistently to everyone under investigation — and, critically, “the case definition must not include the exposure or risk factor you are interested in evaluating.” If you suspect the west wing, do not define a case as illness among people who worked in the west wing. Define it as illness among people who worked in the facility, then analyse by wing. Building the exposure into the definition guarantees a positive finding and tells you nothing.
In hospital outbreak work the definition usually has to be tiered, because the microbiology arrives at different levels of certainty:
- Confirmed — organism identified from a sterile or otherwise definitive site, matching the outbreak phenotype or genotype.
- Probable — compatible clinical syndrome plus epidemiological link (same unit, overlapping stay) without confirmatory isolate.
- Possible / under investigation — pending culture, or a compatible syndrome with no link yet established.
Plot all three on the same curve, shaded differently. CDC’s own materials do exactly this: their Figure 4.7c “shades the individual boxes in each time period to denote which cases have been confirmed with culture results.” Shading lets the outbreak team see the confirmed skeleton and the provisional flesh on one graph, and it stops a curve from silently changing shape as pending cultures resolve.
If your cluster is device-associated, the surveillance definitions do a lot of this work already — see CLABSI, CAUTI and VAE. Note the distinction though: an NHSN surveillance definition is built for standardised inter-facility comparison, not for outbreak case-finding. An outbreak case definition is usually broader at first (to find cases) and tightened later (to analyse them). Do not let the NHSN definition silently become your outbreak definition by default; if it does, say so explicitly in the write-up. A third source sits behind both: where your jurisdiction receives your emergency department feed, syndromic surveillance data can suggest a cluster days before either definition applies — but it is keyed to visit date rather than onset or specimen date, so keep it on a separate axis instead of merging it into the same curve.
Step 2: choose the interval
This is the decision that most often gets made by accident — Excel defaults to one bar per day, and nobody revisits it.
CDC’s stated rule of thumb: “The interval of time should be appropriate for the disease in question, the duration of the outbreak, and the purpose of the graph. If the purpose is to show the temporal relationship between time of exposure and onset of disease, then a widely accepted rule of thumb is to use intervals approximately one-fourth (or between one-eighth and one-third) of the incubation period of the disease shown.”
Note the precise shape of that rule, because it is commonly mis-stated as “a quarter to a third”. The target is approximately one-fourth of the incubation period; the acceptable band runs from one-eighth to one-third.
Worked from CDC’s own stated incubation periods in the same course:
| Pathogen | Incubation period (as stated by CDC) | One-fourth | Practical interval |
|---|---|---|---|
| Salmonellosis | usually 12–36 hours | 3–9 hours | 12-hour bars (CDC’s own worked figure uses 12-hour intervals for this outbreak) |
| E. coli O157 | 1–8 days, usually 2–4 days | ~12–24 hours | 1-day bars |
| Coccidioidomycosis | average 12 days, minimum 7 days | ~3 days | 2–3 day bars |
| Hepatitis A | 15–50 days, average 28–30 days | ~7 days | 1-week bars |
Notice that CDC’s own salmonellosis figure uses 12-hour bars, which is longer than one-fourth of a 12-hour minimum incubation period. The rule is a starting point that gets reconciled against the second constraint — the duration of the outbreak. Too fine an interval on a short outbreak gives you a comb of single-case bars with no discernible shape; too coarse an interval collapses a two-peak propagated pattern into one lump.
Look the incubation period up; do not recall it
The interval rule is only as good as the number you feed it, and incubation periods are exactly the parameter that gets misremembered under time pressure. CDC’s course points investigators at “disease fact sheets available on the Internet or in the Control of Communicable Diseases Manual.” For the organisms an IP actually meets — norovirus, influenza, RSV, pertussis, Legionella, group A Streptococcus, Candida auris, invasive mould — pull the current figure from CDC’s disease page or CCDM at the time of the investigation and record the source in the investigation file. This page deliberately does not reproduce those numbers, because a stale incubation period baked into a reference page is worse than no number at all.
When the pathogen is unknown
Often you have a cluster before you have an organism — three cases of an undifferentiated febrile illness on one ward, cultures pending. CDC’s guidance for this case is empirical rather than formulaic: “Occasionally, you may be asked to draw an epidemic curve when you don’t know either the disease or its incubation time. In that situation, it may be useful to draw several epidemic curves with different units on the x-axis to find one that best portrays the data.”
Do this deliberately, not by trial and error until something looks like an outbreak:
- Plot the same case set at three or four intervals — for a hospital cluster, typically 6-hour, 12-hour, 1-day and 1-week bars.
- Choose the finest interval at which the shape is stable — that is, at which making the bars one step finer does not change the pattern you would report, only the noise.
- Show the reader which interval you chose and why. An interval selected because it produced the most alarming graph is a real and common failure, and it is invisible in the finished figure unless you disclose it.
Once the organism is identified, redraw at the rule-based interval and check whether your provisional reading survives. If the shape changes, the earlier reading was an artefact of binning.
Step 3: symptom onset versus specimen collection date
This is the decision with the largest effect on the finished curve, and it is the one hospital practice is forced into compromising on.
The definition demands onset. Onset is when the biology happened, so onset is the only date that stands in a fixed relationship to exposure — which is the entire basis of Step 5 below. Specimen collection date does not: it is separated from onset by a diagnostic delay that is a property of your organisation, not of the pathogen.
What plotting by specimen date does to the shape
- It shifts the whole curve to the right by the median time from onset to culture. If you then count back one incubation period from the peak, you land after the true exposure window by exactly that delay — the error is systematic, not random, and it does not average out.
- It smears the peak. Onset-to-culture delay varies by patient: an ICU patient with daily blood cultures is sampled within hours; a ward patient whose fever is attributed to something else is sampled two days later. A sharp point-source peak becomes a broad plateau, and a plateau is the signature of a different epidemic pattern. This is how a single contaminated procedure gets misread as a persistent reservoir.
- It imports your own operational rhythm. Specimens are collected when someone orders them; orders cluster on weekday day shifts and around ward rounds. Plot fine intervals by specimen date and you will see a periodicity that is the roster, not the outbreak. The tell is peaks landing on the same weekday, or troughs on both weekend days.
- It responds to your own investigation. The moment you declare a cluster and start screening, detection rises. On a specimen-date curve that appears as a rising limb — the outbreak looks like it is accelerating at precisely the moment you began looking harder.
The colonisation case, where onset does not exist
For MDRO transmission — C. auris, CRE, VRE, carbapenem-resistant Acinetobacter — there is frequently no symptom onset at all. The patient is colonised, not ill. There is nothing to plot on the axis the definition requires.
Be explicit about what you are drawing instead. A curve by first-positive-culture date for a colonisation cluster is a detection curve, not an epidemic curve: it shows when you found cases, which is a joint function of transmission and of testing intensity. Two consequences follow, and both change how the graph should be read:
- Point-prevalence surveys create artefactual spikes. Swab an entire 24-bed unit on a Tuesday and Tuesday gets a bar of six. That bar is a sampling event, not a transmission event. Annotate every PPS date directly on the figure. An unannotated PPS spike is the single most common way a colonisation curve gets misread as an explosive propagated outbreak.
- Absence of bars can mean absence of swabs. A flat stretch before the cluster was recognised usually means nobody was screening, not that nobody was colonised. Say so on the figure rather than letting the reader infer a baseline that was never measured.
Where admission screening exists, the acquisition interval — admission date to first positive — is a better transmission signal than the culture date alone, because it separates importation from in-facility acquisition. Which leads to the stratification that matters most in a hospital.
Stratify present-on-admission against hospital-onset
Shade each bar by whether the case was present on admission or hospital-onset — the same shading device CDC uses for culture confirmation, applied to the distinction that actually decides whether you have an outbreak. A curve rising on imported cases is a community signal arriving at your door. A curve rising on hospital-onset cases is transmission inside your building, and only the second one justifies a unit intervention.
This one graphical choice answers the question the executive team will ask first, and it costs one extra column in the line list.
What to actually record
Capture all four dates per case in the line list, and plot the best available with the choice stated on the figure:
- Date (and hour, for short-incubation pathogens) of symptom onset
- Specimen collection date
- Result / report date
- Admission date, plus unit and room history with dates
Then label the figure honestly: “cases by date of symptom onset”, or “cases by date of first positive culture (onset not recorded)”. A curve whose axis is undocumented is uninterpretable six months later at the review meeting, and it is the version that ends up in a report.
Step 4: read the shape
CDC: “The shape of the epidemic curve is determined by the epidemic pattern (for example, common source versus propagated), the period of time over which susceptible persons are exposed, and the minimum, average, and maximum incubation periods for the disease.”
Point source
“An epidemic curve that has a steep upslope and a more gradual down slope (a so-called log-normal curve) is characteristic of a point-source epidemic in which persons are exposed to the same source over a relative brief period. In fact, any sudden rise in the number of cases suggests sudden exposure to a common source one incubation period earlier.” And the diagnostic constraint: “In a point-source epidemic, all the cases occur within one incubation period.”
That constraint is a test you can apply directly. Take the maximum minus the minimum incubation period for the organism; if your cases span longer than that, it is not a simple point source, whatever the shape suggests. CDC works it explicitly for hepatitis A: incubation 15 to 50 days, so all point-source cases should fall within 50 − 15 = 35 days.
Hospital translations: a single contaminated infusion or flush lot; one contaminated procedure session; a single reprocessing failure on one device; one contaminated multi-dose vial; a single dialysis session. The common thread is a discrete event with a defined attendance list — which is exactly what makes point-source outbreaks tractable, because the exposure roster is retrievable.
Continuous common source
“If the duration of exposure is prolonged, the epidemic is called a continuous common-source epidemic, and the epidemic curve has a plateau instead of a peak.”
Hospital translations: a persistent environmental reservoir. Sink P-traps and drains, ice and water systems, a contaminated ice machine, an inadequately reprocessed endoscope in continuing rotation (see endoscope reprocessing and where it fails), heater-cooler units, a colonised ventilation or water pathway. The plateau ends when the reservoir is removed, and that timing is itself evidence: if cases stop within one incubation period of pulling a specific device from service, you have a strong temporal argument.
Where the plateau coincides with construction or renovation, the reservoir hypothesis should include airborne mould and dust — see infection control risk assessment for construction for what the ICRA should already have been controlling.
Intermittent common source
“An intermittent common-source epidemic (in which exposure to the causative agent is sporadic over time) usually produces an irregularly jagged epidemic curve reflecting the intermittence and duration of exposure and the number of persons exposed.”
Hospital translations: a procedure performed on a fixed schedule; a specific reprocessing cycle or shift; a product lot used intermittently; a colonised healthcare worker present on some rotations and not others. Jaggedness is a prompt to overlay the curve with schedules — theatre lists, dialysis sessions, staff rosters, delivery dates. This is where the epidemic curve stops being a graph and becomes a query against your own operational records.
Propagated
“In theory, a propagated epidemic — one spread from person-to-person with increasing numbers of cases in each generation — should have a series of progressively taller peaks one incubation period apart, but in reality few produce this classic pattern.”
Take that caveat seriously. The textbook propagated curve — clean generational waves — is rare, and it is essentially never visible in a hospital cluster of eight cases. Interventions truncate later generations, susceptibles run out on a closed unit, discharges remove cases from your denominator, and small numbers swamp the signal. Do not reject person-to-person transmission because the curve does not show tidy waves; the absence of the classic pattern is the norm, not evidence against propagation.
Hospital translations: hands and shared equipment, mediated by staff movement between patients. This is the pattern that puts transmission-based precautions, enhanced barrier precautions and donning and doffing competency on the intervention list, and it is the pattern where cohorting patients and staff has the most direct mechanistic rationale.
Read the outliers separately
CDC is emphatic that the cases outside the main body of the curve carry disproportionate information: “An early case may represent a background or unrelated case, a source of the epidemic, or a person who was exposed earlier than most of the cases… Similarly, late cases may represent unrelated cases, cases with long incubation periods, secondary cases, or persons exposed later than most others… On the other hand, these outlying cases sometimes represent miscoded or erroneous data. All outliers are worth examining carefully because if they are part of the outbreak, they may have an easily identifiable exposure that may point directly to the source.”
In practice, chase the earliest case personally before anything else. It is either a data error, an unrelated case that is inflating your outbreak, or the index — and each of those changes the investigation.
Step 5: bound the exposure window
This is the payoff, and it is the step generic explainers omit. A curve that only tells you “there is an outbreak” has not earned the hour it took to build. A curve that tells you “look at the 14th to the 18th” has.
CDC’s method, for an apparent point source of a known disease with a known incubation period, is three steps plus a widening:
- “Look up the average and minimum incubation periods of the disease.”
- “Identify the peak of the outbreak or the median case and count back on the x-axis one average incubation period. Note the date.”
- “Start at the earliest case of the epidemic and count back the minimum incubation period, and note this date as well.”
- “Ideally, the two dates will be similar, and represent the probable period of exposure. Since this technique is not precise, widen the probable period of exposure by, say, 20% to 50% on either side of these dates, and then ask about exposures during this widened period in an attempt to identify the source.”
Three details in that method are worth stating plainly:
- Two counts, two different incubation parameters. Peak uses the average; first case uses the minimum. Using the average for both is a common error and it pushes the first-case estimate too far back.
- Convergence is the validity check. If the two dates land close together, the point-source hypothesis survives and you have a window. If they are far apart, stop — you probably do not have a simple point source, and you should be reading Step 4 again rather than back-calculating.
- The 20–50% widening is not optional. A single-day exposure window from a small case series is false precision, and it will make your interviews miss the real exposure by asking about the wrong day.
CDC’s own worked example
From the same course, a community hepatitis A outbreak: first case onset in the week of 28 October, last in the week of 18 November — a span under one month, inside the 35-day point-source window, so the technique applies. Peak and median both fell in the week of 4 November; counting back one average incubation period (about a month) points to early October. The earliest case was the week of 28 October; counting back the 2-week minimum points to the week of 14 October. The two counts converge on the weeks of 7 and 14 October — “plus or minus a few days” — which turned out to be exactly when a food handler diagnosed in mid-October would have been shedding virus while still working.
Worked example, hospital shape
Illustrative composite. This is not a real outbreak, and no facility, patient, product or person in it is real — it is constructed to show the arithmetic on a hospital-shaped case series.
Six cases of a gastrointestinal illness on a 28-bed surgical ward, onsets recorded from the nursing notes. Suppose the identified organism has a stated incubation period of 1 to 8 days, usually 2 to 4 days.
- Interval: one-fourth of a 2–4 day usual incubation period is roughly 12–24 hours. Use 1-day bars; 12-hour bars would leave six single-case slivers with no shape.
- Span check: onsets run from day 10 to day 15 — six days, well within the 8 − 1 = 7 day point-source span. A point source remains tenable.
- Count back from the peak: peak and median onset both on day 12. Subtract the usual incubation period (3 days) — exposure around day 9.
- Count back from the first case: first onset day 10. Subtract the 1-day minimum — exposure on or after day 9.
- Converge and widen: both counts point at day 9. Widening by roughly 20–50% of the 3-day count gives a window of about day 8 to day 10.
The output is not a graph. It is an instruction to the investigation: pull everything that happened on that ward on days 8 to 10 — meal service, a shared bathroom, agency staffing, a procedure list, visitors — and interview against those three days rather than the whole fortnight. That is the difference a bounded window makes.
The inverse: estimating the incubation period from the curve
When the exposure is known but the organism is not, CDC gives the calculation in reverse: “Subtract the time of onset of the earliest cases from the time of exposure to estimate the minimum incubation period. Subtract the time of onset of the median case from the time of exposure to estimate the median incubation period. These incubation periods can be compared with a list of incubation periods of known diseases to narrow the possibilities.”
This is genuinely useful in a hospital where a single suspect event is already known — one theatre list, one dialysis session, one contaminated lot delivered on a datable day. The derived incubation period narrows the differential before the laboratory confirms it, and it can redirect specimen collection while there is still something to collect.
Where the curve misleads in a hospital
- Small numbers. Most hospital clusters are five to fifteen cases. Curve shape is not diagnosable at that scale — the difference between a plateau and a jagged line at n=8 is noise. Use the curve to bound the exposure window and to time the intervention; do not stake a transmission-route conclusion on the silhouette alone. If you need to say whether a count is genuinely above baseline, that is a question about expected counts, not shape — see the Poisson distribution.
- Ascertainment tracks effort. Cases appear when you look. Any curve that includes the period after cluster recognition is partly a graph of your own screening intensity. Mark the date the investigation started on the figure.
- Patients leave. A hospital population is open. Cases incubating at discharge present elsewhere — at home, in a nursing home, at another hospital — and never reach your line list, which truncates the tail and can make an ongoing outbreak look resolved. For any pathogen with an incubation period longer than the typical length of stay, the curve you can draw is structurally incomplete, and the report should say so.
- Healthcare exposure is rarely a moment. The point-source model assumes a discrete exposure. A patient with a central line has a continuing exposure for as long as the line is in. For device-associated infection the useful “exposure” is often a device-day range rather than a date, and the counting-back method degrades accordingly.
- The window may land where the records are thin. The commonest disappointment: you bound the exposure to three days, and those three days are a weekend with agency staffing and no environmental sampling. That is not a failure of the method; it tells you precisely which records to reconstruct, and it is a defensible reason to escalate for retrospective sampling.
- Do not confuse it with a rate. The curve counts cases, not risk. Six cases in a week on a full 28-bed unit and six on a half-empty one are different events. When you move from describing the outbreak to comparing groups, you need denominators — see prevalence versus incidence. The organism’s resistance pattern belongs with your cumulative antibiogram, not on the curve.
What to hand the outbreak team
One figure and five sentences. The figure: bars at a stated interval, axis labelled with the date type actually plotted, cases shaded by confirmation status and by present-on-admission versus hospital-onset, with PPS dates and the investigation start date annotated. The five sentences: the case definition in force, the interval and why, which date type is plotted and why, the pattern the shape is consistent with, and the bounded exposure window with its widening.
Then update the figure at every meeting and keep the superseded versions. A sequence of curves showing the epidemic flattening after a specific intervention date is the most persuasive evidence you will produce that the intervention worked — and CDC’s own caution applies to reading a live curve: with data only through the upswing “you might conclude that the outbreak is still on the upswing, with more cases to come”, where the same outbreak read a fortnight later “may soon be over”. Do not declare either from a curve you are still in the middle of.
Frequently asked questions
What interval should I use on an epidemic curve?
Approximately one-fourth of the incubation period, within an acceptable band of one-eighth to one-third, per CDC’s Principles of Epidemiology in Public Health Practice. Reconcile that against the outbreak’s total duration: if the rule gives you a comb of single-case bars, widen it; if it collapses distinct peaks into one bar, narrow it. State the interval you chose on the figure.
Should I plot by symptom onset or specimen collection date?
Onset, whenever it exists — it is the only date with a fixed relationship to exposure, and the exposure-window calculation depends on that relationship. Specimen collection date shifts the whole curve right by your diagnostic delay, smears sharp peaks into plateaus, and imports the ordering and staffing rhythm of your own organisation. When onset genuinely does not exist (colonisation, MDRO screening), plot first-positive-culture date and label the figure as a detection curve.
How do I estimate when the exposure happened?
Two counts. From the peak or median case, count back one average incubation period. From the earliest case, count back the minimum incubation period. If the two dates converge, that is your probable exposure period; widen it by 20% to 50% on either side before interviewing, because the technique is not precise. If they do not converge, reconsider whether you have a point source at all.
What does a plateau instead of a peak mean?
Classically, a continuous common source — prolonged exposure to a persistent reservoir rather than a single event. In a hospital that points at water systems, sink drains, ice, reprocessed devices or a colonised environmental niche. But check the artefact first: plotting by specimen collection date rather than onset can flatten a genuine point-source peak into a plateau all by itself.
How do I draw a curve when I do not know the pathogen yet?
Plot the same cases at several x-axis units — CDC’s explicit advice — and choose the finest interval at which the pattern is stable. Disclose which interval you chose. Redraw at the rule-based interval once the organism is identified, and check whether your provisional reading survives.
Is an epidemic curve the same as a histogram?
It is a specific type of histogram: the frequency distribution of cases over time, with adjoining bars whose area is proportional to the case count. That is why a line chart is the wrong display — it implies values between the plotted points that do not exist.
My cluster is only six cases. Is a curve worth drawing?
Yes, but for the exposure window and the timeline, not the silhouette. At single-digit case counts the shape cannot reliably distinguish epidemic patterns. The counting-back arithmetic still works and still narrows your interviews, which is where most of the value is.
Why does my curve dip every weekend?
Almost certainly because you plotted specimen collection or result dates. Cultures are ordered and processed on weekday rhythms. Replot by onset if you have it; if you do not, use a coarser interval that spans the weekly cycle, and note the artefact on the figure.
Sources
The method on this page — the interval rule of thumb, the epidemic curve definition, the classic patterns, and the counting-back procedure for the exposure window — is CDC’s, from Principles of Epidemiology in Public Health Practice, 3rd edition (the SS1978 self-study course), Lesson 4 (Displaying Public Health Data) and Lesson 6 (Investigating an Outbreak). All quotations above are verbatim from those lessons.
Access note: cdc.gov returned HTTP 403 to every direct retrieval attempt made while writing this page, and the RestoredCDC.org mirror returned 404 for these paths. The text was retrieved and quoted from CDC’s own archive at archive.cdc.gov (the www_cdc_gov/csels/dsepd/ss1978/ path), which served the lessons in full. Readers should treat CDC’s live pages as authoritative if and when they become reachable, and should confirm any incubation period against the current CDC disease page or the Control of Communicable Diseases Manual at the time of use rather than relying on a figure reproduced anywhere, including here.
The hospital-specific applications — the present-on-admission stratification, the point-prevalence-survey artefact, the diagnostic-delay effects on curve shape, and the open-population truncation — are reasoned extensions of that method to healthcare-associated infection practice, not direct CDC prescriptions. They are flagged as such where they appear.








