Skip to main content
v2026.11,610 entries · CC-BY 4.0

Prevalence vs. Incidence: Definitions, Formulas, and How They Relate

Prevalence counts existing cases; incidence counts new ones. This guide defines both, works through P ≈ I × duration, and covers risk ratio, rate ratio, odds ratio, and attributable risk.

Ask about Prevalence vs. Incidence: Definitions, Formulas, and How They Relate

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Prevalence is the proportion of a defined population that has a condition at a given point or over a given period. Incidence is the rate at which new cases of that condition occur in a population over time. The two are frequently used interchangeably in casual writing and just as frequently confused in manuscripts — but they answer different questions, are calculated from different denominators, and behave differently over time. This guide defines both precisely, works through the formula that links them, and covers the measures — risk ratio, rate ratio, odds ratio, attributable risk — that are built on top of them.

What is prevalence?

Prevalence is the proportion of a population that has a condition, existing cases and all, measured at (or over) a specific time. It answers “how much of this condition is out there right now?” rather than “how fast is it appearing?” Two variants are distinguished by their time frame:

  • Point prevalence — the proportion of the population with the condition at a single point in time (a specific date, or the moment of a survey). Formula: existing cases at time t ÷ total population at time t.
  • Period prevalence — the proportion of the population that had the condition at any time during a specified interval (a calendar year, for example), including cases present at the start of the period plus new cases that arose during it. Formula: (existing cases at start of period + new cases during period) ÷ average population during period.

Both are proportions — unitless, bounded between 0 and 1 (or expressed as a percentage, or as cases per 1,000 or 100,000 population) — and both count every person who has the condition, regardless of when it started, as long as they had it during the window being measured.

What is incidence?

Incidence counts only new cases — people who did not have the condition at the start of the observation period and developed it during that period. Because it tracks the rate of new occurrence rather than the existing burden, incidence is the measure epidemiology uses to study causation: it tells you how quickly a condition is appearing in a population that was previously free of it. Incidence has two common forms, and mixing them up is one of the most frequent errors in observational research reporting.

Cumulative incidence (risk)

Cumulative incidence — often just called risk — is the proportion of a defined, disease-free population at the start of follow-up that develops the condition over a fixed period:

Cumulative incidence = new cases during the period ÷ population at risk at the start of the period

This is the natural measure in a closed cohort — a fixed group followed for a fixed period with no entries or exits — and it is interpretable as a probability: “an X% risk of developing the condition over this many years.” It breaks down, or at least needs adjustment, when people enter the cohort at different times, are lost to follow-up, or die of unrelated causes before the observation period ends, because the simple proportion no longer reflects each person’s actual time at risk.

Incidence rate (person-time)

Incidence rate expresses new cases relative to the total time each person actually spent at risk, summed across the population, rather than relative to a head count:

Incidence rate = new cases during the period ÷ total person-time at risk

Person-time (commonly person-years) is the sum, across everyone in the study, of the time each individual was observed and disease-free. This handles the two problems cumulative incidence struggles with directly: a participant followed for 2 years before being lost to follow-up contributes 2 person-years rather than being treated identically to someone followed for the full 10; a participant enrolled partway through a study still contributes the time they were actually observed. The tradeoff is interpretability — an incidence rate is not a probability and does not have an intuitive upper bound the way risk does (a rate of “0.08 cases per person-year” is not a percentage), and it implicitly assumes the rate is roughly constant over the follow-up window, which is worth checking rather than assuming.

How prevalence and incidence relate: P ≈ I × D

In a population at steady state — incidence, average disease duration, and the size of the at-risk population all roughly constant over the timeframe considered — prevalence, incidence, and average disease duration are linked by a simple relationship:

Prevalence ≈ Incidence rate × Average duration of disease

(The exact relationship is P/(1−P) = I × D; when prevalence is low, P/(1−P) is close enough to P that the simplified form above is the one used in practice.)

This relationship is the reason prevalence and incidence can move in opposite directions in response to the same underlying change — a result that trips up a lot of otherwise careful readers of health statistics. Consider a treatment that extends survival with a condition without curing it: it does not change incidence (the rate at which new cases arise is unaffected), but it increases average duration D, and increasing D while holding I constant increases prevalence. So an intervention that is unambiguously good for individual patients — they live longer with the condition — can simultaneously make the condition look more common in cross-sectional statistics, purely because more people are living with it at any given moment rather than dying or being cured quickly. The reverse also holds: a cure that removes people from the prevalence pool quickly (short D) can hold prevalence low even while incidence is rising, and a purely preventive intervention that stops new cases (lowers I) without affecting how long existing cases last will lower prevalence only gradually, as the existing caseload ages out. Reading a change in prevalence as a change in incidence — or vice versa — without checking what happened to duration is one of the most common misinterpretations of disease-frequency statistics in both scientific and popular reporting.

When to use prevalence vs. incidence

Use case Right measure Why
Studying causes and risk factors (aetiology) Incidence Only new cases isolate the population actually at risk during the exposure window; prevalent cases mix in survivors of long-past exposure, diluting or confounding an exposure-outcome association (see prevalent-case bias below).
Health-service planning, staffing, budgeting Prevalence Service demand tracks the number of people who currently have the condition and need care, regardless of when they developed it.
Estimating population disease burden at a point in time Prevalence Burden is a snapshot question — how many people, right now, are living with this.
Evaluating whether a prevention programme is working Incidence A falling incidence means new cases are genuinely being prevented; prevalence can stay flat or even rise for the reasons above while incidence is already falling.

Measures built on incidence and prevalence

Comparing incidence or risk between an exposed and unexposed (or treated and untreated) group produces the standard family of comparative effect measures used throughout observational and experimental research:

  • Risk ratio (relative risk) — cumulative incidence in the exposed group ÷ cumulative incidence in the unexposed group. Interpretable directly as “X times the risk.”
  • Rate ratio — the same comparison using incidence rates (person-time denominators) rather than cumulative incidence; used when follow-up time varies across participants.
  • Risk difference (attributable risk) — cumulative incidence in the exposed group minus cumulative incidence in the unexposed group. Unlike the ratio measures, this is expressed in the same units as the outcome and answers “how many additional cases per population does the exposure cause,” which is the figure most directly useful for public-health impact estimates.
  • Odds ratio — the odds of exposure among cases divided by the odds of exposure among controls (or equivalently, odds of outcome in exposed vs. unexposed).

The odds ratio deserves its own note because it is the measure a case-control study is built to produce, and cannot generally be substituted with a true risk ratio. A case-control design fixes the ratio of cases to controls by the researcher’s sampling decision, not by the true prevalence or incidence of the condition in the source population — so cumulative incidence, incidence rate, and therefore risk ratio or rate ratio cannot be validly calculated from a standard case-control sample. The odds ratio, calculated from exposure odds rather than outcome risk, sidesteps that problem. Under the “rare disease assumption” (the condition affects a small fraction of the source population), the odds ratio closely approximates the risk ratio a cohort study of the same population would have produced; when the outcome is common, the two diverge and should not be treated as interchangeable. See case-control study vs. cohort study for the full design-level comparison.

Why prevalence drives diagnostic-test performance

Sensitivity and specificity are properties of a test itself and do not mechanically change with how common a condition is in the population being tested. Predictive values do — and prevalence is exactly the quantity that moves them. Positive predictive value (PPV), the probability that someone who tests positive actually has the condition, rises as prevalence rises and falls as prevalence falls, even when sensitivity and specificity are held perfectly constant. Concretely: a test with 95% sensitivity and 95% specificity, applied in a specialist clinic where prevalence of the condition among the tested population is high, will have a high PPV — most positives really are positives. The identical test, with identical sensitivity and specificity, applied as a population screening tool where prevalence is low, will produce a much lower PPV, because the much larger pool of true negatives generates enough false positives (5% of a very large negative group) to outnumber the true positives (95% of a much smaller positive group). This is precisely why a screening-positive result in a low-prevalence population is routinely followed by a confirmatory test rather than treated as diagnostic on its own — and why the same test can be described, accurately, as both “highly accurate” (a property of sensitivity/specificity) and “not very trustworthy on its own here” (a statement about PPV at low prevalence) depending on which population it’s used in.

Common errors in reporting prevalence and incidence

  • Calling a prevalence an incidence, or vice versa. A cross-sectional survey measures prevalence; it cannot report an incidence unless it specifically asks about condition onset within a defined recall window. Headlines and abstracts routinely blur this — “diabetes is rising” from a prevalence figure alone conflates a possible rise in incidence with a possible rise in survival duration (see the P ≈ I × D relationship above).
  • Leaving the denominator undefined. “Cases per population” is meaningless without specifying which population — the general population, an age-restricted subgroup, a clinic’s patient list, a country’s registered residents — and over what exact time window. Reporting a raw case count without a clearly defined denominator and time frame is not a prevalence or incidence figure at all.
  • Prevalent-case bias (Neyman bias) in cohort studies. Enrolling people who already have the condition (prevalent cases) into a study designed to examine incidence or early risk factors systematically excludes anyone who died or recovered quickly after onset, biasing the sample toward survivors with less severe or slower-progressing disease. Studying incident (newly diagnosed) cases avoids this.
  • Comparing crude rates across populations without age standardisation. Crude incidence or prevalence rates are heavily influenced by a population’s age structure — comparing the crude rate of an age-related condition between an older and a younger population will show a difference driven substantially by age composition rather than by any underlying difference in risk. Direct or indirect age standardisation (adjusting each population’s rate to a common reference age structure) is the standard correction, and its absence is a common, avoidable weakness in cross-population comparisons.

Reporting prevalence and incidence in a manuscript

The STROBE statement (STrengthening the Reporting of OBservational studies in Epidemiology) is the reporting checklist most journals point authors to for cohort, case-control, and cross-sectional studies. It does not prescribe which measure to use, but its expectations are directly relevant to prevalence and incidence reporting: state explicitly how cases were defined and ascertained, define the denominator population precisely (including eligibility criteria), state the exact time window or follow-up period, and report participant flow (including losses to follow-up, which matter more for incidence-rate calculations than for point-prevalence ones). See the EQUATOR Network guide for how STROBE relates to the other major reporting guidelines, and the study design types guide for how design choice (cross-sectional vs. cohort vs. case-control) determines which of prevalence or incidence a study can validly report in the first place. Also report confounding variables considered and how they were handled, and, where relevant, whether rates were age-standardised and against what reference population.

Frequently asked questions

Is prevalence always higher than incidence?

Not necessarily, but for chronic, long-duration conditions it typically is, because prevalence accumulates existing cases over an extended average duration while incidence counts only new cases in a given period. For short-duration, quickly resolving conditions the two figures can be much closer.

Can incidence be higher than prevalence?

Yes, in principle, particularly for conditions with a very short duration (rapid recovery or high fatality) — if D is small, P ≈ I × D can be smaller than I itself even with substantial new-case activity.

What is period prevalence used for if point prevalence already exists?

Period prevalence is useful when a condition is episodic or intermittent (it may not be present at the exact moment of measurement but was present at some point in the interval) or when a study only has the practical ability to survey a population once, over an extended window, rather than at one instant.

Does a case-control study report prevalence or incidence?

Neither, in general. Because a case-control study’s ratio of cases to controls is set by study design rather than reflecting the true frequency of the condition in the source population, it cannot validly produce a prevalence or an incidence estimate for that population — it produces an odds ratio for the exposure-outcome association instead.

Why does prevalence change more slowly than incidence?

Prevalence is a stock (an accumulated pool of existing cases), while incidence is a flow (the rate new cases enter that pool). A change in incidence only fully shows up in prevalence once the existing pool of cases, accumulated under the old incidence rate, has had time to resolve or turn over — which can take years for a chronic condition.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →