Skip to main content
v2026.11,610 entries · CC-BY 4.0

Nested Case-Control Studies: Risk-Set Sampling and Density Sampling

How risk-set (incidence-density) sampling draws matched controls from a cohort’s risk set at each case’s failure time, and why the resulting odds ratio estimates the incidence rate ratio — with a worked matched-set example.

Ask about Nested Case-Control Studies: Risk-Set Sampling and Density Sampling

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

A nested case-control study does not sample controls from the whole cohort at once — it samples them from the risk set at the exact moment each case occurs. That single design choice, called risk-set sampling (or, equivalently, incidence-density sampling), is what separates a nested case-control design from a standard case-control study built from scratch, and it is also what lets the resulting odds ratio estimate the incidence rate ratio directly rather than merely approximate a risk ratio under the rare-disease assumption. This guide stays inside that mechanism specifically: how a risk set is defined, how controls are drawn from it, how matching works inside that sampling frame, and a worked matched-set example. For the broader design comparison — cases vs. controls, retrospective exposure ascertainment, recall bias — see case-control study design and case-control vs. cohort study, which already cover that ground.

The problem: measuring exposure on an entire cohort is expensive

A prospective cohort study gets its statistical strength from following every enrolled member forward in time and observing who develops the outcome. That strength has a cost: if the exposure of interest requires a lab assay, a biobanked specimen, a manual chart abstraction, or an expensive imaging read, measuring it on the full cohort — most of whom will never develop the outcome — is often infeasible. A cohort of 50,000 followed for a decade to accumulate 200 incident cases does not need exposure data on all 50,000 people to answer “does this exposure predict this outcome”; it needs exposure data on the 200 cases and on a much smaller, carefully chosen comparison group. Nested case-control design exists to make that comparison group statistically valid without paying for full-cohort measurement.

What a “risk set” actually is

The risk set at a given point in time is the set of cohort members who are still under follow-up and still outcome-free at that instant — enrolled, not yet censored, and not yet an event. It changes continuously: as follow-up proceeds, people leave the risk set by developing the outcome, by being censored (lost to follow-up, study end, competing event), or, in some designs, simply by aging further from a matching window. Every time a case occurs, that case had a specific risk set at the moment of failure — everyone else in the cohort who was still eligible to become a case at that same point in time. Risk-set sampling means: for each case, define that risk set, then randomly sample one or more controls from it.

Risk-set sampling vs. sampling controls once from a fixed pool

This is the mechanical difference that the phrase “density sampling” is naming. A standard case-control study, even one built inside a cohort, is often described loosely as drawing controls from “the non-cases.” Risk-set (incidence-density) sampling is more specific and more consequential in three ways:

  • Sampling is time-anchored. Each case gets its own risk set, defined at its own failure time — not one shared pool of “everyone who never became a case,” which would only be knowable in hindsight, after the whole cohort’s follow-up is complete.
  • Sampling is with replacement across cases, and controls can later become cases. A cohort member sampled as a control for one case remains in the risk set for later cases and can be sampled again, and — because being a control at time t only requires being event-free at time t, not staying event-free forever — that same person can go on to become a case themselves later in follow-up. This is expected, not a design flaw; excluding future cases from eligibility as controls is what would bias the estimate.
  • The sampling fraction tracks person-time, not head count. Because risk sets shrink and are redrawn continuously, a person who is at risk for a long stretch of the study has more opportunities to be sampled as a control than someone who enrolls late or is censored early — which is exactly proportional to how person-time works in the full-cohort analysis this design is standing in for.

Contrast this with a case-cohort design, which samples a single subcohort once, at or near baseline, independent of when any case occurs, and reuses that same subcohort as the comparison group for every outcome under study. Case-cohort trades away the time-anchoring in exchange for a subcohort that can be reused across multiple outcomes without redrawing it — but it requires a weighted analysis (inverse-probability or Barlow-style weights) rather than the simpler matched analysis risk-set sampling supports. If a candidate is genuinely asking about that design rather than this one, it is a different comparison than a standard case-control study, and worth treating as a distinct topic.

Why the odds ratio estimates the incidence rate ratio, not just approximates the risk ratio

In a standard (non-nested) case-control study, the odds ratio approximates the relative risk only under the rare-disease assumption — the outcome has to be uncommon enough that the odds of exposure among non-cases closely tracks the true prevalence of exposure in the source population. See how to interpret an odds ratio for that mechanic in the standard design.

Risk-set sampling changes this. Because controls are drawn from the actual risk set present at each case’s failure time — the same set of people who were, at that instant, “at risk” of becoming that case — the resulting odds ratio from a matched analysis (typically conditional logistic regression, conditioning on the matched sets the sampling created) is a valid estimator of the incidence rate ratio the full cohort would have produced, without needing to invoke the rare-disease assumption as an approximation. This is the core efficiency argument for the design: it is not just “cheaper than measuring the whole cohort,” it is cheaper while targeting the same estimand a full-cohort survival or Poisson analysis would target, rather than a looser stand-in for it.

Matching inside a risk-set-sampled design

Matching is not required for risk-set sampling to work statistically, but it is used routinely for two practical reasons: it controls confounding by the matched factor without having to model it, and it keeps the comparison biologically or clinically sensible (comparing a case’s exposure to controls who were plausible substitutes for that case). The matching factor that is structurally built into risk-set sampling itself is time — calendar time, study time, or age at risk, depending on which time scale the analysis uses. Beyond that baseline time-matching, studies commonly add one or two further matching factors from a small, deliberately short list (sex, enrollment site, a strong known confounder) — matching on too many factors shrinks the eligible risk set for each case and can introduce its own problems (overmatching, where a factor correlated with the exposure itself is matched away, taking real signal with it).

Worked example: building one matched set

Illustrative example — a hypothetical walk-through, not drawn from a real published study or dataset; the numbers exist only to make the mechanics concrete.

A cohort of 8,000 participants is enrolled and followed for up to 12 years for a rare outcome. By year 6, participant #4,187 develops the outcome — this is a case, with a failure time of 6.0 years since enrollment. The risk set for this case is every other cohort member who, at the 6-year mark, is still enrolled, still outcome-free, and not yet censored: suppose that comes to 5,400 people. The study protocol calls for risk-set sampling with a 1:4 case-to-control ratio, matched on age at the 6-year mark (within 2 years) and sex. Filtering the 5,400-person risk set to those two matching criteria narrows it to, say, 310 eligible candidates; four are then drawn at random from that 310 and assigned to case #4,187’s matched set. Exposure history (in this illustration, a specific occupational exposure requiring a manual records review) is then abstracted only for the case and these four controls — not for the other 5,396 people who were technically eligible but not selected. That same process repeats independently for every other case the cohort produces, and the resulting case/matched-controls sets are analyzed together with conditional logistic regression, which conditions the model on each matched set so that only within-set exposure contrasts contribute to the odds ratio.

Choosing the number of controls per case

The statistical efficiency of a matched ratio does not scale linearly. Moving from 1 control per case to 2 or 3 produces a meaningful precision gain; moving from 4 to 8 produces very little additional gain relative to the added measurement cost, because the marginal information a matched analysis extracts from additional controls in the same set diminishes quickly. In practice, most nested case-control studies settle on somewhere between 1:1 and 1:4 unless the outcome is so rare that only a handful of cases will ever accumulate, in which case a higher ratio can be worth the added cost to stabilize the estimate.

What risk-set sampling does not fix

Risk-set sampling addresses efficiency and the estimand risk-set sampling targets (the rate ratio) — it does not, by itself, address confounding by factors that were not matched or adjusted for, misclassification in how exposure is measured for cases versus controls, or bias introduced if the exposure measurement itself (a chart review, an assay run years after specimen collection) differs systematically in quality between cases and controls. Those remain the same threats they are in any observational design and need the same handling: adjustment in the conditional model, sensitivity analysis, and, where biospecimens are involved, processing cases and controls identically and, ideally, blind to case status.

Reporting a nested case-control study

STROBE’s observational-study checklist treats nested case-control as a variant of the case-control design for reporting purposes, with the added expectation that the source cohort, the definition of the risk set, the matching factors, and the case:control ratio are all stated explicitly — a reader needs to be able to reconstruct exactly how each matched set was assembled. See the full STROBE checklist for the item-by-item requirements, and cohort study design for the source-cohort side of the design this guide assumes as a starting point.

Frequently asked questions

Is a nested case-control study the same as a case-cohort study?

No. Both draw a comparison group from inside an existing cohort, but a nested case-control study draws a new, time-matched risk set for every case, while a case-cohort study draws one subcohort at baseline and reuses it as the comparison group for every case and every outcome under study, at the cost of needing a weighted rather than a simple matched analysis.

Can a person be both a control and, later, a case?

Yes, and this is expected under risk-set sampling rather than an error to correct. Serving as a control at one case’s failure time only requires being event-free and under follow-up at that moment; it says nothing about that person’s status afterward.

Does the rare-disease assumption still apply in a nested case-control study?

No — that is one of the design’s main advantages over a standard case-control study. Because controls are sampled from the actual risk set at each case’s failure time, the odds ratio from the matched analysis estimates the incidence rate ratio directly, without relying on the outcome being rare to justify the approximation.

How many controls per case should a nested case-control study use?

Most studies use somewhere between 1 and 4 controls per case; statistical efficiency gains taper sharply past about 4 controls per case relative to the added cost of measuring exposure on each additional control, so a higher ratio is usually only worth it when very few cases will ever accumulate.

What has to be reported for a nested case-control study to meet STROBE?

The source cohort, how the risk set was defined at each case’s failure time, the matching factors and case:control ratio used, and the analysis method (typically conditional logistic regression) — see the full STROBE checklist for the complete item list.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.