Written and maintained by CASRAI Editorial Board
Last updated
Berkson’s bias is not about how many people you sampled — it is about which people your sampling frame let in. It is the specific selection-bias mechanism that shows up when a study draws its participants from a hospital, clinic, or referral population rather than the general population, and the exposure and the disease under study — independently of each other — each raise the odds of being admitted. Because admission itself now depends on both, an association can appear between the exposure and the disease among admitted patients that does not exist in the population those patients were drawn from, or a real association can be masked entirely. It was first described by the biostatistician Joseph Berkson in 1946, and it remains one of the more common ways a well-conducted hospital-based or clinic-based study reaches a wrong answer without any error in the statistics.
What creates Berkson’s bias
The mechanism needs three things, all of which are ordinary in clinical research rather than exotic:
- A sample drawn from people who were admitted, referred, or enrolled into a specific setting — a hospital ward, a specialist clinic, a disease registry fed by referrals — rather than sampled directly from the general population.
- Two variables of interest — typically an exposure and a disease, or two diseases — that are genuinely independent (or have some true relationship you’re trying to measure) in the general population.
- Admission to the sampled setting that is influenced, separately, by each of the two variables. Either one, on its own, makes admission more likely; having both makes it even more likely; having neither makes it comparatively unlikely.
Under those conditions, being in the sample carries information. A patient who lacks one of the two conditions had to have something else account for their admission — and that something else is disproportionately the other condition. The result is a negative association between the two variables inside the admitted sample, or in other configurations a positive one, that reflects nothing about their real-world relationship. This is the same causal structure as collider bias in general — admission is a variable with arrows pointing into it from both the exposure and the disease, and conditioning the analysis on that variable (by only ever looking at admitted patients) manufactures the spurious link. Berkson’s bias is the specific, sampling-restriction version of that general mechanism.
The classic example: Berkson’s 1946 hospital study
Berkson introduced the problem in “Limitations of the Application of Fourfold Table Analysis to Hospital Data,” Biometrics Bulletin, 1946, using hospital in-patient data and the illustrative pairing of diabetes and cholecystitis (gallbladder disease) — two conditions with no established causal relationship to each other. His point was that a patient without diabetes who is nonetheless in the hospital is more likely to have some other admitting condition, such as cholecystitis, than a person in the general population is — simply because the non-diabetic patient needed a reason to be there. That pattern produces a spurious negative association between the two conditions among hospitalized patients, even though nothing links them outside the hospital. (The paper is frequently miscited as a 1949 Biological Bulletin article; the original is the 1946 Biometrics Bulletin piece.)
The following table is an illustrative worked example built to demonstrate the mechanism numerically — the prevalence and admission-rate figures are round, hypothetical values chosen for clarity, not Berkson’s own published data. Start with a hypothetical population of 100,000 people in which diabetes and gallbladder disease are genuinely independent, each affecting 10% of the population:
| Group | Population count | Admission rate | Admitted count |
|---|---|---|---|
| Diabetes + gallbladder disease | 1,000 | 40% | 400 |
| Diabetes only | 9,000 | 20% | 1,800 |
| Gallbladder disease only | 9,000 | 15% | 1,350 |
| Neither condition | 81,000 | 3% | 2,430 |
In the general population, the odds ratio between the two conditions is exactly 1.0 — they are independent by construction. Restrict the analysis to the 5,980 admitted patients, and the odds ratio becomes (400 × 2,430) / (1,800 × 1,350) ≈ 0.40: diabetic patients now appear to have roughly 60% lower odds of gallbladder disease than non-diabetic patients, purely as an artifact of who was admitted. A researcher working only from the hospital chart data, with no way to see the general-population numbers, would have no way to know the true odds ratio is 1.0 rather than 0.40 unless they recognized the sampling structure itself as the problem.
Why the association reverses inside the hospital sample
The intuition generalizes beyond the specific numbers above: whenever admission is driven independently by two conditions, presence of one condition becomes indirect evidence against needing the other to explain the admission. A patient admitted with a clear diagnosis of Condition A didn’t need Condition B to get through the door, so among admitted patients, having A slightly lowers the estimated chance of also having B, compared to the general population. The distortion gets larger as the two admission-driving effects get stronger and as the base admission rate for people with neither condition gets smaller — both of which are common in real hospital and clinic populations, where healthy people with neither condition mostly aren’t there at all.
Modern equivalents: registry and EHR-based studies
Berkson described the mechanism using paper hospital charts, but nothing about it is specific to 1946-era data collection. The same structure shows up anywhere a study’s sampling frame is a population selected by clinical contact rather than drawn from the general population:
- Disease registries fed by referral. A registry that enrolls patients referred to a tertiary center captures a population whose referral itself was driven by disease severity, comorbidity burden, or specialist availability — not a random cross-section of everyone with the condition. Two comorbidities that are unrelated in the general population can appear associated (or dissociated) within the registry purely because of what drove referral.
- EHR-derived case-control and cohort studies. A study built entirely from a single health system’s electronic health records inherits that system’s admission and encounter patterns as its sampling frame. Patients captured in the records are, by definition, people who interacted with that system — and if the exposure and outcome under study each independently make a healthcare encounter more likely, the same collider-conditioning distortion applies to an EHR cohort exactly as it applied to Berkson’s paper charts.
- Multi-condition claims or billing data. Administrative claims data captures diagnoses that were coded during a billable encounter. If the presence of one condition increases the likelihood of an encounter where a second, unrelated condition also gets coded, the two conditions can show a spurious co-occurrence pattern in the claims data that isn’t present in the underlying patient population.
The fix is the same across all three: the risk comes from the sampling frame, not from the vintage of the data source, so it has to be addressed at the design or analysis stage, not assumed away because the data happen to be electronic and large.
How to tell if your study is at risk
Three questions flag the risk before analysis starts:
- Is the sample defined by clinical contact? Hospital admission, clinic referral, registry enrollment, or an EHR/claims encounter are all clinical-contact sampling frames. A sample drawn by population-based random sampling, or by a design that enrolls before the health event of interest (e.g., a birth cohort enrolled at birth, followed forward), is not at risk from this specific mechanism.
- Do both variables of interest independently affect the odds of being in the sample? If only one of the two variables plausibly influences admission or enrollment, the collider-conditioning mechanism doesn’t apply in the same way — it needs both incoming arrows.
- Is the comparison group (e.g., hospital controls) drawn from the same clinical-contact population as the cases? Using hospital-based controls admitted for unrelated conditions doesn’t avoid the problem; it can reproduce it, since the controls’ presence in the hospital is itself selected on the same admission mechanism.
Berkson’s bias vs. related quantitative-analysis pitfalls
Berkson’s bias is easy to conflate with adjacent concepts that produce similar-looking distortions through different mechanisms:
| Concept | What actually happens | How it differs from Berkson’s bias |
|---|---|---|
| Collider bias (general) | Conditioning — by sampling or by statistical adjustment — on a common effect of two variables creates a spurious association between them. | Berkson’s bias is the specific case where the conditioning happens through sample selection into a clinical setting, rather than through a covariate added to a model. |
| Selection bias (general) | Any systematic difference between who is sampled and who should have been sampled to answer the research question. | Selection bias is the umbrella category; Berkson’s bias is one specific, well-characterized mechanism within it, defined by the dual-cause admission structure. |
| Confounding | A third variable causes both the exposure and the outcome, biasing their apparent relationship even with no selection involved at all. | A confounder has arrows pointing out to both variables; the collider behind Berkson’s bias has arrows pointing in from both. The fix for confounding (adjust for it) is the opposite of the fix for Berkson’s bias (don’t condition on the collider). |
| Immortal time bias | A period of follow-up during which the outcome could not have occurred by design gets misclassified into the wrong exposure group. | A time-classification error, not a sample-composition error — it doesn’t require a clinical-contact sampling frame at all. |
Guarding against it
- Prefer a population-based sampling frame when the research question is about a general-population relationship. If that’s not feasible, be explicit that any hospital- or clinic-based estimate describes the clinical-contact population, not the general population, and say so in the paper’s limitations.
- Choose controls carefully in a hospital-based case-control design. Controls admitted for a condition known to be unrelated to both the exposure and the outcome under study reduce, though don’t always eliminate, the risk; controls drawn from the general population avoid it entirely where that’s practical.
- Check whether the suspicious association appears in any population-based data on the same question. An association strong in hospital-based data and absent or reversed in population-based data is a signature worth investigating as a Berkson’s-bias candidate before treating it as a real finding.
- Map the admission mechanism as a causal diagram before analysis, the same discipline used to check for collider bias generally — draw the arrows from the exposure and the outcome into “was sampled,” and ask whether restricting to the sampled group creates a path between them that wasn’t there before.
Frequently asked questions
Is Berkson’s bias the same thing as selection bias?
Berkson’s bias is a specific form of selection bias, not a synonym for it. Selection bias covers any systematic mismatch between the sample and the population of interest; Berkson’s bias names one particular mechanism — differential admission driven independently by both variables under study — that produces that mismatch in hospital- and clinic-based samples specifically.
Does using a large hospital dataset fix the problem?
No. Like other forms of collider bias, Berkson’s bias is a structural artifact of which patients ended up in the sample, not a statistical-power problem. A larger hospital-based sample produces a more precisely estimated wrong answer, not a less biased one, because every additional patient was selected through the same distorting admission mechanism.
Can Berkson’s bias make a real association disappear instead of creating a fake one?
Yes. The mechanism can work in either direction depending on how the admission probabilities line up across the four exposure-disease combinations — it can manufacture an association where none exists, exaggerate a real one, attenuate a real one, or in some configurations reverse its direction entirely. The table-based check in the “how to tell if your study is at risk” section above applies regardless of which direction the resulting distortion runs.
Does Berkson’s bias affect randomized controlled trials?
Rarely in the way described here, because randomization happens after enrollment, not through it. It becomes relevant to a trial’s external validity if the trial’s own enrollment criteria function like a hospital admission filter — for example, requiring a comorbidity-linked referral to reach the trial site — but it is primarily a concern for observational hospital- and clinic-based designs, not for the randomized comparison itself.
Related quantitative-analysis topics
Berkson’s bias sits inside a small family of sample-composition problems covered in more general form elsewhere on CASRAI: see Collider Bias for the underlying causal-diagram mechanism, Selection Bias for the broader category of sampling mismatches, Sampling Bias for a fuller taxonomy of how samples stop representing their target population, Spurious Correlation for the general phenomenon of associations that don’t reflect a real underlying relationship, Confounding Variable for the third-variable structure that Berkson’s bias is often confused with, and Immortal Time Bias for a differently-mechanized bias common in the same clinical-data settings.








