Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Snowball Sampling: Definition, Method, and When to Use It

How snowball (chain-referral) sampling works, its variants including respondent-driven sampling (RDS), strengths and well-documented biases, ethical issues in referral-based recruitment, and how to report it in a Methods section.

Snowball sampling (also called chain-referral sampling) is a non-probability sampling method in which existing study participants recruit future participants from among their own acquaintances. A researcher starts with a small number of initial contacts — the seeds — who meet the study’s eligibility criteria, then asks each of them to refer other eligible people they know. Those referrals become the next wave of participants, who in turn refer further contacts, and the sample grows outward in successive waves like a snowball rolling downhill.

It is the standard method for reaching hidden, stigmatized, or hard-to-identify populations for which no public list or sampling frame exists — people who share an illegal, highly private, or socially sensitive characteristic or experience, members of an informal or dispersed professional network, or any group that can only realistically be located through personal trust and referral. The method was formally described by Patrick Biernacki and Dan Waldorf in their 1981 paper “Snowball Sampling: Problems and Techniques of Chain Referral Sampling” (Sociological Methods & Research, 10(2), 141–163), based on their work recruiting a sample of ex-opiate addicts who were not in treatment and therefore invisible to any institutional list.

This guide covers how snowball sampling actually works, its main variants (including respondent-driven sampling), when it is the right choice, its well-documented biases and limitations, the ethical issues specific to referral-based recruitment, and how to report it in a Methods section.

How Snowball Sampling Works

Snowball sampling proceeds in three basic steps:

  1. Identify seeds. The researcher recruits a small initial group of eligible participants — often through a gatekeeper, a community organization, an advertisement, or the researcher’s own existing contacts. Seed selection matters: a narrow or homogeneous set of seeds tends to produce a narrow, homogeneous final sample, because referral chains generally stay within existing social networks.
  2. Ask for referrals. Each seed is asked to identify or introduce other people who meet the study’s eligibility criteria. Depending on the study design, this might mean the participant directly contacts their acquaintance, provides contact information to the researcher (with the acquaintance’s prior consent to be contacted), or hands the researcher’s contact card to the referral.
  3. Repeat across waves. Each new participant is, in turn, asked for further referrals. The sample expands in waves (sometimes called referral chains or “generations”) until the researcher reaches a target sample size, the referral chains stop producing new eligible contacts, or the sample reaches thematic/informational saturation (a common stopping rule in qualitative snowball studies).

When to Use Snowball Sampling

Snowball sampling is the appropriate choice when a probability sampling frame is not available or not feasible, and one or more of the following applies:

  • The population is hidden or stigmatized. No public registry, membership list, or administrative record identifies members — for example, people engaged in an illegal or heavily stigmatized behavior, or people who share a private and sensitive characteristic they may not disclose to strangers.
  • The population is rare or dispersed. Members exist in low absolute numbers spread across a wide geographic or social area, making random sampling from a general population prohibitively expensive (screening thousands of people to find a handful of eligible respondents).
  • Trust is a precondition for participation. Potential participants are more likely to take part, and to answer honestly, when approached through someone they already trust rather than a stranger or an institution.
  • The research is exploratory or qualitative. Early-stage or qualitative studies (interviews, ethnography, grounded theory) often prioritize depth and access over statistical representativeness, which fits snowball sampling’s strengths.

It is not the right choice when the research question requires statistically representative, generalizable estimates of a population with a known or constructible sampling frame — in that situation, a probability method such as stratified sampling or simple random sampling is the appropriate tool.

Variants of Snowball Sampling

Several named variants refine the basic chain-referral idea:

  • Linear snowball sampling. Each participant refers exactly one further contact, producing a single linear chain. Rarely used in practice because it grows the sample too slowly and is highly vulnerable to a single broken link ending the chain.
  • Exponential non-discriminative snowball sampling. Each participant refers as many contacts as they can, and every eligible referral is enrolled, so the sample grows geometrically wave over wave.
  • Exponential discriminative snowball sampling. Each participant may refer multiple contacts, but the researcher enrolls only one (or a limited subset) per referring participant, balancing growth speed against over-concentration within any single participant’s network.
  • Respondent-driven sampling (RDS). A more rigorous refinement introduced by sociologist Douglas Heckathorn in “Respondent-Driven Sampling: A New Approach to the Study of Hidden Populations” (Social Problems, 44(2), 174–199, 1997). RDS adds a structured incentive system (participants are typically compensated both for their own participation and for successfully recruiting others, up to a capped number of referrals each) and a mathematical weighting model that uses information about each participant’s network size and the recruitment pattern to adjust for the fact that people with larger networks are more likely to be recruited. Under specific assumptions, RDS can produce population-level estimates with calculable confidence intervals — something ordinary snowball sampling cannot do — which is why it is the preferred method in public-health surveillance of hidden populations (for example, WHO- and CDC-supported studies of people who inject drugs, sex workers, and men who have sex with men in HIV surveillance).

Strengths and Limitations

Strengths

  • Often the only feasible way to reach a hidden population. When no sampling frame exists and cold outreach is impractical or unsafe, referral through trusted contacts may be the only realistic access route.
  • Low cost and logistically simple relative to building or purchasing a sampling frame, especially for populations that are geographically dispersed.
  • Higher response and disclosure rates than approaching strangers, because referred participants arrive with a degree of pre-established trust in the researcher via the person who referred them.

Limitations

  • No known probability of selection. Because participants are not drawn from a defined sampling frame with known selection probabilities, standard formulas for statistical inference and margin of error do not apply to ordinary (non-RDS) snowball samples, and findings cannot be assumed to generalize to the wider population.
  • Homophily and network bias. People tend to refer others similar to themselves (in demographics, attitudes, or the specific characteristic under study), so the sample can systematically over-represent certain sub-groups within the population and under-represent people who are more isolated or who belong to a different social cluster.
  • Seed bias. The characteristics of the initial seeds can shape the entire referral tree; a poorly chosen or overly narrow set of seeds skews the whole sample from the outset.
  • Exclusion of network isolates. Anyone in the target population who is not connected, even indirectly, to the seeds’ social networks has essentially zero chance of being sampled.
  • Gatekeeping and masking effects. Participants may deliberately withhold referrals (to protect a contact’s privacy or their own) or refer only people they believe will make the group “look good,” further skewing the sample.

Snowball Sampling vs. Purposive and Convenience Sampling

All three are non-probability sampling methods and are sometimes confused with one another:

  • Snowball sampling relies specifically on participant-to-participant referral to build the sample outward in waves — the defining feature is the chain-referral mechanism, not just researcher discretion.
  • Purposive sampling (also called judgment sampling) has the researcher directly select participants based on specific characteristics relevant to the research question, without relying on referral chains. Snowball sampling is sometimes described as a special case of purposive sampling in which the selection criterion at each step is “known to, and vouched for by, a prior participant.”
  • Convenience sampling recruits whoever is easiest to reach — on the basis of proximity, availability, or willingness — with no systematic referral structure and no deliberate matching to specific characteristics.

See the broader taxonomy in CASRAI’s Sampling Methods: Probability and Non-Probability Types Explained dictionary entry for how all of these fit alongside probability methods such as stratified sampling.

Ethical Considerations in Snowball Sampling

Referral-based recruitment raises confidentiality issues that do not arise with other sampling methods, and Institutional Review Boards (IRBs) and Research Ethics Committees typically scrutinize them closely:

  • Indirect disclosure of participation or status. When Participant A refers Participant B, A necessarily learns (or confirms) that B meets the study’s eligibility criteria — which, for a study of a stigmatized characteristic or behavior, can amount to A learning something sensitive about B that B did not choose to disclose directly to A. Protocols should minimize this: for example, having participants pass along study contact information rather than passing personal information to the researcher, so the referred person controls whether and when to make contact.
  • Consent of the referred, not just the referrer. The referring participant’s consent to take part in the study does not substitute for the referred person’s own informed consent; each new participant must independently consent before any data are collected from them.
  • Avoiding coercion within the chain. Referral incentives (including RDS-style dual compensation) should be structured so that participants do not feel pressured to recruit, and referred individuals do not feel pressured to participate because a friend or relative asked them to.
  • Confidentiality of the referral network itself. Depending on the topic, the pattern of who-referred-whom can itself be sensitive information (it can reveal social or criminal associations); researchers should consider whether network data need to be collected, stored, or reported at all, and if so, how to de-identify it.

How to Report Snowball Sampling in a Methods Section

A Methods section using snowball sampling should specify:

  • How and why the seeds were identified, and their relevant characteristics (this lets readers assess potential seed bias).
  • The referral mechanism used (participant-to-participant contact, researcher-mediated, contact-card distribution, etc.).
  • The number of waves/referral generations and the final sample size, ideally with a simple diagram or table showing growth across waves.
  • Any incentive structure offered for successful referrals, and whether it differed from the incentive for the participant’s own participation.
  • Whether respondent-driven sampling weighting was applied and, if so, the software/estimator used (for example, RDS Analyst or the RDS package in R) and the network-size question used to derive weights.
  • An explicit acknowledgment that findings are not statistically generalizable to the broader population unless RDS (or an equivalent weighted design) was used and its assumptions are defensible for this population.

Worked Example

The following is an illustrative composite, not a real study, included to show how the method works in practice. A researcher studying the support needs of informal (unpaid) caregivers for family members with a rare degenerative condition has no registry of caregivers to sample from — the condition is rare enough that no single clinic or patient registry captures more than a handful of cases, and many caregivers do not identify with any formal caregiver organization. The researcher recruits five seed participants through two rare-disease patient advocacy groups, interviews each one, and at the end of each interview asks whether the participant knows other caregivers in a similar situation who might be willing to talk. Over four waves, the sample grows to 34 participants. The researcher reports the seed sources, the number of waves, and the final sample size in the Methods section, and explicitly notes that because all five seeds came through advocacy-group contacts, caregivers with no connection to any advocacy organization are likely under-represented in the resulting sample — a limitation acknowledged directly rather than glossed over.

Frequently Asked Questions

Is snowball sampling a probability or non-probability sampling method?

Non-probability. Because participants are recruited through referral chains rather than drawn from a defined sampling frame with a known probability of selection, standard statistical formulas for margin of error and confidence intervals do not apply to an ordinary snowball sample. Respondent-driven sampling (RDS) is a specific, more rigorous variant that adds network-size weighting specifically to allow calculable population estimates, but it requires its own set of design assumptions to be defensible.

What is a real-world example of when snowball sampling is used?

It is widely used in public-health and social-science research on hidden or hard-to-reach populations — for example, studies of people who inject drugs, sex workers, undocumented migrants, or people with a rare and highly stigmatized condition — where no sampling frame exists and participants are more likely to take part when approached through someone they already trust.

What is the difference between snowball sampling and respondent-driven sampling?

Respondent-driven sampling (RDS) is a structured refinement of snowball sampling: it adds a capped, dual-incentive recruitment structure and a mathematical weighting model based on each participant’s network size, which allows researchers to calculate population-level estimates with confidence intervals. Ordinary snowball sampling has neither the structured incentive cap nor the weighting model, and its results are not treated as statistically generalizable.

What is the main limitation of snowball sampling?

Homophily and network bias: participants tend to refer people similar to themselves, so the final sample tends to over-represent tightly connected sub-groups and under-represent anyone outside the recruitment chains that happen to form, including people entirely disconnected from the seeds’ social networks.

How many waves of referral are needed for snowball sampling?

There is no fixed number. Researchers typically continue recruiting waves until they reach a predetermined target sample size, the referral chains stop producing new eligible contacts, or — in qualitative studies — the data reach thematic/informational saturation, meaning additional interviews are no longer surfacing new themes.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →