Written and maintained by CASRAI Editorial Board
Last updated
“How many interviews is enough?” is the question every qualitative researcher gets asked by a reviewer, an IRB, or a funder — and “until we reach saturation” has been the default answer for decades. The problem is that saturation is a judgment you can only make after the data is in, which makes it a weak thing to write into a proposal or a methods section before you’ve collected anything. Kirsti Malterud, Vibeke Siersma, and Ann Dorrit Guassora’s 2016 information power model was proposed as a more defensible replacement: instead of asking “did new codes stop appearing,” it asks a set of questions you can actually answer at the design stage. This guide covers both — what saturation claims, why it has real documented problems as a stopping rule, and how to use information power to write a methods-section justification that holds up before a single interview happens.
What “Data Saturation” Actually Means
The idea originates in grounded theory: Barney Glaser and Anselm Strauss’s 1967 The Discovery of Grounded Theory described theoretical sampling continuing until new data stopped adding new properties to the core categories — what they called theoretical saturation. That is a narrower, more specific claim than how the term is used today. Most qualitative studies outside grounded theory borrow the word loosely to mean something closer to thematic or code saturation: you keep interviewing until a new transcript stops surfacing codes that aren’t already in the codebook. See the grounded theory guide for the original theoretical-saturation criterion in its native context.
The landmark empirical reference for code saturation is Guest, Bunce, and Johnson’s 2006 study in Field Methods, “How Many Interviews Are Enough? An Experiment with Data Saturation and Variability.” Working with 60 interviews from a relatively homogeneous population, they found the great majority of codes had already appeared by the twelfth interview, with only marginal additions after that. That single finding is the source of the “12 interviews” heuristic that circulates widely in qualitative methods teaching — including in disciplines and sample types that look nothing like Guest, Bunce, and Johnson’s original, relatively homogeneous study population.
Why “We Reached Saturation” Is a Weaker Justification Than It Looks
Saturation is intuitive and it isn’t wrong, but four documented problems limit how much weight it can carry on its own, especially as an a priori sample-size justification:
- It’s retrospective, not predictive. You cannot know in advance when new codes will stop appearing — which means “we will continue to data saturation” is a real stopping rule but not a plannable one. A grant budget, an IRB recruitment cap, or a pre-registration that needs a concrete number can’t be satisfied by a criterion that’s only checkable after the fact.
- The operational definition is unstable. “No new codes” and “no new meaning” are not the same stopping point. A study can hit code saturation — the codebook itself stops growing — well before it reaches what some methodologists call meaning saturation, where the depth, nuance, and range within each existing code stops growing. Hennink, Kaiser, and Marconi’s 2017 replication in Qualitative Health Research is the standard citation for this distinction: they found code saturation arriving commonly cited as around the single digits of interviews in their dataset, well before meaning saturation, which took roughly double that. The exact counts are dataset-specific, but the direction of the finding — codes stop before meaning does — is the reason “we reached saturation” needs a definition attached to mean anything specific.
- It’s sensitive to who is coding and how. Saturation is observed by an analyst applying a codebook, not measured by an instrument. Two coders working the same transcripts with different granularity, different a priori sensitizing concepts, or different stopping tolerances can report different saturation points from identical data.
- The critique has its own name in the literature. O’Reilly and Parker’s 2013 paper “Unsatisfactory Saturation” (Qualitative Research) argues the term is used inconsistently across studies, is frequently invoked without being operationally defined at all, and sits awkwardly inside research traditions — narrative inquiry, some strands of discourse analysis, single-case designs — where the underlying idea of “no new codes across cases” was never really the goal in the first place.
Seeing the Problem Directly: A Reproducible Simulation
The numbers below come from a seeded random simulation run for this page, not from a real study. It’s an illustration of how “no new codes” stopping rules behave, not a claim about any actual dataset.
The simulation models 25 underlying codes with different “prevalence” — some common, some rare — and simulates 30 interviews in which each code has a fixed chance of surfacing in any given interview (deterministic PRNG, fixed seed 20260826; independently rerunnable). Two common a priori stopping rules were applied to the same simulated sequence:
- Naive rule (“stop after one interview with zero new codes,” not evaluated before interview 10): triggered at interview 11, with 16 of the 23 codes that eventually appeared already captured — missing 7 of them, i.e. roughly 30% of the eventual codebook was still to come.
- Stronger a priori rule, in the spirit of Francis et al.’s (2010) recommendation to fix a run length in advance (“stop after three consecutive interviews with zero new codes,” not evaluated before interview 10): triggered at interview 18, with 20 of 23 codes captured — better, but still missing 3 codes that went on to appear at interviews 19, 21, and 26.
Neither rule is unreasonable, and the longer run length clearly performs better. But both were fooled by an ordinary, unremarkable lull — several interviews in a row that happened not to surface anything new, followed by another genuine addition much later. That’s the concrete version of the “retrospective, not predictive” problem above: any fixed run-length rule is a bet about how long a lull can plausibly last, decided by convention rather than by anything you can determine about your specific study before you start.
Malterud’s Alternative: Information Power
Malterud, Siersma, and Guassora proposed information power in “Sample Size in Qualitative Interview Studies: Guided by Information Power” (Qualitative Health Research, 26(13), 2016, pp. 1753–1760). The core claim: the more information a sample holds that’s relevant to the actual study, the fewer participants are needed — and how much information a sample holds can be assessed across five dimensions before you recruit anyone:
- Study aim — narrow vs. broad. A tightly scoped research question needs less data to answer than a sprawling, exploratory one.
- Sample specificity — dense vs. sparse. A purposively selected, information-rich, homogeneous sample yields more per participant than a broad, loosely defined one.
- Use of established theory — applied vs. not. A study building on a well-developed theoretical framework needs fewer participants than one starting from an unstructured, purely exploratory position.
- Quality of dialogue — strong vs. weak. Rich, probing interviews with genuine reflection produce more usable information per interview than short, superficial ones.
- Analysis strategy — case-focused vs. cross-case. An in-depth narrative or single/few-case analysis needs fewer participants than a cross-case strategy explicitly looking for variation across many cases.
Strong performance across several dimensions — a narrow aim, a dense purposive sample, real theoretical grounding, high-quality dialogue, case-focused analysis — means a defensibly rigorous study can run on a considerably smaller sample than a broad, exploratory, cross-case one. That is the opposite of a fixed number: it’s a rubric for reasoning about sample size design, not a target headcount.
Writing the Methods-Section Case: Saturation vs. Information Power
The practical difference shows up in what you can honestly write before data collection starts.
A saturation-only justification, written in advance, can only say something like: “Recruitment will continue until data saturation is reached, anticipated at approximately N interviews based on comparable studies.” That’s a real, defensible sentence — but it’s a prediction borrowed from someone else’s study, not a property of yours, and reviewers increasingly know it.
An information-power justification lets you argue from your own study’s design instead of an analogy:
“This study addresses a narrowly defined aim [state it], drawing on a purposively selected, information-dense sample of participants who share [the specific relevant characteristic]. The interview guide was developed from an established theoretical framework [name it], and semi-structured, in-depth interviews were used to maximize dialogue quality. Given the combination of a narrow aim, high sample specificity, applied theory, and an in-depth (rather than cross-case) analysis strategy, a sample of N participants was judged to hold sufficient information power for this analysis; saturation of the coding categories relevant to the study aim was also confirmed during analysis.”
Note the structure: information power justifies the planned sample size before data collection, and saturation is kept as a secondary, confirmatory check made during analysis — not as the sole a priori justification. That combination answers both the “why did you plan for N” question and the “how do you know N was enough” question, which a saturation-only account can only really answer to the second one, and only in hindsight.
Where Saturation Still Has a Real Role
None of this makes saturation wrong — it makes it insufficient as the only justification, particularly for the a priori sample-size statement a proposal or protocol needs. Theoretical saturation remains the correct, native stopping criterion inside grounded theory specifically, where it’s tied to the constant comparative method rather than bolted on afterward — see the grounded theory guide. And confirming that no new codes emerged during analysis is still worth reporting as a rigor check alongside an information-power design rationale, per the COREQ/SRQR reporting guidelines for qualitative research, which expect sample-size adequacy to be addressed explicitly either way.
Common Mistakes
- Citing “12 interviews” as a universal rule. That number comes from one study (Guest, Bunce & Johnson 2006) with a specific, fairly homogeneous population and a specific research question. It is not a general threshold, and reviewers familiar with the literature will recognize it as a borrowed number rather than a reasoned one.
- Using “saturation” without defining which kind. State explicitly whether you mean code saturation, meaning saturation, or theoretical saturation — they are not interchangeable and produce different stopping points from the same data.
- Treating information power as a formula that outputs a number. It’s a structured judgment across five dimensions, not a calculator. Two studies can reasonably reach different sample sizes from the same aim if their sample specificity or dialogue quality differ.
- Dropping saturation checking entirely. Information power justifies the plan; it doesn’t replace confirming, during analysis, that the plan held up. Report both.
Frequently Asked Questions
What is data saturation in qualitative research?
The point in data collection at which additional interviews or observations stop producing meaningfully new codes, themes, or theoretical properties. It originates in grounded theory’s theoretical saturation and is now used more loosely across qualitative traditions as thematic or code saturation.
How many interviews are needed to reach saturation?
There is no universal number. The widely cited “12 interviews” figure comes from one 2006 study of a specific, relatively homogeneous population; other studies have found saturation much earlier or considerably later depending on sample heterogeneity, research aim breadth, and which type of saturation (code vs. meaning) is being measured.
What is information power in qualitative research?
A sample-size framework proposed by Malterud, Siersma, and Guassora (2016) that judges sample adequacy across five dimensions — study aim, sample specificity, use of established theory, dialogue quality, and analysis strategy — rather than by counting interviews until no new codes appear.
Can I use both saturation and information power in the same study?
Yes, and it’s the stronger approach: use information power to justify the planned sample size before data collection, then report during analysis whether saturation of the relevant coding categories was also observed, as a confirmatory check.
Does information power replace saturation entirely?
No. It replaces saturation as the sole a priori justification for a planned sample size. Saturation — particularly theoretical saturation within grounded theory specifically — remains a legitimate concept for describing what happened during analysis.








