Skip to main content
v2026.11,772 entries · CC-BY 4.0

Maximum Variation Sampling: Choosing Dimensions and Defending the Range

How to choose variation dimensions that do real evidentiary work (not just demographic ones), build and report the sampling matrix, and write the specific methods-section justification that an achieved range counts as maximum variation rather than convenience diversity.

Ask CASRAI · included with Regulatory Radar

Ask about Maximum Variation Sampling: Choosing Dimensions and Defending the Range

Ask CASRAI answers research-administration questions about this guide and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Written and maintained by CASRAI Editorial Board

Last updated

Maximum variation sampling (also called heterogeneity sampling) is a purposive sampling strategy: instead of narrowing a sample to a single well-defined group, a researcher deliberately selects cases that span the widest credible range on a small set of dimensions relevant to the research question, then looks for patterns that hold despite that variation. A theme that recurs across a genuinely heterogeneous group is harder to dismiss as an artifact of one setting, one role, or one type of participant — that is the entire evidentiary logic of the technique. For the full family of purposive strategies this sits inside (homogeneous, typical case, extreme case, critical case, expert sampling), see Purposive Sampling: Choosing Cases on Purpose. This guide focuses on the two places maximum variation sampling most often goes wrong in practice: picking dimensions that don’t actually do any evidentiary work, and describing the achieved sample as “diverse” without showing why that diversity was deliberate rather than a post-hoc description of whoever agreed to participate.

What Maximum Variation Sampling Is Actually Claiming

The technique makes a specific, checkable claim: the sample was constructed, in advance, to cover the range on named dimensions, and any theme found across that range is more credible evidence of a shared, cross-cutting phenomenon than the same theme found in a narrow or accidental sample would be. That claim only holds if two things are both true and both demonstrated, not just asserted:

  • The dimensions were chosen before recruitment, for a reason tied to the research question — not selected afterward because they happened to be the axes along which the achieved sample varied.
  • The achieved range is reported against the possible range, so a reader can see how much of the relevant variation was actually captured, not just that “participants varied in background.”

Skip either one and the write-up reads as convenience sampling wearing a maximum-variation label — a distinction covered in more general terms in Convenience Sampling: When It’s Defensible and When It Isn’t. The rest of this guide works through both requirements in order: choosing dimensions, then defending the range.

Choosing Variation Dimensions — Not Just Demographic Ones

The most common weakness in a maximum-variation methods section is a dimension list that is easy to measure but not actually theoretically load-bearing: age, gender, and years of experience, chosen because they were on the intake form, not because the research question predicts they matter. Demographic dimensions are legitimate when the research question is genuinely about how a phenomenon differs by demographic group — but for most study questions, the dimensions that do real evidentiary work fall into three other categories:

  • Structural or positional dimensions — the participant’s role, level of institutional authority, or position in a process (e.g., principal investigator vs. research coordinator vs. participant; centralized vs. distributed service model).
  • Experiential dimensions — exposure duration, severity, timing, or prior experience with the phenomenon under study (e.g., newly diagnosed vs. long-term patients; first policy cohort vs. a cohort several years into implementation).
  • Contextual dimensions — the setting, environment, or governing conditions the participant operates under (e.g., institution type, regulatory environment, resource level, urban vs. rural setting).

A dimension earns a place in the matrix by answering one question: if this dimension were held constant instead of varied, would a reviewer reasonably suspect the finding might not generalize past that one value? If yes, vary it and say why in the methods section. If the honest answer is “probably not,” it’s a tag for the participant table, not a sampling dimension — padding the matrix with dimensions that don’t do evidentiary work only shrinks the number of participants available per cell.

In practice, most defensible maximum-variation designs use two to four dimensions. Beyond that, a small qualitative sample (commonly 12–25 participants for this strategy) runs out of people to fill the resulting cells, which is the subject of the next section.

Building and Reporting the Sampling Matrix

Once dimensions and their levels are fixed, cross them into a matrix — one axis per dimension, one cell per combination of levels. The matrix does two jobs: it disciplines recruitment (target cells guide who gets approached, rather than recruitment happening first and dimensions being read off the result), and it becomes the actual evidence, in the methods section, of how much of the possible range was covered.

Illustrative composite example (not a real study) — a researcher studying how early-career faculty experience a new institutional data-management mandate sets three a priori dimensions: discipline (lab science, humanities, social science, applied/professional — 4 levels), institution type (R1 research university, teaching-focused four-year, community college — 3 levels), and career stage (pre-tenure 0–5 years, tenured 6+ years — 2 levels). Crossed, that is 4 × 3 × 2 = 24 possible cells. With a target sample of 18 participants, one participant per distinct cell fills 18 of the 24 cells — 75.0% cell coverage, leaving 6 combinations unrepresented (a realistic outcome: full factorial saturation is rarely achieved or necessary for maximum variation sampling to do its job). Per-dimension, the achieved allocation still touches every level of every dimension: discipline counts of 5/4/5/4 across its four levels, institution-type counts of 7/6/5 across its three levels, and career-stage counts of 9/9 — each dimension’s least-represented level still has at least 4 participants, so no single level of any dimension is riding on one person’s account. Reporting that allocation — the matrix itself, or a summary table of it, in the methods section or an appendix — is what turns “a diverse sample” into a checkable claim.

The Specific Justification a Methods Section Needs

This is the part that separates a defensible maximum-variation claim from an after-the-fact diversity narrative. A methods section (or an IRB/ethics submission describing the sampling plan) needs to do all of the following, not just describe the sample as varied:

  1. State that the dimensions were set a priori, before recruitment, and tie each one to the research question or a specific theoretical expectation — “we expected experience of the mandate to differ by discipline because compliance burden differs by data type” is a defensible sentence; “participants came from a range of disciplines” is not.
  2. Report the possible range, not just the achieved sample’s description. Name the full set of levels on each dimension (or the theoretical min/max on a continuous one) so a reader can judge coverage — the 24-cell / 18-participant / 75% example above is the kind of statement this produces; “a mix of disciplines and institution types” is not.
  3. Name the theme that held across the range — the actual analytic payoff of the strategy. A maximum-variation methods section that never returns, in the results, to “and this pattern appeared across all four disciplines and all three institution types” has done the sampling work but not the reporting work that makes the strategy pay off.
  4. Disclose what was NOT varied, and why. Every study bounds its sampling frame somewhere — institution country, language, a minimum tenure, an eligibility window. Naming the boundary is not a weakness to hide; it is what keeps the maximum-variation claim honest about what range it actually covers.
  5. Distinguish the recruitment channel from the sampling logic. A snowball or convenience-recruited sample that happens to turn up varied participants is not maximum variation sampling unless the dimensions were set first and recruitment was actively steered to fill under-represented cells (follow-up targeted recruitment, screening questions keyed to the matrix, purposive selection from a larger volunteer pool). If recruitment wasn’t steered by the matrix, name the method for what it actually was — see Convenience Sampling and Snowball Sampling for the honest alternative framings.

A reviewer who has seen the weak version many times is specifically checking for items 2 and 5 above — a range that’s asserted but not shown, and a recruitment method that doesn’t match the sampling claim. Both are checkable from the write-up alone, which is exactly why they’re worth getting right rather than treated as boilerplate.

Sample-Size Implications

Maximum variation sampling typically needs a larger N than a homogeneous or typical-case purposive design pursuing the same depth of understanding, because credibility depends on the theme recurring across multiple cells, not just on reaching thematic saturation within one group. That doesn’t mean pursuing full factorial coverage of every cell — as the worked example above shows, a defensible design can leave a meaningful share of cells empty as long as every level of every dimension is represented by more than a single participant. Prioritize which dimensions and levels are essential to the argument and which are secondary before recruitment starts, so that if time or access runs out, the gaps that remain are the ones that matter least. For the general mechanics of justifying a sample size (qualitative or quantitative) in a methods section, see Justifying Sample Size in a Manuscript and Purposive Sampling‘s “How many cases is enough?” section, which covers saturation-based reasoning that applies here too.

When Maximum Variation Sampling Is the Wrong Choice

  • The research question is about one well-bounded group, not about what holds across a heterogeneous one — use homogeneous sampling instead, which holds variation constant on purpose rather than maximizing it.
  • The goal is statistical generalization to a defined population — maximum variation sampling is a non-probability technique; it cannot support population inference the way stratified sampling or probability sampling can, even though the matrix logic looks superficially similar to stratification.
  • The theory is still emerging and sampling needs to follow what the data reveals, cell by cell, rather than being fixed in advance — that is theoretical sampling, covered in Grounded Theory Methodology, not maximum variation.

Frequently Asked Questions

How is maximum variation sampling different from convenience sampling?

Convenience sampling selects whoever is easiest to reach and does not claim to control which dimensions vary. Maximum variation sampling sets specific dimensions in advance and recruits (or screens volunteers) to fill the range on those dimensions deliberately. A sample can look diverse under either approach; only the second can honestly claim the diversity was constructed rather than incidental.

How many dimensions should a maximum variation design use?

Most defensible designs use two to four. Each additional dimension multiplies the number of matrix cells, and a small qualitative sample runs out of participants to fill them long before it runs out of dimensions worth varying.

Does maximum variation sampling produce a representative sample?

No. It is a non-probability, purposive technique aimed at capturing the range of a phenomenon and testing whether a theme holds across that range — not at producing a sample whose composition mirrors a defined population’s. That is what stratified sampling is for.

What’s the difference between maximum variation sampling and stratified sampling?

Stratified sampling divides a known population into strata and samples (often randomly) within each stratum, in proportions tied to the population, to support statistical inference. Maximum variation sampling has no population frame to sample from proportionally — it deliberately seeks the extremes and breadth of a phenomenon among whoever is eligible and reachable, for analytic contrast rather than representativeness.

How large does a maximum variation sample need to be?

There is no fixed formula. It depends on the number of dimensions and levels chosen and on reaching the point where the cross-cutting theme is no longer producing meaningfully new variation — the same saturation logic used across qualitative sampling generally, covered in Purposive Sampling.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.