Inclusion and exclusion criteria are the rules that decide, before recruitment starts, which
individual units — people, records, sites, documents, whatever the unit of analysis is — are
eligible for a study and which are not. Getting them right is a design act, not paperwork: every
criterion you add narrows who the study’s findings can speak for, and every criterion you leave
out risks pulling in cases so different from each other that no single effect or pattern can be
detected. This page gives you a working method for drafting criteria (a worksheet you can apply to
any study), two fully worked examples — one clinical, one survey-based — and a dedicated look at
the most common way criteria go wrong: making them so tight the study can no longer generalize to
anyone. For the step that comes before this one, see CASRAI’s guide to
how to write a research question; for what
comes after, see power analysis and
sample size calculation and simple random
sampling.
Inclusion criteria and exclusion criteria are two sides of one boundary
The two terms describe the same boundary from opposite directions, not two independent lists:
- Inclusion criteria are the characteristics a unit must have to be eligible —
they define the target population positively. “Adults aged 18-65 with a confirmed diagnosis of
type 2 diabetes” is an inclusion criterion. - Exclusion criteria are the characteristics that remove an otherwise-eligible
unit from the study — they carve exceptions out of the population the inclusion criteria already
defined. “Currently enrolled in another interventional trial” is an exclusion criterion.
A criterion should live on only one list. Writing “must not have condition X” as an inclusion
criterion and then separately listing “has condition X” as an exclusion criterion states the same
rule twice and is a sign the criteria set was drafted without a single pass to de-duplicate it —
worth checking for directly when you review a draft protocol.
Both lists exist to answer one question for every candidate unit: is this case similar enough
to the others to be combined with them, and different enough from the general population that the
research question actually applies to it? That is a definitional act — it is closely related to,
but distinct from, the sampling mechanics covered in CASRAI’s guide to
simple random sampling, which assumes the eligible
population is already defined and addresses how you select from within it.
Why criteria exist: the trade-off they’re actually managing
Every inclusion or exclusion criterion trades internal validity against external validity —
see CASRAI’s internal vs. external validity
comparison and guide to generalizability
for the fuller treatment of that trade-off. Tighter criteria (a narrower age band, excluding
anyone with a comorbidity, requiring a specific baseline severity) reduce the noise between units
and make it more likely a real effect, if one exists, is detectable — that is a gain in internal
validity. But every one of those same restrictions also shrinks the population the result can
honestly be said to describe. A study result about “otherwise-healthy adults aged 40-55 with no
comorbid conditions” is not evidence about anyone outside that description, no matter how
statistically robust the finding is within it. Drafting criteria is the act of deciding, explicitly
and in advance, where on that trade-off a given study needs to sit — see the
dedicated section below on what happens when that trade-off is made
without thinking about it.
A worksheet for drafting criteria
Work through each category below and ask, for every candidate criterion you’re considering, the
single test in the right-hand column. A criterion that fails the test — one you can’t justify
against the research question itself — is a candidate for cutting.
| Category | What it controls for | Typical criteria | The justification test |
|---|---|---|---|
| Demographic | Who the population is (age, sex, occupation, role) | Age range; specific job title or role; student vs. faculty status | Does the research question name this group specifically, or would including a wider group still let you answer it? |
| Clinical / behavioral | Baseline condition, severity, prior exposure or treatment history | Confirmed diagnosis; symptom severity threshold; prior treatment naive | Is this condition part of what the question is asking about, or a convenience filter to get a “cleaner” sample? |
| Temporal | When the unit’s data or experience occurred | Enrolled within the last 12 months; diagnosis within a defined window | Does the timeframe matter to the question (e.g., a policy or protocol change), or is it arbitrary? |
| Geographic / institutional | Setting, site type, jurisdiction | Single country; specific institution type (academic medical center vs. community clinic) | Is the setting itself a variable the question depends on, or just where recruitment happens to be feasible? |
| Data availability / completeness | Whether the unit can actually be measured | Complete baseline data on file; able to read/write the study language; internet access for an online survey | Is this a genuine measurement requirement, or does excluding on it introduce a bias (e.g., excluding non-native speakers systematically excludes recent immigrants)? |
| Ethical / safety | Whether participation itself is safe or appropriate | Capacity to consent; not currently pregnant (for a teratogenic exposure); no contraindication to a study procedure | Is there a real, documented safety or capacity basis, or is this criterion doing double duty as a convenience filter? |
Two drafting habits catch most of the errors that show up later, during review or at analysis:
- Write a one-line rationale next to every criterion as you draft it. If you
cannot state in one sentence what the criterion protects against or why the population needs that
boundary, it is very likely a convenience criterion masquerading as a methodological one, and it
should either be justified properly or dropped. - Pre-specify criteria before you see any data, and do not revise them once
recruitment or data collection is underway based on how the data are turning out. A criterion
added after seeing preliminary results — even with a plausible-sounding rationale — is a form of
the same after-the-fact rationalization problem CASRAI covers in its guide to
HARKing: it changes
what counts as “the study” after some of the answer is already known.
Worked example: a clinical study
Illustrative example — this is a constructed, generic scenario for demonstration, not a
description of any real trial or institution. A study is testing whether a structured
medication-adherence coaching program improves adherence among adults recently diagnosed with
hypertension.
| Criterion | Type | Rationale |
|---|---|---|
| Age 18-75 | Inclusion | Defines the adult population the intervention is designed for; upper bound excludes an age range where comorbidity burden would confound adherence measurement. |
| New diagnosis of hypertension within the past 6 months | Inclusion | The intervention targets early habit formation; a “new diagnosis” population is the one the research question is actually about. |
| Able to read and understand the coaching materials’ language | Inclusion | Genuine measurement requirement — the intervention cannot be delivered otherwise. Flagged for a translated-materials sub-study rather than silently excluding non-English speakers long-term. |
| Current enrollment in another adherence-related interventional study | Exclusion | Controls for contamination between interventions, which would make it impossible to attribute any adherence change to this program specifically. |
| Cognitive impairment that would prevent independent completion of the coaching program | Exclusion | Safety/feasibility criterion — documented basis, not a convenience filter, since the intervention requires independent task completion by design. |
| Life expectancy under 12 months per treating clinician | Exclusion | The outcome (adherence over a 6-month follow-up) is not meaningfully measurable in this group. |
Notice what is not on this list: no exclusion for common comorbidities like
well-controlled type 2 diabetes or obesity, and no exclusion based on baseline adherence itself.
Excluding on either would have quietly redefined the study population as “healthier and more
adherence-prone than the real population this program will actually be used with” — the exact
problem covered next.
Worked example: a survey study
Illustrative example, generic and not describing a real institution or dataset. A
study is surveying research administrators about the length of time it takes to close out a
federal grant after the period of performance ends.
| Criterion | Type | Rationale |
|---|---|---|
| Currently employed in a sponsored-programs or research-administration role | Inclusion | Defines the population with direct, first-hand knowledge of the closeout process being studied. |
| Personally closed out at least one federally funded award in the past 24 months | Inclusion | Ensures respondents are answering from recent, direct experience rather than general impression — reduces recall bias. |
| Employed at a U.S. institution receiving federal funds directly (not solely as a subrecipient) | Inclusion | The research question is specifically about prime-recipient closeout timelines, which follow a different process than subrecipient closeout. |
| Respondent works exclusively in a purely administrative-support role with no closeout responsibility | Exclusion | Removes respondents who could complete the survey but would be answering about a process they don’t actually perform. |
| Incomplete survey response (fewer than 80% of items answered) | Exclusion | A data-quality threshold set and published in advance, not applied selectively after seeing which responses look favorable. |
The same worksheet categories apply to a survey as to a clinical study — the “demographic”
category here is professional role rather than age, and “clinical/behavioral” becomes “direct,
recent experience with the process being studied,” but the underlying logic (does this criterion
serve the question, or just convenience?) does not change by design type.
How over-tight criteria destroy generalizability
The single most common failure in criteria drafting is not too few criteria — it is too many,
each individually defensible, that compound into a population so narrow the result cannot be
applied to anyone the study was actually meant to inform. This shows up in a few recurring
patterns:
- The healthy-volunteer effect. Stacking exclusions for every comorbidity,
concurrent medication, and borderline lab value produces a sample healthier and more compliant
than the real population the intervention or finding will eventually be applied to. The result can
be internally valid and still misleading in practice, because the people it was tested on are not
representative of who it will be used on. - The explanatory-vs-pragmatic trade-off. A tightly restricted study (an
“explanatory” design, in the terminology clinical trial methodology uses) answers “can this work
under ideal conditions,” while a more loosely restricted, more representative study (a “pragmatic”
design) answers “does this work under real-world conditions.” Neither is wrong — but a study drawn
up with explanatory-tight criteria and then discussed as if it answers the pragmatic question is a
generalizability error, not a data error. Say explicitly which question your criteria were built
to answer. - Criteria that quietly select on the outcome. An exclusion criterion that
correlates with the outcome being measured (for example, excluding anyone with “poor prior
adherence” from an adherence study) does not just narrow the population — it can bias the
estimated effect itself, because it removes exactly the cases most informative about the
question. - Convenience criteria dressed as methodology. A criterion whose real purpose
is “easier to recruit” or “easier to measure” rather than anything about the research question
(excluding anyone without reliable internet access for a study with no online component, for
instance) narrows the sample without buying any real gain in validity. The one-line-rationale habit
in the worksheet above is specifically aimed at catching these.
The practical check: after your criteria list is complete, describe the population it defines
in one plain sentence, then ask whether that sentence still matches the population your research
question was actually about, or whether it has quietly become “the population that was easiest to
study.”
Reporting your criteria
Whatever the study type, eligibility criteria belong in the methods section, stated in enough
detail that another researcher could apply them independently and reach the same enrollment
decision on a given case. Established reporting guidelines make this an explicit, checked item
rather than an optional courtesy: the CONSORT statement for randomized trials requires reporting
eligibility criteria for participants (see CASRAI’s
CONSORT statement guide), and the
STROBE guideline for observational studies has an equivalent reporting item covering how the study
population was defined and selected. Reviewers and readers use this section to judge how far a
study’s findings can be extended beyond the sample actually studied — it is doing real
methodological work, not just satisfying a template.
Common mistakes
- Circular criteria. Defining eligibility in terms of the outcome you’re about
to measure (“must show improvement” as an inclusion criterion for a study measuring improvement)
makes the finding true by definition rather than by evidence. - Undefined thresholds. “Clinically significant” comorbidity, “regular” social
media use, “adequate” organizational capacity — any criterion using a qualitative threshold needs
an operational definition (a specific scale, cutoff, or documented source) or two people applying
it independently will not agree on borderline cases. - Revising criteria mid-study without documentation. Legitimate protocol
amendments happen, but an undocumented, informal loosening or tightening of criteria partway
through recruitment breaks comparability between early- and late-enrolled units and should be
logged as a formal, dated amendment, not applied quietly. - Treating “exclusion” as the inverse of every inclusion criterion by default.
Exclusion criteria should identify specific, additional reasons to remove an otherwise-eligible
unit — not simply restate the inclusion list in the negative, which adds length without adding
information. - Setting criteria after recruitment has already started. Even a well-justified
new criterion introduced mid-recruitment can bias the sample if it wasn’t applied uniformly to
everyone from the start; document the change and its effective date explicitly.
Frequently asked questions
What is the difference between inclusion and exclusion criteria?
Inclusion criteria state the characteristics a unit must have to be eligible; exclusion
criteria identify specific, otherwise-eligible units that should still be removed. They define one
population boundary from two directions, not two separate populations.
How many inclusion and exclusion criteria should a study have?
There is no fixed number — the right count is however many are needed to define the population
the research question is actually about, and no more. Every additional criterion should pass the
justification test in the worksheet above; a long list of criteria that each individually make
sense can still, in combination, define a population too narrow to be useful.
Can a criterion be listed as both an inclusion and an exclusion criterion?
It shouldn’t be. A criterion belongs on one list; restating the same rule on both lists (as a
positive inclusion requirement and again as its negative exclusion equivalent) is redundant and
usually a sign the list wasn’t reviewed as a whole before finalizing.
Where do inclusion and exclusion criteria go in a study protocol or manuscript?
In the methods section, as a distinct, clearly labeled subsection, stated specifically enough
that another researcher could apply the same criteria to a candidate case and reach the same
eligibility decision. Reporting guidelines including CONSORT and STROBE treat this as a required,
checked reporting item, not an optional detail.
What happens if a participant stops meeting the criteria after enrollment?
This should be decided and documented in the protocol before enrollment starts, not improvised
case by case. Common approaches are: retain and analyze as enrolled (an “intention-to-treat”-style
approach in intervention studies), or document a formal, pre-specified withdrawal/discontinuation
rule. What should not happen is a silent, undocumented removal of the case from the dataset after
the fact.







