Written and maintained by CASRAI Editorial Board
Last updated
Quantitative research has an established toolkit for judging rigor: internal validity, external validity, reliability, and objectivity. Naturalistic and qualitative inquiry does not map cleanly onto those criteria, because it does not share the same assumptions about a single, observer-independent reality. Lincoln and Guba addressed that gap directly in Naturalistic Inquiry (1985), proposing four parallel criteria for evaluating the trustworthiness of qualitative findings: credibility, transferability, dependability, and confirmability. Each criterion has a specific, well-documented set of techniques that satisfy it — this guide maps all four, with a checklist you can use directly in a methods section.
The four criteria and their quantitative parallels
Lincoln and Guba built each criterion as a direct answer to a specific quantitative concern, which is why reviewers and committee members familiar with quantitative rigor tend to recognize the logic quickly once it’s named.
| Qualitative criterion | Quantitative parallel | Core question it answers |
|---|---|---|
| Credibility | Internal validity | Do the findings represent a credible interpretation of the data, one the participants themselves would recognize? |
| Transferability | External validity / generalizability | Can a reader judge whether these findings apply to their own context? |
| Dependability | Reliability | Would the process, if repeated with the same participants and context, be recognizable and consistent? |
| Confirmability | Objectivity | Are the findings shaped by the data and participants, not by the researcher’s unexamined assumptions? |
Credibility: techniques that satisfy it
Credibility is the criterion reviewers scrutinize most, since it stands in for internal validity. No single technique establishes it — Lincoln and Guba’s own recommendation is to combine several.
- Prolonged engagement. Spending enough time in the field to learn the culture, build trust with participants, and get past initial distortions in what people say or do because a researcher is present.
- Persistent observation. Going deep rather than broad on the specific features of the phenomenon that matter to the research question, once prolonged engagement has established what those features are.
- Triangulation. Cross-checking findings across multiple data sources, methods, investigators, or theoretical lenses. See Triangulation in Research for the four types (data, investigator, method, theory) and how to report which one you used.
- Peer debriefing. Working through the analysis with a disinterested colleague who probes for taken-for-granted interpretations, hidden assumptions, and gaps in the logic connecting data to claims.
- Negative case analysis. Actively searching for and accounting for data that don’t fit the emerging pattern, then revising the working hypothesis until it accounts for the exceptions rather than discarding them.
- Referential adequacy. Setting aside a subset of raw data (recordings, documents) unanalyzed, then checking preliminary findings against it later as an independent check.
- Member checking. Taking findings, interpretations, or transcripts back to participants to confirm they recognize the account as accurate. This is usually the single technique reviewers ask about first. See Member Checking in Qualitative Research for the variants (transcript review vs. synthesized-findings review) and a usable protocol.
Transferability: techniques that satisfy it
Transferability is not the researcher’s job to prove in the way generalizability is proven statistically — it’s the researcher’s job to provide enough information that a reader can judge transferability to their own setting.
- Thick description. Detailed, contextualized description of the setting, participants, and circumstances — not just what happened, but the context that gives it meaning, so a reader in a different setting can assess the fit. See Ethnographic Writing Conventions for what distinguishes thick description from a thin summary in practice.
- Purposive sampling documentation. Stating explicitly why participants and sites were selected, and what range of variation they represent, so a reader can judge how far the findings might extend.
- Boundary conditions. Being explicit about what the study does not claim to cover — the settings, populations, or time periods the findings should not be assumed to transfer to.
Dependability: techniques that satisfy it
Dependability accepts that conditions in naturalistic settings change; the standard isn’t identical replication, it’s that the research process is documented well enough that the trajectory of decisions is traceable and defensible.
- Audit trail. A documented record of methodological decisions and their rationale: raw data, analysis notes, process notes (design and methodological decisions), and materials relating to intentions and dispositions (reflexive notes, the evolving research plan). A reader or examiner should be able to follow how a specific finding traces back to specific data.
- Dependability audit. An external reviewer examines the process and the audit trail (not just the product) to assess whether the procedures used were consistent with accepted qualitative practice and could reasonably produce the findings reported.
- Code-recode procedure. Coding the same dataset at two points in time, then comparing the two coding passes to check consistency in how categories are being applied.
Confirmability: techniques that satisfy it
Confirmability shifts the objectivity question from “is the researcher neutral?” (usually an impossible standard in interpretive work) to “can the findings be traced to the data and participants, rather than to the researcher’s own predispositions?”
- Confirmability audit. Using the same audit trail built for dependability, an external reviewer traces specific findings and interpretations back through the analysis to the raw data, checking that conclusions are grounded rather than invented.
- Reflexivity. An ongoing, documented account of the researcher’s own position, assumptions, and role in shaping the data and its interpretation — usually kept as a reflexive journal alongside fieldwork and analysis. See Reflexivity and Positionality Statements for a template with worked examples.
- Triangulation. The same cross-checking used for credibility also supports confirmability: findings that hold up across multiple sources or methods are less plausibly an artifact of one researcher’s viewpoint.
Practical checklist: what to build, and what it lets a reader verify
| Criterion | Minimum technique(s) to include | What it lets a reader verify |
|---|---|---|
| Credibility | Member checking or peer debriefing, plus triangulation | The interpretation is recognizable to participants and holds up under independent scrutiny. |
| Transferability | Thick description of setting, participants, and sampling rationale | Enough context to judge fit with a different setting. |
| Dependability | Audit trail (raw data, process notes, decision log) | The path from data to finding is traceable, not asserted. |
| Confirmability | Reflexive journal plus audit trail | Findings are grounded in data and participants, not the researcher’s unexamined position. |
A methods section that names all four criteria and states, for each, exactly which technique(s) were used and where the evidence lives (an appendix, a supplementary file, an available-on-request audit trail) is doing real trustworthiness work. Naming the criteria without naming the technique that satisfies each is the most common gap reviewers flag.
Common mistakes
- Treating the four criteria as a checklist to name, not techniques to actually do. Writing “credibility was established through triangulation” without specifying which type of triangulation, or what was actually cross-checked against what, reads as a formula rather than evidence.
- Using member checking as a rubber stamp. Sending a full transcript back with no structured question tends to get a generic “looks fine” response. A synthesized-findings review with specific prompts produces a more genuine credibility check — see the member checking guide linked above for the distinction.
- Conflating dependability and confirmability. Both rely on the audit trail, but they answer different questions: dependability asks whether the process was consistent and defensible; confirmability asks whether the findings are traceable to data rather than to the researcher. Address both explicitly rather than treating one audit-trail mention as covering both.
- Skipping boundary conditions. A transferability discussion that only describes the study setting, without stating what it does not claim to generalize to, leaves the reader to guess — and guessing wrong is what transferability is meant to prevent.
Frequently asked questions
Are credibility, transferability, dependability, and confirmability the same as validity and reliability?
They are parallel, not identical. Lincoln and Guba designed the four criteria specifically because qualitative and quantitative research rest on different assumptions about reality and knowledge, so the quantitative terms (internal/external validity, reliability, objectivity) don’t transfer directly. The four qualitative criteria answer analogous questions using methods appropriate to interpretive, naturalistic inquiry rather than statistical ones.
Do I need to use every technique under every criterion?
No. Reviewers generally expect at least one well-executed, clearly reported technique per criterion, not an exhaustive list. Combining two or three credibility techniques (for example, triangulation plus member checking) is common because credibility gets the most scrutiny, but a single technique per criterion, done rigorously and reported specifically, is defensible.
What is an audit trail, concretely?
It’s the documentation that lets someone else follow your process: raw data files, coding decisions and their rationale, memos, changes to the interview guide or sampling plan and why they were made, and a record of how themes or categories evolved. It supports both dependability (was the process sound?) and confirmability (do findings trace back to data?).
Where did these four criteria come from?
Lincoln, Y. S., and Guba, E. G., Naturalistic Inquiry (Sage Publications, 1985). The framework remains the most widely cited trustworthiness model in qualitative methods texts and is commonly taught alongside later refinements, including Guba and Lincoln’s own subsequent work adding “authenticity” criteria in evaluation contexts.
Is thick description just a longer methods section?
No — length isn’t the point. Thick description is contextualized detail that conveys meaning, not just an inventory of events. A thin description states what happened; a thick description gives a reader enough of the surrounding context and significance to interpret it themselves. See the ethnographic writing conventions guide linked above for concrete examples of the distinction.








