Skip to main content
v2026.11,610 entries · CC-BY 4.0

Coding Qualitative Interview Data: A Worked-Example Guide

A practical, example-driven guide to coding qualitative interview transcripts – inductive and deductive approaches, first- and second-cycle coding methods, codebook development, inter-coder reliability, saturation, and ethics.

Ask about Coding Qualitative Interview Data: A Worked-Example Guide

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Coding is the foundational analytic step in most qualitative research: it is the process of attaching short labels — codes — to segments of data (a phrase, a sentence, a paragraph of a transcript) so that similar ideas can be found, compared, and grouped across an entire dataset. A code is not a finding and it is not yet a theme. It is a retrieval tag, applied close to the data, that makes the later analytic work of building categories and themes possible. Confusing a code with a theme — treating a label like “time pressure” as if it were already an analytic claim about the dataset — is one of the most common weaknesses reviewers flag in qualitative manuscripts.

This guide focuses on the mechanics of coding itself, with worked examples. For the surrounding steps, see CASRAI’s guides on designing semi-structured interviews (how the data you are about to code gets collected), note-taking methods for research (capturing observations and analytic memos as you go), and thematic analysis (what happens after coding, when codes are developed into themes). If your study builds theory from the ground up rather than testing a prior framework, see the dictionary entry on grounded theory, where coding plays an even more central, iterative role.

Coding Approaches: Inductive, Deductive, and Hybrid

Before opening a transcript, a researcher has to decide where the codes are going to come from. There are three broad approaches, and most real studies use some blend of the first two.

  • Inductive (open) coding — codes emerge from the data itself. The researcher reads the transcript with few or no preconceptions about what the codes will be and names segments based on what is actually said. This is the default starting point in grounded theory and in reflexive approaches to thematic analysis, and it is the better choice when a topic is under-theorized or when the research question is genuinely exploratory.
  • Deductive (a priori) coding — codes are defined before coding begins, drawn from an existing theoretical framework, a conceptual model, or the structure of the interview guide itself (each interview question or probe becomes a candidate code). This is faster, produces more consistent results across coders, and is the natural choice when a study is explicitly testing or applying an existing framework — but it risks missing anything the framework did not anticipate.
  • Hybrid (a priori plus emergent) coding — the most common approach in applied research: the researcher starts with a small set of deductive codes tied to the research questions, then adds inductive codes as unanticipated content appears in the data. The codebook is treated as a living document rather than a fixed instrument from the first transcript.

The choice connects directly to a study’s underlying reasoning: see CASRAI’s guide to inductive vs. deductive vs. abductive reasoning for how this decision fits into the broader logic of a research design.

First-Cycle and Second-Cycle Coding Methods

Johnny Saldaña’s The Coding Manual for Qualitative Researchers is the standard reference for organizing the many named coding methods into two broad “cycles,” and the terminology below follows that framing, which is now widely used across disciplines regardless of whether a study cites Saldaña directly.

First-cycle methods are applied directly to the raw data, generating the initial pool of codes:

  • Open (initial) coding — breaking the data into discrete segments and assigning a provisional label to each, without trying to fit segments into a predetermined structure.
  • In-vivo coding — using the participant’s own words or short phrases as the code itself, preserving their language rather than the researcher’s paraphrase.
  • Descriptive coding — summarizing the topic of a segment in a short noun phrase (e.g., “onboarding delay”), useful for cataloguing a large, topically varied dataset.
  • Process coding — using gerunds (“-ing” words) to code observable or conceptual action (e.g., “negotiating access,” “delaying disclosure”).
  • Values coding — capturing a participant’s values, attitudes, and beliefs about a topic.
  • Emotion coding — labelling the emotions a participant reports or expresses.
  • Versus coding — capturing binary conflicts or tensions participants describe (e.g., “autonomy vs. oversight”).

Second-cycle methods are applied to the first-cycle codes themselves, not the raw data, and are how a large, flat list of codes gets organized into structure:

  • Pattern coding — grouping first-cycle codes into a smaller number of categories or explanatory constructs.
  • Focused coding — a term from Charmaz’s constructivist grounded theory: identifying the most frequent or analytically significant first-cycle codes and using them to sift through the remaining data.
  • Axial coding — a term from Strauss and Corbin’s approach to grounded theory: reassembling data along the properties and dimensions of a category, and specifying relationships (conditions, context, consequences) between categories.
  • Theoretical coding — a later-stage grounded theory method for specifying how the categories identified through focused/axial coding relate to one another as an integrated theory.

Not every study needs all of these. Most projects use one or two first-cycle methods suited to the research question, then move to pattern or focused coding to build categories, and only projects explicitly building theory (rather than describing a phenomenon) need axial or theoretical coding.

Worked Examples: Coding Two Contrasting Excerpts

The excerpts below are illustrative composites written for this guide. They are not transcripts from a real study, and they are not attributed to any real participant, institution, or dataset. They exist to show the mechanics of coding — how a raw line of transcript becomes a code, how codes cluster into a category, and how categories build toward a theme — not to report a real finding.

Excerpt 1: Inductive (Open) Coding

Context: a hypothetical interview exploring early-career researchers’ experiences balancing research and administrative workload. No interview guide codes are applied in advance — codes are generated purely from what is said.

“Honestly, the grant reporting takes up more of my week than the actual experiments do right now. [1] I sat down last Tuesday to finish an analysis I’d been looking forward to for weeks, and by eleven I was still filling in a budget justification instead. [2] Nobody tells you in your PhD that this is what the job actually looks like. [3] I don’t resent the reporting itself, I get why funders need it — I resent that nobody protected the time for it, so it just eats into evenings and weekends. [4]

Segment Code applied Coding reasoning
[1] admin-time-exceeds-research-time Descriptive code capturing the specific, comparable claim being made (reporting > experiments), not a broader label like “workload,” which would lose the comparison.
[2] interrupted-by-admin (process coding) The gerund form captures the action — being pulled off analysis mid-task — rather than just the topic, which matters because this is a pattern, not a one-off complaint.
[3] “nobody tells you” (in-vivo) Kept in the participant’s own words because the phrase itself signals a recurring frustration — a gap between training and the realities of the role — that a paraphrase would flatten.
[4] unprotected-time (values coding) Distinguishes the participant’s stance (reporting is legitimate) from their actual complaint (no protected time for it) — an important distinction to preserve for later analysis, not collapse into a single “frustration with reporting” code.

From codes to category: across multiple similar transcripts, admin-time-exceeds-research-time, interrupted-by-admin, and unprotected-time would likely cluster into a category such as “administrative burden displacing research time,” distinct from a separate category about training/preparedness that the in-vivo code “nobody tells you” might feed into instead. Only once that category is checked against enough of the dataset, and related to other categories (e.g., how it connects to reported burnout or turnover intentions), does it become candidate material for a theme — see CASRAI’s guide to thematic analysis for how that later step works.

Excerpt 2: Deductive (A Priori) Coding

Context: a hypothetical interview guide built around a study of researchers’ willingness to share raw data. The interview guide already defines codes tied to a theoretical framework of data-sharing barriers: trust-barrier, incentive-barrier, technical-barrier, and norms-barrier. The coder applies these predefined codes to new transcript text, in the same style as a semi-structured interview built around a fixed probe structure.

“I’d share the raw dataset tomorrow if I trusted people not to reanalyze it badly and publish something misleading with my name still attached to the original collection. [1] There’s also just no credit for it — my tenure file doesn’t have a line for ‘made data available,’ it has a line for papers. [2] And practically, our lab doesn’t have anyone who knows how to structure it properly for a repository, so even if I wanted to, it would take a week I don’t have. [3]

Segment Code applied Coding reasoning
[1] trust-barrier Matches the a priori definition directly: concern about downstream misuse of the data, not a resourcing or incentive issue.
[2] incentive-barrier Matches the predefined code for lack of formal credit or career recognition for data sharing.
[3] technical-barrier Matches the predefined code for lack of skills/capacity to prepare data for deposit — note this is coded separately from the incentive issue in [2] even though both appear in the same sentence, because they are analytically distinct barriers that a funder or repository would need to address differently.

From codes to category to theme: because the codes here were defined by the theoretical framework before coding began, the “category” step is mostly already done — the framework specified that trust-barrier, incentive-barrier, and technical-barrier are all instances of the higher-order construct “barriers to data sharing.” The analytic work that remains is deductive-hybrid: does the data reveal a barrier the framework did not anticipate (which would need a new, inductively generated code), and which barriers recur most strongly or interact with each other across the dataset — that pattern, not any single code, is what becomes a reportable theme.

Read side by side, the two excerpts show why the choice of approach matters: inductive coding in Excerpt 1 let a specific, unanticipated complaint (interrupted analysis time) surface in the participant’s own terms; deductive coding in Excerpt 2 efficiently sorted a dense, multi-barrier answer into a pre-validated framework but would have missed a genuinely new barrier if one had been mentioned. A hybrid strategy — a priori codes for the framework, with an “other/emergent” code kept open for anything that doesn’t fit — is how most applied qualitative studies get the benefit of both.

Building and Maintaining a Codebook

A codebook is the auditable record of what each code means and how it should be applied. At minimum, each entry needs:

  • Code name — short, consistent, ideally not ambiguous with another code in the set.
  • Definition — a plain-language description of what the code refers to.
  • Inclusion criteria — what kind of statement qualifies for this code.
  • Exclusion criteria — what looks similar but should NOT get this code (this is what actually prevents two coders from drifting apart, and it’s the field most codebooks skip until a disagreement forces them to write it).
  • Exemplar quote — a real (or, if illustrative, clearly labelled) example of text that earns the code, so a second coder has a concrete anchor, not just an abstract definition.

Codebooks are iterative, not fixed at the outset, particularly under an inductive or hybrid approach: new codes get added, overlapping codes get merged, and definitions get sharpened as disagreements surface. Because the codebook itself is part of the analytic record, most rigorous studies version it — keeping dated copies (or a version-controlled file) so that a reviewer or a second coder can see exactly which definition was in force when a given transcript was coded, and so the researcher can trace when and why a code’s meaning changed partway through analysis.

Multiple Coders and Inter-Coder Reliability

When more than one person codes the same data, two quite different methodological traditions offer different answers to “how do we know the coding is trustworthy?” — and a well-written methods section states which one it is following and why, rather than treating the choice as self-evident.

The reliability-statistics position, more common in coding-reliability approaches to content and thematic analysis and in mixed-methods/health-services research, treats independent double-coding and a quantitative agreement statistic as the evidence of trustworthy coding. Two (or more) coders code the same subset of transcripts independently, without discussing their decisions in advance, and agreement is calculated statistically:

  • Cohen’s kappa — the most commonly reported statistic for agreement between exactly two coders, correcting raw percent agreement for the agreement expected by chance.
  • Krippendorff’s alpha — a more flexible alternative that handles more than two coders, missing data, and different levels of measurement, which is why it is increasingly preferred in larger coding teams.

Interpretation guidance varies by field and by the specific statistic, but published coding-reliability studies commonly treat values in the region of roughly 0.61-0.80 as “substantial” agreement and above roughly 0.80 as “near-perfect,” with anything meaningfully below that prompting codebook revision rather than simply being reported as-is.

The consensus/reflexive position, more common in reflexive thematic analysis, constructivist grounded theory, and other interpretivist traditions, explicitly rejects treating a kappa or alpha score as evidence of coding quality. The argument, most closely associated with Braun and Clarke’s writing on reflexive thematic analysis, is epistemological rather than practical: if meaning in qualitative data is understood as co-constructed by the researcher rather than as a fixed, objectively “findable” fact waiting to be measured, then two coders reaching the identical label is not actually the goal — a single, well-positioned, reflexive researcher’s interpretation, developed and checked through discussion, memoing, and engagement with the data over time, is treated as the more meaningful form of rigor than statistical agreement between coders working in isolation. In this tradition, multiple coders (where used at all) work toward consensus coding — discussing disagreements to reach a shared, negotiated interpretation — rather than measuring how often they agreed before discussing anything.

Neither position is simply “more rigorous” than the other; they answer to different underlying assumptions about what qualitative data is and what coding is meant to establish. The methodological error is not picking one — it is failing to state which one a study is using, or reporting a kappa score inside a study that has otherwise adopted an interpretivist, reflexive framework where that statistic doesn’t logically fit.

Recognizing and Reporting Saturation

Saturation is the point at which coding additional data stops producing new codes or meaningfully revising existing ones — the codebook has stabilized and further transcripts are confirming, not extending, the existing structure. In practice, saturation should be judged and reported specifically, not asserted as a blanket claim:

  • State what kind of saturation is being claimed — code saturation (no new codes appearing) is a lower bar than meaning saturation (no new nuance within existing codes/themes), and reviewers increasingly expect the distinction to be made explicit.
  • Report roughly how many transcripts were coded before new codes stopped appearing, and ideally show the trajectory (e.g., codes added per transcript, tapering toward zero) rather than a single end-point number.
  • Be honest when saturation was not fully reached — because of a fixed recruitment budget, participant access constraints, or a deliberately small, purposive sample — and say so, rather than asserting saturation as a formality. A stated, well-justified limitation is more credible to a reviewer than an unsupported saturation claim.

Software for Coding Qualitative Data

Qualitative data analysis software organizes coding — it does not perform the analysis. All of the tools below let a researcher attach codes to segments of text (and often audio, video, or image data), retrieve every segment tagged with a given code, visualize code co-occurrence, and manage a codebook with version history. None of them decide what a segment means or whether two codes should merge; that judgment remains the researcher’s.

  • NVivo — widely used commercial QDA software with strong support for mixed-methods projects and large datasets.
  • ATLAS.ti — commercial QDA software with a strong network/relationship-visualization feature set, popular in grounded theory and network-analysis-oriented projects.
  • MAXQDA — commercial QDA software with particularly strong mixed-methods and quantitative-content-analysis integration alongside standard coding tools.
  • Dedoose — browser-based, subscription-priced QDA tool built for collaborative, multi-coder teams working across locations.
  • Taguette — free, open-source coding software aimed at researchers and students who need core tag-and-retrieve functionality without a commercial license.

Rigor and the Audit Trail

Because qualitative coding involves interpretive judgment, the standard for rigor is not “was the coding objective” but “is the analytic process visible and traceable” — an audit trail a second reader could follow from raw data to final theme. The core practices:

  • Memoing — writing short analytic notes at the moment a coding decision is made (why this segment got this code, what it reminds the coder of, a tension with an earlier decision) rather than only after the fact. Memos, not the codes alone, are usually where the actual analytic thinking is preserved. See CASRAI’s guide to note-taking methods for structured approaches to capturing this kind of running analytic note.
  • Reflexivity — explicitly documenting how the researcher’s own position, assumptions, and prior expectations may have shaped which codes were noticed and how they were defined, rather than presenting the coder as a neutral instrument.
  • Negative-case analysis — deliberately searching for and reporting data that contradicts an emerging code or theme, rather than only retaining confirming examples. A theme that survives active testing against disconfirming cases is far more credible than one built solely from supporting quotes.
  • A decision log — a running record of codebook changes (codes added, merged, split, redefined) with dates and rationale, distinct from the memos themselves, so the evolution of the analysis can be reconstructed later. This plays a similar role to triangulation in strengthening the credibility of a qualitative finding, though the two are not the same thing — triangulation compares across data sources or methods, while an audit trail documents the analytic process within a single method.

Ethics: Anonymizing Transcripts and Quotes

Interview data is identifiable data even when a participant is not named, and coding is typically the stage at which a researcher is reading transcripts most closely — making it the natural point to apply anonymization decisions consistently, not the point to defer them. In practice this means:

  • Removing or replacing direct identifiers (names, employers, precise locations, dates that could be triangulated with other public information) from the working transcript before it is widely circulated among a coding team, consistent with what was promised in the study’s informed consent process.
  • Checking that quotes selected as codebook exemplars — and any quotes eventually used in a manuscript — do not themselves make a participant identifiable through distinctive phrasing, an unusual role, or a rare combination of details, even if no name is attached.
  • Confirming what participants actually consented to regarding direct quotation, since consent to be interviewed does not automatically include consent to be quoted verbatim in a publication.
  • Storing transcripts and the coded dataset securely, with access limited to the coding team, for as long as the data governance plan specifies. See CASRAI’s guide to data anonymisation in research for the underlying techniques and standards, and the comparison of anonymization vs. pseudonymization for how the two differ in practice.

Frequently Asked Questions

What is the difference between a code and a theme?

A code is a short label applied to a specific segment of data during initial analysis; a theme is a broader, patterned meaning built by grouping and interpreting multiple codes across the dataset, usually after an intermediate step of grouping codes into categories. A code describes; a theme claims something analytically. See CASRAI’s guide to thematic analysis for the full process from codes to themes.

Do I need more than one coder?

Not necessarily. A single, reflexive researcher coding with a clear audit trail (memos, a decision log, negative-case analysis) is an accepted and, in some traditions, preferred approach. Multiple coders are more common where a study needs to demonstrate coding consistency to a specific audience (e.g., a funder, a clinical or mixed-methods context) or where the dataset is large enough that splitting the coding load across a team is practically necessary.

What counts as an acceptable inter-coder reliability score?

There is no single universal threshold, and the answer depends first on whether a study’s methodological tradition treats reliability statistics as appropriate at all (see the “Multiple Coders” section above). Where a kappa or alpha statistic is reported, values in roughly the 0.61-0.80 range are commonly described as “substantial” agreement in the literature, with anything notably lower typically prompting the coders to revise the codebook and re-code rather than report a low score as final.

How many transcripts do I need before I can stop coding?

There is no fixed number; the honest answer is “until saturation is reached and justified,” and that point varies with how homogeneous the sample is, how narrow the research question is, and how deep each interview goes. Report the trajectory of new codes per transcript rather than asserting a round number as a target set in advance.

Should I code by hand or use software?

Either is methodologically valid for smaller datasets; software becomes practically necessary once a dataset is large enough that manual retrieval of every segment carrying a given code is unmanageable, or where a team needs to code collaboratively and compare coding in one place. Software organizes and retrieves; it does not interpret the data for the researcher.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →