Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Thematic Analysis: A Step-by-Step Guide to Braun and Clarke’s Six Phases

How thematic analysis works: the six-phase Braun and Clarke framework, the three schools of thematic analysis (coding reliability, codebook, reflexive), inductive vs. deductive coding, and how to report it in a methods section.

Thematic analysis is a method for identifying, analyzing, and reporting patterns (themes) within qualitative data — interview transcripts, focus-group discussions, open-ended survey responses, field notes, or documents. It is one of the most widely used qualitative analysis methods across psychology, health research, education, social science, and increasingly business and UX research, largely because it is flexible enough to work with almost any qualitative dataset and any theoretical framework, rather than being tied to a single epistemological position the way grounded theory or phenomenology are.

The method’s dominant modern reference point is Virginia Braun and Victoria Clarke’s 2006 paper “Using Thematic Analysis in Psychology” (Qualitative Research in Psychology, 3(2), 77–101), which set out a six-phase process that has since become the standard citation across disciplines far beyond psychology. Braun and Clarke have since revised and clarified their own approach substantially — most notably in a 2019 paper distinguishing their method, which they now call reflexive thematic analysis, from other approaches that also go by the name “thematic analysis.” Understanding that distinction matters more than most guides to the topic acknowledge, because conflating the different schools is one of the most common methodological errors reviewers flag in qualitative manuscripts.

What Counts as a “Theme”?

A theme in thematic analysis is not simply a topic that came up repeatedly, or a summary of everything participants said about a subject. Braun and Clarke are explicit that a theme should capture a coherent and meaningful pattern relative to the research question — something the researcher has actively constructed through interpretation, not passively discovered lying in the data waiting to be counted. This is the single most common conceptual error in student and early-career thematic analysis: treating themes as pre-existing labels for domains of content (“participants talked about family,” “participants talked about work”) rather than as analytic claims about what is going on across the dataset in relation to the question being asked.

When to Use Thematic Analysis (and When Not To)

Thematic analysis is a good fit when the goal is to understand patterns of meaning across a dataset without committing to a specific theoretical apparatus for generating that understanding. It is less suited to research questions that require a different, more specialized analytic logic:

  • Grounded theory is the better choice when the goal is to generate a novel theoretical model or explanatory framework from the data itself, with its own iterative sampling and theoretical-saturation logic. See CASRAI’s grounded theory entry for how the two methods differ in aim and process.
  • Discourse analysis is the better choice when the research question is about how language itself constructs meaning, identity, or power relations — the focus is on language-in-use, not on identifying patterns of meaning across a dataset. See CASRAI’s guide on how to conduct a discourse analysis.
  • Narrative analysis is the better choice when the unit of analysis is the story itself — its structure, sequence, and how it is told — rather than themes extracted across multiple participants’ accounts. See CASRAI’s guide on narrative analysis as a qualitative research method.
  • Content analysis (particularly quantitative content analysis) is the better choice when the goal is to systematically count the frequency of pre-defined categories across a large body of text, rather than to interpret patterns of meaning.

Thematic analysis is compatible with almost any research paradigm — positivist, interpretivist, or critical — because it is a method, not a methodology tied to a single set of ontological/epistemological commitments. See CASRAI’s research paradigm guide for how that distinction plays out in a methodology section.

The Three Schools of Thematic Analysis

Not all thematic analysis is the same analytic approach, even though many published studies use the label loosely. Braun and Clarke’s own later work (2019, 2021) distinguishes three broad approaches, and naming the one actually used is now considered good methodological practice in a manuscript’s methods section:

  • Coding reliability thematic analysis (associated with Richard Boyatzis’s 1998 approach) treats themes as fixed, pre-defined units, uses a structured codebook applied by multiple coders, and reports inter-rater reliability statistics (such as Cohen’s kappa) as evidence of coding consistency. It assumes there is a single correct coding of the data that independent coders should converge on.
  • Codebook thematic analysis (including approaches such as framework analysis and template analysis) also uses a structured, often matrix-based codebook developed early in the process, but treats coding as more interpretive than the reliability approach — codebook development can be iterative, and it doesn’t necessarily require multiple independent coders reaching statistical agreement.
  • Reflexive thematic analysis — Braun and Clarke’s own current approach — explicitly rejects the idea that coding should be checked for reliability between independent coders, on the grounds that qualitative coding is inherently a subjective, interpretive act shaped by the researcher’s theoretical position and analytic engagement with the data. Instead of inter-rater reliability, reflexive TA relies on researcher reflexivity, depth of engagement with the data, and the coherence of the resulting analysis as its quality markers.

These are not interchangeable labels for the same procedure. A manuscript that reports “thematic analysis” without specifying which of the three it used, or that reports inter-rater reliability scores while citing Braun and Clarke’s reflexive approach (which explicitly disavows that quality criterion), is presenting an internally inconsistent methods section — a pattern peer reviewers in qualitative research increasingly flag.

Braun and Clarke’s Six-Phase Framework

The original 2006 framework, still the most commonly cited structure regardless of which of the three schools a researcher ultimately follows, sets out six phases. Braun and Clarke are explicit that the process is recursive, not strictly linear — researchers move back and forth between phases rather than completing each one once and moving on permanently.

  1. Familiarization with the data. Repeated, active reading of the full dataset (transcribing recordings personally, where feasible, is itself considered part of familiarization, not just a preparatory step) while noting initial ideas.
  2. Generating initial codes. Systematically coding interesting features of the data across the entire dataset, collating data extracts relevant to each code. Coding at this stage should be comprehensive and inclusive rather than pre-filtered to what looks obviously relevant.
  3. Searching for themes. Collating codes into candidate themes, gathering all data extracts relevant to each candidate theme. This is where a mass of codes starts to be organized into a smaller number of broader patterns.
  4. Reviewing themes. Checking candidate themes against the coded extracts (do the extracts actually support the theme, and is the theme internally coherent?) and against the entire dataset (does the developing thematic map accurately represent the whole dataset, not just the extracts already coded?). Themes are commonly split, merged, or discarded at this stage.
  5. Defining and naming themes. Refining the specifics of each theme, identifying the “essence” of what each theme is about, and generating clear, concise names and definitions for each.
  6. Producing the report. Selecting vivid, compelling data extracts, relating the analysis back to the research question and existing literature, and writing up a coherent account — not just a list of themes with illustrative quotes attached.

Inductive vs. Deductive Thematic Analysis

Thematic analysis can be conducted inductively (data-driven, where codes and themes emerge from the content of the data itself, with minimal reference to pre-existing theory) or deductively (theory-driven, where the analysis is guided by an existing theoretical framework, coding the data against categories derived from that framework). Many real analyses combine both: a largely inductive first pass supplemented by deductive coding against a specific theoretical concept the study is testing or extending. Whichever approach is used, it should be stated explicitly in the methods section, since it materially affects how the resulting themes should be interpreted.

Semantic vs. Latent Themes

A related distinction is between semantic (explicit) and latent (interpretive) themes. Semantic thematic analysis identifies themes within the explicit, surface meaning of what participants said, without looking for anything beyond what was directly stated. Latent thematic analysis goes further, examining the underlying ideas, assumptions, and ideologies that are theorized to shape or inform the semantic content — this level of analysis typically involves more interpretive work and is closer to a constructionist than a purely descriptive analytic stance. As with the inductive/deductive choice, most published studies sit somewhere on a spectrum rather than at a pure endpoint, and stating where on that spectrum the analysis sits is part of a transparent methods section.

Software for Thematic Analysis

Qualitative data analysis software does not perform thematic analysis automatically — the interpretive work of generating and reviewing themes remains the researcher’s, not the software’s — but it substantially eases the mechanics of coding a large dataset: tagging extracts, retrieving all data linked to a given code, building and reorganizing a codebook, and keeping an audit trail of how codes evolved into themes. The most commonly used tools are NVivo, ATLAS.ti, MAXQDA, and the lighter-weight Delve, alongside manual approaches (color-coding printed transcripts, spreadsheet-based coding) that remain entirely valid for smaller datasets. See CASRAI’s NVivo guide for a detailed look at how one of the major platforms supports the coding-to-theme workflow described above.

Common Mistakes in Thematic Analysis

  • Themes as topic summaries, not analytic claims. The single most common issue: presenting “domain summaries” (everything said about a given subject) as though they were themes, rather than identifying a coherent pattern of meaning across the data.
  • Not specifying which school of thematic analysis was used. Citing Braun and Clarke’s six phases while also reporting inter-rater reliability statistics mixes reflexive TA with coding reliability TA in a way the two approaches’ own developers treat as methodologically incompatible.
  • Coding too narrowly, too early. Filtering out data that doesn’t look immediately “interesting” during initial coding, rather than coding inclusively across the full dataset and letting theme development do the narrowing later.
  • Treating theme prevalence as the main quality marker. A theme’s value comes from what it captures about the research question, not from how many participants mentioned it or how many times it was coded — a theme voiced by a minority of participants can still be analytically important.
  • Insufficient reviewing. Skipping phase four (reviewing themes against both the coded extracts and the full dataset) and moving straight from initial candidate themes to a final write-up, which risks a thematic map that doesn’t actually represent the dataset as a whole.

Reporting Thematic Analysis in a Manuscript

A methods section reporting thematic analysis should specify, at minimum: which of the three schools (coding reliability, codebook, or reflexive) was followed and why; whether coding was inductive, deductive, or a combination; whether the analysis operated at a semantic or latent level; who conducted the coding (a single researcher, or multiple coders, and if multiple, how disagreements were resolved — noting that this question does not apply the same way under a reflexive-TA approach, which does not treat coder disagreement as an error to be reconciled); and how many phases of Braun and Clarke’s framework, or an equivalent process, were followed. Reporting a bare citation to Braun and Clarke (2006) with no further detail is increasingly treated by reviewers as under-specified, given how much the field’s own understanding of what “thematic analysis” means has diversified since that original paper.

Frequently Asked Questions

Is thematic analysis qualitative or quantitative?

Thematic analysis is a qualitative method. It works with non-numerical data (text, transcripts, open-ended responses) and produces an interpretive account of patterns of meaning, not statistical output. Some coding-reliability variants do report reliability statistics (like Cohen’s kappa) as a supplementary quality check, but the core analytic work stays qualitative.

What’s the difference between thematic analysis and content analysis?

Content analysis, especially in its quantitative form, typically involves systematically counting the frequency of pre-defined categories across text. Thematic analysis is more interpretive: themes are constructed through engagement with the data’s meaning rather than counted from a pre-set coding frame, though codebook thematic analysis sits closer to content analysis on this spectrum than reflexive thematic analysis does.

How many themes should a thematic analysis produce?

There is no fixed number. Braun and Clarke explicitly caution against treating theme count as a quality marker; the right number is whatever set of coherent, non-overlapping patterns best answers the research question for that specific dataset, typically resulting in a handful of major themes (often with sub-themes), not dozens of loosely related codes reported as themes.

Do you need special software to do thematic analysis?

No. Software such as NVivo, ATLAS.ti, MAXQDA, or Delve makes coding a large dataset more manageable and auditable, but thematic analysis was originally developed and is still routinely conducted with manual methods (highlighting transcripts, index cards, spreadsheets) for smaller datasets.

What sample size does thematic analysis need?

There is no universal minimum, and Braun and Clarke have explicitly pushed back against fixed sample-size rules for thematic analysis (including against blanket application of “saturation” as a stopping criterion, particularly for reflexive TA). The right size depends on the richness and focus of the research question, the depth of each data item, and the scope of the analytic claim being made, and should be justified on those grounds rather than by an arbitrary number.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →