Skip to main content
v2026.11,610 entries · CC-BY 4.0

Content Analysis: Coding Text and Media Systematically

A practical guide to content analysis: the manifest/latent distinction, quantitative vs. qualitative variants, unitizing, building and piloting a coding frame, and measuring inter-coder reliability with Cohen’s kappa and Krippendorff’s alpha.

Ask about Content Analysis: Coding Text and Media Systematically

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Content analysis is a research method for systematically classifying the content of text, images, or other recorded communication into categories, so that patterns can be counted, compared, and interpreted. It sits between quantitative and qualitative traditions: it applies a rule-based coding system to unstructured material, but the categories themselves — and the judgment calls coders make when applying them — are inescapably interpretive. A researcher applying content analysis to interview transcripts, news coverage, open-ended survey responses, policy documents, or social media posts is asking a version of the same question: what is actually in this material, how consistently can it be classified, and does the classification hold up when someone else applies it?

This guide covers the mechanics that make content analysis rigorous rather than impressionistic: the manifest/latent distinction, how to build and test a coding frame, how to choose a unit of analysis, and how to measure and report inter-coder reliability — plus how content analysis relates to the other major approaches to analyzing text, discourse analysis, narrative analysis, and thematic analysis, which are frequently confused with it.

Manifest content vs. latent content

The single most consequential design decision in a content analysis is whether coders are classifying manifest content or latent content.

  • Manifest content is what is literally, visibly present in the material — the surface-level, countable features. Coding whether a news article mentions a named funding source, whether a survey response contains a specific keyword, or how many times a policy document uses the word “shall” versus “may” are all manifest-content decisions. Manifest coding is more mechanical, generally faster to train coders on, and tends to produce higher inter-coder reliability because there is less room for disagreement about what counts.
  • Latent content is the underlying meaning, tone, or theme that a coder has to infer rather than simply observe — whether a passage expresses skepticism about a policy, whether an interview response conveys frustration, whether an image’s composition signals authority. Latent coding requires more interpretive judgment, takes longer to train, and reliability is harder to achieve and must be checked more carefully, but it captures meaning that manifest coding alone would miss entirely.

Most real studies use both: a manifest layer (counts, presence/absence, named entities) alongside a smaller number of latent, interpretive categories, with reliability reported separately for each because latent categories almost always show lower agreement.

Quantitative vs. qualitative content analysis

“Content analysis” is sometimes used loosely enough to cover two approaches that differ in what they’re trying to produce as output.

Dimension Quantitative content analysis Qualitative content analysis
Primary output Frequencies, proportions, and statistical relationships between categories A structured description of meaning within categories, often with representative quotations
Coding frame Fixed before coding begins; not revised mid-study Often revised iteratively as coding surfaces categories the initial frame missed
Sampling logic Usually a defined, often probability-based sample of a larger corpus Often purposive or the full available corpus, with less emphasis on statistical generalization
Reliability approach A formal, numeric inter-coder reliability coefficient is close to mandatory Reliability is often addressed through consensus coding, audit trails, or peer debriefing rather than (or in addition to) a coefficient
Typical claim made “Category X appeared in 34% of coded units, significantly more often in Y than Z” “Category X captures a recurring way participants described Y, illustrated by these examples”

These are poles on a spectrum, not a strict binary — many published studies are explicitly mixed, using a quantitative frequency count to establish which qualitative categories are worth exploring in depth, or vice versa.

Unitizing: choosing the unit of analysis

Before a single line of material can be coded, the researcher has to decide what counts as one codeable “unit.” This decision — unitizing — shapes everything downstream, including how reliability is calculated, since reliability coefficients compare coder decisions unit by unit.

  • Physical units — a whole document, article, image, or video, coded as a single case (e.g., “does this press release mention the funder by name: yes/no”).
  • Syntactical units — a word, sentence, or paragraph, used when the research question needs finer granularity than a whole document allows.
  • Referential units — every instance where a particular object, person, or concept is mentioned, regardless of how it’s phrased (e.g., every reference to “the committee,” whether by name, pronoun, or role).
  • Propositional units — a single assertion or claim, used when the research question is about what is being claimed rather than what is being mentioned.
  • Thematic units — a passage of any length that expresses a single idea or theme, common in latent and qualitative content analysis where meaning doesn’t align neatly with sentence or paragraph boundaries.

A common design mistake is picking a unit that’s too coarse for the research question (coding whole documents when the real variation happens within a document) or too fine to code reliably (asking coders to make latent judgments at the individual-sentence level, where context from surrounding sentences is stripped away). The unit of analysis should be decided and piloted before full-scale coding starts, not adjusted partway through, since a mid-study change invalidates reliability figures calculated under the old unit.

Building and testing a coding frame

The coding frame — the full set of categories, their definitions, and the decision rules for applying them — is the instrument of a content analysis in the same sense a survey questionnaire is the instrument of a survey. A weak coding frame produces unreliable, uninterpretable results no matter how carefully the coding itself is executed.

1. Decide deductive, inductive, or both

A deductive (a priori) coding frame is built from existing theory or prior literature before looking at the data — appropriate when the research question is testing a known framework. An inductive coding frame is built up from the data itself, typically by closely reading a subset of the material first and letting categories emerge — appropriate for exploratory questions where no adequate existing framework exists. Many studies combine both: a deductive skeleton of categories the researcher already expects to matter, refined and supplemented inductively after an initial read of the corpus.

2. Write explicit category definitions and decision rules

Each category needs a written definition specific enough that two people reading it would classify the same borderline case the same way. A weak definition (“mentions funding”) invites disagreement; a strong one specifies exactly what counts and gives a worked example of an edge case (“codes as ‘funder mentioned’ only when a specific named funding body or grant number appears; general references to ‘funding’ or ‘support’ without a named source do not count”). This document — sometimes called a codebook — is what makes the coding frame reusable by someone other than the person who designed it.

3. Make categories exhaustive and, where the design calls for it, mutually exclusive

Exhaustive means every unit in the corpus has somewhere to go — in practice this usually means including an explicit “none of the above” or “other” category, since real material routinely contains cases the original frame didn’t anticipate. Mutually exclusive means a unit can only be coded into one category within a given variable — required for frequency counts and most statistical analysis, though some qualitative designs deliberately allow multiple, overlapping codes per unit (multi-labeling) when the research question calls for it.

4. Pilot the frame on a subset before full coding

Piloting on a small, representative sample — with two or more coders working independently and then comparing results — is where most coding-frame problems surface: categories that overlap in practice, definitions that don’t cover common real cases, or a unit of analysis that’s the wrong grain size. Pilot coding should produce a reliability check of its own; a frame that fails reliability at the pilot stage should be revised before it’s used on the full corpus, not patched afterward.

5. Train coders and code the full corpus

Coder training means working through the codebook together, coding a shared practice set, and resolving disagreements by refining the codebook rather than simply overruling one coder. Once training reliability is acceptable, coders proceed — usually with a portion of the corpus double-coded throughout the study, not just at the start, so reliability can be reported for the actual data used in the analysis rather than only for a pilot phase.

Inter-coder reliability: measuring whether the coding frame actually works

Inter-coder reliability (also called inter-rater reliability in this context) asks a narrow, specific question: when two or more coders independently apply the same coding frame to the same material, how much do they agree? It is the mechanism that turns “I read this and categorized it this way” into a defensible, reportable claim about what the material contains, rather than one person’s impression.

Simple percent agreement (the proportion of units where coders agreed) is intuitive but overstates reliability, because it doesn’t correct for agreement that would happen by chance alone — a real problem when a category is rare or common, since coders can agree at a high rate just by both defaulting to the majority code. Chance-corrected coefficients are the standard for reporting:

  • Cohen’s kappa — the most widely reported coefficient for two coders working with nominal (categorical) data, correcting observed agreement for the agreement expected by chance given each coder’s marginal distribution.
  • Krippendorff’s alpha — more flexible than kappa: it handles more than two coders, missing data, and multiple levels of measurement (nominal, ordinal, interval, ratio) within the same framework, which is why it’s often preferred for larger or more complex coding teams.
  • Holsti’s coefficient — an older, simpler agreement measure that, like raw percent agreement, does not correct for chance; generally considered a weaker choice than kappa or alpha for reporting in a published study, though still seen in older literature.

There is no single universal cutoff for “acceptable” reliability — the threshold that counts as adequate depends on the field, the coefficient used, and how the categories are being used downstream — but a study should state which coefficient it used, on which unit of analysis, on what proportion of the corpus, and what threshold it treated as acceptable, rather than reporting a bare agreement percentage with no coefficient at all. This is also why the unit-of-analysis decision above matters so much: a reliability coefficient is only meaningful relative to the specific units coders were comparing.

For the related but distinct question of how inter-rater reliability compares to test-retest reliability (the same coder or instrument at two points in time, rather than two coders at the same point in time), see Test-Retest vs. Inter-Rater Reliability. For the broader relationship between reliability and validity as evaluation criteria, see Reliability vs. Validity.

A step-by-step content analysis procedure

  1. Define the research question and the corpus. What material is actually in scope — which documents, time range, publications, or communication channel — and why.
  2. Select a sampling strategy if the corpus is too large to code in full, and document it the same way any other sampling decision would be documented (see Sampling Methods).
  3. Choose the unit of analysis (physical, syntactical, referential, propositional, or thematic — see above) and confirm it fits the research question.
  4. Build the coding frame: deductive, inductive, or mixed; written definitions and decision rules for every category; an “other” category to keep the frame exhaustive.
  5. Pilot the frame on a representative subset with two or more independent coders, calculate a reliability coefficient, and revise categories that perform poorly.
  6. Train coders to consistency on the finalized frame before full coding begins.
  7. Code the full corpus, with a defined portion double-coded throughout for ongoing reliability checks, not only at the pilot stage.
  8. Report reliability for the actual data analyzed — coefficient used, sample of units it was calculated on, and the value achieved.
  9. Analyze and interpret — frequencies and statistical tests for quantitative designs; thick description and representative examples for qualitative designs; often both.

Content analysis vs. discourse analysis, narrative analysis, and thematic analysis

These four approaches are all applied to text, and published work sometimes uses the terms loosely, but they answer different questions and produce different kinds of output.

Method Core question Typical output
Content analysis What categories does this material contain, and how often/in what form? A coded, often countable classification of the corpus against a pre-defined or emergent coding frame
Thematic analysis What patterns of meaning (themes) run across this data set? A set of named, defined themes with supporting extracts, without the emphasis on a fixed, reliability-tested coding frame
Discourse analysis How does language construct meaning, power, or identity in this specific context? An interpretive account of how language is doing social/rhetorical work, not a category count
Narrative analysis How is this story structured, and what does that structure convey? An analysis of plot, sequence, and storytelling structure within accounts

In practice, content analysis is the most rule-based and reliability-focused of the four, which is exactly why it is taught first in most qualitative-methods curricula: the coding-frame-and-reliability discipline it requires transfers directly into thematic, discourse, and narrative work, even where those methods don’t require a formal reliability coefficient of their own.

Software for content analysis

Manual coding on paper or in a spreadsheet is workable for small corpora, but most studies beyond a modest sample size use dedicated qualitative data analysis software to manage the coding frame, track which coder applied which code to which unit, and calculate inter-coder reliability directly. See NVivo: Qualitative Data Analysis Software and the comparison of NVivo vs. ATLAS.ti vs. MAXQDA for how the major platforms differ, including their built-in reliability-coefficient tools. AI-assisted coding is an increasingly common addition to this workflow rather than a replacement for a human-defined coding frame — see AI in Qualitative Coding for how that fits in and where it still requires human reliability checks.

Frequently asked questions

Is content analysis qualitative or quantitative?

Both traditions exist under the same name — see the comparison above. Read a study’s methods section for which variant it used before assuming; the term alone doesn’t tell you.

What’s the difference between content analysis and thematic analysis?

Content analysis is built around a coding frame that is tested for inter-coder reliability, and frequently produces counts. Thematic analysis is built around identifying and reporting patterns of meaning (themes) without the same emphasis on a fixed, reliability-tested frame or on counting. See the comparison table above for the fuller picture, and Thematic Analysis: A Step-by-Step Guide for the full method.

How many coders does a content analysis need?

At minimum two, for at least a subset of the corpus, so that inter-coder reliability can be calculated. A single coder working alone can still classify material, but without a second independent coder there is no way to check whether the classification is reproducible rather than idiosyncratic to that one person’s judgment.

What counts as acceptable inter-coder reliability?

There is no single universal number — it depends on the coefficient, the field, and how the coded data will be used. What matters is reporting which coefficient was used, on what unit, on what share of the corpus, and being transparent about the value achieved, rather than citing a specific cutoff as if it were universally agreed.

Can content analysis be used on interview transcripts?

Yes — content analysis is commonly applied to interview data alongside thematic analysis, and the two are sometimes combined. See Coding Qualitative Interview Data: A Worked-Example Guide for a walked-through example of coding interview transcripts specifically.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →