Written and maintained by CASRAI Editorial Board
Last updated
CAQDAS — Computer-Assisted Qualitative Data Analysis Software — is a category of tools (NVivo, ATLAS.ti, MAXQDA, Dedoose, Taguette, Quirkos, and others) built to organize, code, and retrieve unstructured qualitative material: interview transcripts, focus-group recordings, field notes, open-ended survey responses, documents, images. What these tools have in common matters more than their individual feature lists: every one of them is fundamentally a retrieval and bookkeeping system for coded segments. None of them reads a transcript and tells you what it means. This page sets out, concretely, what a CAQDAS workflow actually automates, where the boundary sits, and which specific tasks researchers routinely — and wrongly — expect the software to do for them.
What CAQDAS actually does
Strip away the differences in interface and pricing between individual packages, and the core operations are the same across the category:
Coding and retrieval
You attach a label (a code, or “node” in NVivo’s terminology) to a segment of source material — a sentence, a paragraph, a clip. The software’s job from that point on is to keep every coded segment linked back to its exact location in the source, so you can retrieve everything tagged with a given code, across an entire dataset, in seconds. That retrieval speed is the single biggest practical advantage over manual coding on paper or in a spreadsheet, and it is the whole reason the category exists.
Organizing cases and attributes
Sources can be grouped into cases (typically one per participant) and attached to categorical attributes — age, site, condition, role — so you can filter or compare coded material across subgroups without manually re-sorting anything.
Querying and counting
Once material is coded, CAQDAS tools support text-search queries, word-frequency counts, and matrix or cross-tabulation queries that show how codes distribute across cases or attributes. A matrix query can tell you, instantly, that a code appears in 14 of 20 interviews and disproportionately among one participant group. It cannot tell you why, or whether that distribution is analytically meaningful.
Inter-coder reliability statistics
When more than one person codes the same material, most CAQDAS packages calculate an inter-coder agreement statistic (commonly Cohen’s kappa or a similar coefficient) automatically from the coding comparison. This is a genuine, valuable automation — the arithmetic is exactly the kind of thing software should do — but the statistic only describes how consistently two coders applied a scheme; it says nothing about whether the scheme itself was the right one.
Memos and the audit trail
Analytic memos and annotations attached to codes and sources give you a timestamped record of how your interpretation developed. The software stores and links these; it does not write them, and a project with no memos gets none of the audit-trail benefit regardless of how sophisticated the underlying tool is.
The honest boundary: what CAQDAS does not do
Every operation above is mechanical: tagging, linking, counting, cross-referencing. None of it is interpretation. The boundary that matters for a CAQDAS workflow is this one: the software organizes and retrieves coded segments; it does not decide what those segments mean, and it does not build the analytic argument that turns codes into a finding. A theme, in the sense the qualitative-methods literature uses the word, is an interpretive claim about a pattern of meaning across a dataset — see CASRAI’s thematic analysis guide for the full Braun and Clarke process. A code is not a theme, and no query, however elaborate, converts one into the other automatically. That step is done by the researcher, working from the coded material the software helped assemble.
This is not a shortcoming specific to any one product — it is true of the category by design, and it is true regardless of how much AI-assisted coding a current release adds (see CASRAI’s AI in qualitative coding entry): an AI suggestion is still a suggestion a researcher reviews, not an interpretation the researcher is relieved of making.
Tasks researchers wrongly expect CAQDAS to automate
Most of the frustration researchers report with CAQDAS tools traces back to a small, recurring set of mismatched expectations — not a software defect, but a misunderstanding of what the category was ever built to do.
1. Expecting AI-suggested or auto-coding to be publication-ready without review
Current versions of several major tools offer AI-assisted coding suggestions. These are a starting draft, not a finished coding pass — a researcher who accepts suggested codes wholesale, without independently checking a sample against the source material, has effectively delegated an interpretive judgment to a pattern-matching system that has no access to the study’s actual research question or theoretical framework.
2. Expecting the software to generate themes from codes
A matrix query or a cluster diagram can show you which codes co-occur or how frequently a code appears. It cannot tell you that three co-occurring codes actually represent one underlying theme, or that a frequently-coded segment is analytically less important than a rare one that captures a pivotal deviant case. Moving from codes to themes is manual, interpretive work — the visualization is an input to that thinking, not a substitute for it.
3. Expecting a query result to explain causation
A cross-tabulation showing that a code appears more often in one participant group than another is a description of the coded data, not an explanation of why the difference exists. That “why” question is answered by going back to the source material and reasoning about it — the query only tells you where to look.
4. Expecting an inter-coder reliability statistic to validate the coding scheme itself
A high kappa score means two coders applied the same scheme consistently. It says nothing about whether the scheme captured the right concepts in the first place — two coders can agree perfectly while both missing something the data actually shows. Reliability and validity are different properties, and the software only automates the former.
5. Expecting the tool to write the analysis
Memos, code definitions, and query exports are raw material for a findings section, not a findings section. The CAQDAS project file organizes what you found; turning that into a written argument that answers the research question is done outside the software, by the researcher, in the same way it always was.
6. Expecting word-frequency counts to substitute for close reading
A word-frequency query is a useful sanity check — it can flag a term you under-coded — but qualitative meaning frequently lives in phrasing, context, and what is left unsaid, none of which a frequency count captures. Treating a word cloud as an analysis result rather than a prompt for closer reading is a common and avoidable error.
7. Expecting the software to triangulate across data sources for you
Importing interview transcripts, survey responses, and documents into the same project makes it mechanically easier to query across them — it does not automatically reconcile what they say. Triangulation is a deliberate analytic comparison the researcher performs; the software’s contribution is keeping all the material in one searchable place while that comparison happens.
A CAQDAS workflow that respects the boundary
A workflow that treats the software as retrieval infrastructure rather than an analyst tends to follow the same shape regardless of which package is in use:
- Set up the project before importing everything. Decide, from the research question and study design, whether the coding scheme will be primarily deductive (codes drawn from an existing framework), inductive (codes emerging from the data), or a hybrid — see CASRAI’s guide to coding qualitative interview data for worked examples of all three. Importing material before this decision tends to produce a code list built ad hoc, which is harder to defend later.
- Keep the code tree disciplined. A code list that grows without periodic review — merging near-duplicates, retiring codes that turned out not to be analytically useful, writing a one-sentence definition for each code as it is created — becomes unmanageable well before a project ends. The definition matters most for team projects, where two coders working from an unwritten, shared understanding of a code name reliably drift apart.
- Run coding comparison early, not just at the end. Checking inter-coder agreement on a small batch of material before the full dataset is coded surfaces scheme problems while they are still cheap to fix, rather than after every transcript has been coded against a scheme two coders were quietly interpreting differently.
- Use AI-assisted suggestions as a first pass, reviewed against source. Spot-check a meaningful sample of AI-suggested codes against the actual transcript before trusting the tool’s suggestions for the rest of the material.
- Write memos as you go, not retrospectively. An audit trail reconstructed after coding is finished is a much thinner record than one built alongside it — and reviewers and committees increasingly expect to see evidence of that trail, not just the final code structure.
- Plan the export step. A native project file is a working analysis environment, not an archival format. If the coded dataset needs to remain usable independent of the software — for data sharing, deposit, or simply outliving a license — plan an explicit export (a codebook, a structured export of coded segments) rather than treating the project file itself as the long-term record.
Choosing among CAQDAS tools
The tools differ meaningfully on pricing model, platform support, collaboration features, and specific query/visualization capabilities — CASRAI’s NVivo vs. ATLAS.ti vs. MAXQDA comparison and Dedoose vs. Taguette guide cover those differences directly. None of that variation changes the boundary described above: every package in the category automates organization and retrieval, and none of them automates interpretation. Choosing a tool is a workflow and budget decision, not a decision about how much of the analytic work you get to skip.
Frequently asked questions
Does qualitative analysis software do the analysis for you?
No. It organizes, codes, and retrieves data, and automates mechanical tasks like inter-coder reliability arithmetic and cross-tabulation. Interpreting coded material into themes and findings remains the researcher’s work.
Can AI-assisted coding in NVivo or similar tools replace manual coding?
Current AI-assisted coding features generate suggestions a researcher reviews and confirms, not a finished, publication-ready coding pass. Treat them as a first draft to check against the source material, not a substitute for reading it.
Why doesn’t a high inter-coder reliability score guarantee good analysis?
The statistic measures consistency between coders applying the same scheme, not whether the scheme itself captured the right concepts. Two coders can agree with each other while both missing something the data shows.
What’s the difference between a code and a theme?
A code is a retrieval tag applied close to the data. A theme is an interpretive claim about a pattern of meaning across the dataset, built by the researcher from coded material — see CASRAI’s thematic analysis guide for the full process.
Do I need CAQDAS software for a small qualitative study?
Not necessarily — manual coding on paper or in a spreadsheet remains viable for small datasets. Dedicated CAQDAS tools become more valuable as a dataset or coding team grows, since they keep coded segments linked to context and calculate reliability statistics that manual coding cannot produce without separate calculation.
For the broader research-methods and software landscape this fits into, see CASRAI’s research tools and software hub and the qualitative research dictionary entry.








