Skip to main content
v2026.11,610 entries · CC-BY 4.0

Data Collection Methods: A Complete Guide for Researchers

A complete reference to quantitative, qualitative, and mixed-methods data collection techniques, how to choose between them, and the ethics, consent, and data protection steps required before collection begins.

Ask about Data Collection Methods: A Complete Guide for Researchers

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Data collection is the systematic process of gathering information to answer a defined research question. It sits between two other decisions a researcher has already made — the research question and the study design — and one that follows from it: the analysis. The method you choose to collect data determines what kind of analysis is even possible afterward, which is why “which method should I use” is really a downstream question, not a starting point.

This guide covers the full range of quantitative, qualitative, and mixed-methods data collection techniques, how to compare them on practical grounds (sample size, cost, time, what each produces), and the research-administration requirements — ethics review, informed consent, data protection law, and data management planning — that have to be settled before collection begins, not retrofitted afterward.

What Is Data Collection, and Where Does It Sit in the Research Process?

Data collection is the stage where a research design becomes actual observations, responses, measurements, or records. It follows the research question and study design (see sampling methods and the broader design literature on this site) and precedes analysis. The chain runs: research question → data type needed → collection method → analysis technique the data supports. Choosing a method out of order — picking a familiar tool before deciding what kind of evidence the question actually requires — is one of the most common design errors in student and early-career research, and it is usually irreversible once collection has started.

Primary vs. Secondary Data

Primary data is collected first-hand by the researcher for the specific study at hand — a survey fielded for this project, interviews conducted for this project, an experiment run for this project. Secondary data already exists, collected by someone else for a different original purpose: government statistics, administrative records, a prior study’s public dataset, electronic health records, or an existing repository. Secondary data is faster and cheaper to obtain but constrains the researcher to whatever variables, definitions, and population the original collection used, and it raises a distinct consent question covered later in this guide — the people in the data may never have consented to this use of it.

The Decision Logic

  1. What does the research question actually require? A question about prevalence or association across a population points toward quantitative, numeric data. A question about meaning, process, or lived experience points toward qualitative, textual/observational data. A question with both dimensions may need both.
  2. What data type does that imply? Numeric/structured (counts, scale scores, measurements) vs. textual/unstructured (transcripts, field notes, documents) vs. a combination.
  3. Which method produces that data type, at a feasible cost, in the population you can actually reach? This is where the comparison table below is useful.
  4. What analysis does that method’s output support? A structured survey with closed items supports statistical analysis; an unstructured interview transcript supports qualitative coding, not a t-test. Confirming this backward-compatibility before collection starts avoids ending up with data that cannot answer the question it was collected for.

Quantitative Data Collection Methods

Quantitative methods produce numeric or categorical data intended to be counted, measured, or statistically analyzed.

Surveys and Questionnaires

The most widely used quantitative instrument. Surveys can be self-administered (respondents complete it alone, on paper or online), interviewer-administered (a researcher reads items aloud, in person or by phone), or delivered purely online through a survey platform. Each mode trades off cost, response rate, and susceptibility to bias differently — interviewer-administered surveys tend to have higher completion rates but more social-desirability bias, while self-administered online surveys are cheap to scale but suffer lower and more selective response rates. See the dedicated guide on research questionnaire design for wording, ordering, and response-scale construction.

Structured Observation

The researcher records predefined, standardized categories of behavior or events as they occur, using a coding sheet or checklist fixed in advance. Unlike qualitative observation (below), the categories are closed and decided before data collection starts, which is what makes the output countable.

Experiments and Quasi-Experiments

Experiments manipulate an independent variable under controlled conditions and randomly assign participants to conditions, which is what supports a causal claim. Quasi-experiments share the manipulation but lack random assignment (e.g., comparing pre-existing groups), which weakens but does not eliminate the ability to infer causation. Both produce the numeric outcome data that inferential statistics are built to analyze.

Physiological and Instrument Measurement

Direct measurement using calibrated instruments — biometric sensors, lab assays, imaging, environmental sensors — rather than self-report. This removes recall and social-desirability bias but introduces its own error sources: instrument calibration, measurement precision, and inter-rater reliability when a human reads the instrument.

Existing Datasets and Administrative Data

Government statistics, institutional records, claims data, or previously collected research datasets, reused for a new analytical purpose. This is secondary data collection in the sense that the researcher is “collecting” it from an existing source rather than generating it, and it carries the consent-for-reuse question addressed in the research-administration section below.

Register and Records Extraction

Structured extraction of predefined fields from clinical, administrative, or institutional records — for example pulling diagnosis codes, dates, or outcome fields from an electronic health record or disease registry against a fixed extraction protocol. This differs from simply “using an existing dataset” in that the researcher defines and applies the extraction protocol themselves, which needs to be piloted and documented like any other instrument.

Qualitative Data Collection Methods

Qualitative methods produce textual, narrative, or observational data intended to capture meaning, process, and context rather than to be counted. See the dictionary entry on qualitative data collection techniques for a shorter reference version of the categories below.

In-Depth Interviews

One-on-one conversations that range from structured (a fixed script, asked identically to every participant — closer to a spoken questionnaire), to semi-structured (a guide of open-ended questions and probes, flexible in order and follow-up — see the practical guide to designing semi-structured interviews), to unstructured (a starting topic with the conversation left to develop, common in ethnographic and exploratory work).

Focus Groups

Guided group discussions, typically 6–10 participants, that generate data through participant interaction as much as through individual responses — useful for surfacing shared norms, disagreement, and language a one-on-one interview would not produce, but vulnerable to dominant-participant effects and social conformity within the group.

Participant and Non-Participant Observation

In participant observation the researcher takes part in the setting being studied while recording field notes; in non-participant observation the researcher records what happens without taking part. The distinction affects both the depth of access (participation typically yields richer access) and the risk of the researcher’s presence changing what is being observed (an observer effect, covered below).

Ethnography and Fieldwork

Extended immersion in a social setting, combining participant observation, informal interviewing, and document collection over a sustained period to produce a holistic account of a culture or group. See the guide on ethnographic research method for fieldwork and reporting conventions.

Document and Archival Analysis

Systematic analysis of existing texts — institutional records, correspondence, media, policy documents, historical archives — as the primary data source rather than as background reading. The researcher did not generate the documents but applies a defined method (often qualitative coding) to analyze them as data.

Diary and Elicitation Methods

Participants record their own experiences, in their own words, over time (solicited diaries, day reconstruction) or respond to a stimulus — a photograph, object, or prior artifact — that elicits richer reflection than a direct question would (photo-elicitation, object-elicitation).

Think-Aloud Protocols

Participants verbalize their thought process in real time while completing a task, most common in usability, cognitive, and instrument-testing research, producing a process-level account that a post-hoc interview about the same task would not reliably reconstruct.

Mixed-Methods Data Collection

Mixed-methods designs deliberately combine quantitative and qualitative collection within a single study, and the sequencing of that combination is a design decision in its own right — see the fuller treatment at mixed methods research. Three sequencing patterns account for most mixed-methods designs:

  • Explanatory sequential — quantitative data collected and analyzed first, with a follow-up qualitative phase (e.g., interviews) used to explain or contextualize the quantitative results.
  • Exploratory sequential — qualitative data collected first to explore a phenomenon or generate constructs, which then inform the design of a subsequent quantitative instrument (e.g., a survey built from interview themes).
  • Convergent — quantitative and qualitative data collected in parallel, analyzed separately, then merged or compared during interpretation.

Sequencing matters because it determines which strand drives instrument design for the other: an exploratory sequential survey is only as good as the interview data that generated its items, and an explanatory sequential interview guide is only as good as the quantitative results it is meant to explain.

Comparing Data Collection Methods

Method Data produced Typical sample size Cost / time Key strength Key limitation Analysis it enables
Self-administered survey Numeric / categorical Dozens to thousands Low–moderate Scales cheaply, standardized Low response rate, self-report bias Descriptive and inferential statistics
Interviewer-administered survey Numeric / categorical Tens to hundreds Moderate–high Higher completion, clarifies items Interviewer/social-desirability bias Descriptive and inferential statistics
Structured observation Numeric / categorical Varies by setting Moderate No reliance on self-report Observer effects, coding-scheme limits Frequency counts, inferential statistics
Experiment / quasi-experiment Numeric Tens to hundreds Moderate–high Supports causal inference Artificial setting, feasibility limits ANOVA, regression, group comparisons
Physiological / instrument measurement Numeric Varies High (equipment) Objective, no recall bias Calibration and instrument error Statistical analysis of measured values
Existing dataset / administrative data Numeric or structured records Often large Low Fast, inexpensive, often large-N Fixed to original variables/definitions Secondary statistical analysis
In-depth interview Textual / narrative ~10–30 (often saturation-driven) High (per participant) Depth, nuance, participant meaning Not generalizable, time-intensive Thematic / qualitative coding
Focus group Textual / narrative 3–6 groups of 6–10 Moderate Captures group norms and interaction Dominant-voice and conformity effects Thematic analysis
Participant observation / ethnography Field notes, narrative 1 setting, extended time Very high (time) Contextual, holistic understanding Researcher bias, low transferability Thematic analysis, thick description
Document / archival analysis Textual Varies by corpus Low–moderate Unobtrusive, historical reach Limited to what was recorded/retained Content or thematic analysis

Sampling and Its Effect on Data Collection

The collection method and the sampling strategy are linked, not independent choices. Probability sampling (simple random, stratified, cluster) gives every unit in the population a known chance of selection and is what supports statistical generalization to that population; it pairs naturally with surveys and structured methods aimed at population-level claims. Non-probability sampling (convenience, purposive, snowball) does not support the same generalization but is often the only feasible approach for qualitative and hard-to-reach-population research, where the goal is depth or theoretical insight rather than population estimates. See sampling methods for the full probability/non-probability breakdown, and generalizability in research for how the sampling strategy determines what a study can and cannot claim beyond its sample.

Data Quality: Validity, Reliability, and Common Sources of Error

Every collection method is vulnerable to specific, well-documented sources of error:

  • Validity — whether the method actually measures the construct it claims to measure. A poorly worded survey item can produce reliable but invalid data (consistently measuring the wrong thing).
  • Reliability — whether the method produces consistent results across repeated administration, raters, or items.
  • Measurement error — random or systematic inaccuracy in how a variable is captured, from instrument imprecision to ambiguous question wording.
  • Non-response bias — when those who don’t respond to a survey differ systematically from those who do, skewing results even with an adequate sample size.
  • Social desirability bias — respondents answering in a way they believe is more socially acceptable rather than accurately, most acute in interviewer-administered methods and sensitive-topic surveys.
  • Observer effects — a participant’s behavior changing because they know they are being observed, relevant to both structured observation and participant observation.

Every method above is vulnerable to a different mix of these: self-administered surveys minimize interviewer-driven social desirability bias but maximize non-response bias; observation minimizes self-report bias but maximizes observer effects; interviews minimize non-response but are labor-intensive to scale and depend heavily on interviewer skill and rapport.

The Research-Administration Layer: What Has to Be Settled Before You Collect Anything

Method selection is only half the design problem. Before any data collection involving human participants (or identifiable data about them) begins, a set of ethical, legal, and administrative requirements has to be in place — and unlike the method itself, these are not optional based on convenience.

Ethics and IRB/REC Approval

Research involving human participants generally requires review and approval by an institutional review board (IRB) or research ethics committee before data collection starts, not after. In the United States this review is governed by the Common Rule (45 CFR 46); equivalent frameworks (research ethics committees, RECs) apply internationally. Starting collection before approval is a compliance failure, not a paperwork formality, and can invalidate the data for publication.

Informed Consent

Participants must be given adequate information to decide voluntarily whether to take part, and that consent has to be documented appropriately to the method and risk level — see the guide to the four principles of informed consent and the broader informed consent reference. For secondary data (existing datasets, records, archives), the critical question is whether the original consent actually covers this new use — consent obtained for one study does not automatically extend to a different research purpose, and reusing identifiable data without that coverage typically requires a fresh ethics determination or a documented waiver.

GDPR, HIPAA, and Personal-Data Handling

Where data collection involves identifiable personal data, data protection law applies alongside (not instead of) ethics review. Under the EU/UK GDPR, Article 9(1) prohibits processing “special category” data (health, genetic, biometric, and other sensitive categories) except under specific conditions, one of which — Article 9(2)(j) — provides a research-specific derogation subject to appropriate safeguards. Article 4(5) defines pseudonymisation, and Article 89(1) names it as an accepted safeguard for research processing. In the US, HIPAA’s Privacy Rule governs protected health information and operates as a legally distinct regime from the Common Rule/IRB system — a study can require both HIPAA authorization (or a waiver) and IRB approval, with separate and sometimes divergent criteria. See the full walkthrough at GDPR and data protection compliance in research.

Data Management Plans and Funder Requirements

Most major funders now require a data management plan (DMP) submitted before funding is awarded, specifying how data will be collected, documented, stored, secured, and eventually shared or preserved. This is not a post-hoc formality — the collection method itself (what format, what metadata, what identifiers) should be chosen with the DMP’s storage and sharing commitments in mind. See the practical DMP template and structure guide for what funders expect in this section.

Secure Storage, Transfer, and Retention

Collected data — especially identifiable data — needs a defined secure-storage plan (encryption, access control), a secure transfer method if data moves between sites or team members, and a retention schedule consistent with both funder/institutional policy and data protection law’s storage-limitation principle. These decisions should be made before collection starts, since retrofitting security controls onto data already collected in an insecure format is far harder than designing the collection instrument around them from the outset.

Anonymisation vs. Pseudonymisation

Anonymisation irreversibly removes the ability to link data back to an individual; properly anonymised data generally falls outside data protection law’s scope entirely. Pseudonymisation replaces identifying fields with a key or code but keeps the data re-identifiable by whoever holds the key, which means pseudonymised data remains personal data under GDPR (Recital 26) even though it carries a lower risk profile. See the comparison at anonymisation vs. pseudonymisation and the practical guide to data anonymisation in research for techniques and when each is actually required.

Data-Sharing Agreements

Where data will be collected by, or shared with, an external partner — a collaborating institution, a clinical site, a commercial vendor — a data-sharing (or data transfer/use) agreement should specify permitted uses, security obligations, and retention/destruction terms before any data actually moves. This is distinct from, and in addition to, participant-level consent.

Collecting Data Across Jurisdictions

Multi-country data collection can trigger overlapping and sometimes conflicting legal regimes at once — GDPR’s restrictions on transferring personal data outside the EU/EEA (Article 44), sector-specific laws like HIPAA in the US, and each jurisdiction’s own local ethics review requirements. A collection instrument, consent form, and storage plan designed for a single-jurisdiction study will often need to be reworked, not just translated, before it can be deployed across borders.

Designing and Piloting the Instrument

Whatever method is chosen, the instrument itself — the questionnaire, interview guide, coding sheet, or extraction protocol — needs to be drafted with the eventual analysis in mind, then tested before full deployment. A pilot study run with a small sample from the target population (or a close proxy) is the standard way to catch ambiguous wording, unworkable question order, technical failures, and unrealistic time burdens before they compromise the full dataset. Piloting also produces an early estimate of variability that can sharpen the sample-size calculation for quantitative designs, and it is the point at which consent materials and data-handling procedures get their first real-world test as well, not just the instrument content.

Frequently Asked Questions

What is the difference between a data collection method and a research design?

Research design (e.g., cross-sectional, longitudinal, experimental) is the overall architecture that determines what a study can claim; data collection method (e.g., survey, interview, observation) is the specific technique used to gather the data within that design. A single design can use multiple collection methods, and the same method can be used within different designs.

Can I combine multiple data collection methods in one study?

Yes — this is standard practice, not just formal “mixed-methods” research. A survey might be supplemented with follow-up interviews, or structured observation combined with document analysis, as long as each method’s output is analyzed appropriately and the combination is planned rather than improvised after the fact.

Which data collection method is most reliable?

No single method is universally most reliable — reliability depends on what is being measured and how the method is executed. A well-piloted structured survey can be highly reliable for stable, well-defined constructs; a semi-structured interview conducted by a trained interviewer can be highly reliable for capturing consistent themes across participants. The relevant question is fit to the research question, not a ranking of methods in the abstract.

Do I need IRB approval before piloting my instrument?

In most institutional settings, yes — if the pilot involves human participants and the data will be used or reported in any way (including to refine the instrument), it generally falls under the same review requirement as the main study. Confirm this with your institution’s IRB or ethics office before collecting any pilot data; treating a pilot as automatically exempt is a common and avoidable compliance mistake.

What is the difference between primary and secondary data collection?

Primary data collection generates new data specifically for the current study (a survey, interview, or experiment run for this project). Secondary data collection uses data that already exists for another original purpose (administrative records, a prior dataset, government statistics). Secondary data is faster and cheaper to access but constrains the researcher to the original variables and raises distinct consent-for-reuse questions.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →