Written and maintained by CASRAI Editorial Board
Last updated
Two surveys that ask the same questions of the same population can still produce different answers if they are administered in different modes. That gap has two distinct causes, and confusing them leads to the wrong fix. A mode selection effect happens when the kind of person who chooses (or is more likely to complete) one mode differs systematically from the kind of person who chooses another — an internet-only sample skews younger and more educated than a mail sample of the same frame, for reasons that have nothing to do with the questions themselves. A mode measurement effect happens when the same respondent would answer differently depending on which mode presented the question — a phone interview yields more socially desirable answers on a sensitive item than a self-administered web form does, even holding the respondent constant. Mixed-mode design exists to expand coverage and control cost, but every mode combination has to manage both effects, and they require different fixes.
Where mode measurement effects come from
Measurement effects trace to a small number of structural differences between how modes present information and how respondents process it.
Interviewer presence and social desirability
Interviewer-administered modes (telephone, face-to-face) consistently produce more socially desirable responding on sensitive topics — substance use, income, health behaviors, attitudes with a normative “right answer” — than self-administered modes (web, mail), because respondents moderate answers in front of another person in a way they do not on a form. See social desirability bias for the underlying mechanism and how it interacts with question sensitivity independent of mode.
Visual vs. aural presentation
Visual modes (web, mail, paper) let a respondent see the full response scale at once, re-read the question, and revise an answer before submitting. Aural modes (telephone) present options serially and rely on the respondent’s memory of the scale, which measurably shifts response distributions — aural administration tends to produce more primacy effects for visually presented lists and more recency effects for orally read lists, because the last-heard option is freshest in memory. This is a presentation-format effect, not a respondent-quality difference, and it persists even when the same person answers the same item across modes.
Acquiescence and extreme responding
Telephone surveys tend to show higher acquiescence (a tendency to agree regardless of content) and more extreme-category selection than self-administered modes, plausibly because responding quickly and agreeably reduces the interpersonal cost of the interaction. Self-administered modes give respondents more time and less social pressure, which tends to produce more differentiated, less extreme response patterns.
Missing data and item nonresponse
Self-administered modes usually show higher item-level missingness (a respondent can skip a question a phone interviewer would have prompted them on) but lower unit-level social-desirability distortion. Interviewer modes show the reverse: an interviewer can probe an unclear or blank answer, which raises completion but can also introduce interviewer-specific measurement variance across the interviewer pool.
Two design strategies, and why they are not interchangeable
Once a study commits to more than one mode, there are two structurally different ways to build the instrument, and picking one has real consequences for whether mode and measurement can be separated later.
Unimode (unified-mode) construction
Unimode construction, the approach Don Dillman set out in Internet, Phone, Mail, and Mixed-Mode Surveys: The Tailored Design Method (4th ed., with Jolene D. Smyth and Leah Melani Christian, Wiley, 2014), deliberately writes every question to be as functionally equivalent as possible across the modes in use — the same number of response categories, the same category order, no mode-exclusive question types (no matrix grids read aloud, no “select all that apply” phrasing that behaves differently spoken vs. read), and consistent instructions. The goal is not identical presentation, which is impossible across a web page and a phone call, but equivalent stimulus: a respondent should be choosing among functionally the same set of answers regardless of mode. Unimode construction reduces measurement effects at the cost of sometimes producing a version that is not optimal for any single mode — a compromise question that works acceptably everywhere rather than ideally on one platform.
Mode-specific (optimized) design
The alternative lets each mode use its own best practice — a web form uses drop-downs and dynamic skip logic, a phone script uses shorter category lists suited to memory, a paper form uses visual layout cues unavailable on a call — and then attempts to adjust for the resulting measurement differences statistically after the fact (weighting, calibration against a reference mode, or mode as a covariate in analysis). This produces a better within-mode respondent experience but leaves a genuine measurement gap that has to be modeled rather than designed away, and that adjustment is only as good as the assumptions behind it.
Neither approach eliminates mode effects; unimode construction designs against them up front, mode-specific design defers the problem to analysis. The right choice depends on whether the study can tolerate a compromise instrument or would rather carry a statistical adjustment burden later.
Separating mode effects from selection effects
Because selection and measurement effects both show up as “the results differ by mode,” a raw comparison of respondents-by-mode cannot tell you which one is responsible — and the fix for one does nothing for the other. Three designs actually separate them:
Randomized mode-assignment experiments
Randomly assigning sampled units to a mode (rather than letting respondents self-select) holds the respondent population constant across modes by design, so any remaining difference in answers is attributable to measurement, not selection. This is the cleanest design and the one methodological literature on mode effects is mostly built from, but it is not always operationally available — a sampling frame with only postal addresses cannot randomize a phone-eligible subset that does not exist.
Sequential (mode-choice) designs as a natural comparison
A common operational pattern — “web push to mail,” where every sampled unit first receives a web invitation and non-responders are followed up by mail — is not a controlled experiment, but comparing early web responders to late mail responders on observable characteristics gives a rough read on how much of the mode gap is selection (who responded to which wave) versus measurement (how the same kind of respondent answered differently). It is suggestive, not dispositive, without further modeling.
Re-interview or mixed-mode calibration studies
Re-administering the same instrument to a subsample in a second mode isolates within-person measurement effects directly, since the same respondent’s mode-one and mode-two answers can be compared without any selection confound. This is the standard approach when a study needs to combine or trend data collected in different modes over time (for example, moving a repeated survey from telephone to web) and needs an empirical adjustment factor rather than an assumption.
Practical guidance for a mixed-mode study
- Decide unimode vs. mode-specific before writing a single item — retrofitting unimode construction onto an instrument already optimized per-mode usually means rewriting most of it.
- Budget for the confound, not just the modes. If mode assignment will not be randomized, plan a design (sequential response-wave comparison at minimum, a calibration subsample if resources allow) that gives some empirical purchase on how much of any observed mode gap is selection versus measurement, rather than reporting a pooled result with no way to characterize it.
- Report mode as a design feature, not an afterthought. AAPOR’s Standard Definitions (10th edition, 2023) reorganized its response-rate framework by sampling frame rather than by mode, but disclosure of which modes were used, in what sequence, and at what proportion of final respondents still belongs in any methods write-up, alongside the same non-response disclosure any single-mode survey owes.
- Treat sensitive items differently from neutral ones. Measurement-effect size on socially desirable topics is typically larger than on factual or neutral items, so a study asking about substance use, income, or normatively charged attitudes has more at stake in the mode decision than a study asking about product feature usage.
Frequently asked questions
Does mixed-mode survey design always introduce mode effects?
Any study using more than one mode has the potential for both selection and measurement effects; whether they are large enough to matter depends on the topic (sensitive items are more exposed than neutral ones) and how different the modes are (web and mail, both self-administered and visual, produce smaller measurement gaps than web and telephone).
Is unimode construction the same as making every mode identical?
No. Identical presentation across a web page, a phone script, and a paper form is not achievable. Unimode construction targets equivalent response options and question structure — the same underlying stimulus — not identical visual or aural delivery.
Can weighting fix a mode effect after data collection?
Weighting can address known differences in who responded by mode (a selection-effect fix) if the variables driving mode choice are measured and included. It cannot correct a genuine measurement effect — the same respondent answering differently depending on mode — because that difference is not a function of sample composition.
How is a mode effect different from a general response bias?
Response biases such as social desirability, acquiescence, or extreme responding can occur within a single mode. A mode effect specifically refers to those biases (or coverage/selection patterns) varying systematically between modes in the same study, which is what makes combined-mode data harder to interpret than single-mode data.
For the sampling, administration, and wording decisions that sit alongside mode choice, see survey research methods and questionnaire design. Related instrument-construction topics: leading questions in surveys, question order effects, skip logic and branching, and Likert scale design. See the research methods pillar for the full sub-cluster.








