A survey lives or dies on the questionnaire itself. Sampling can be sound and the response rate respectable, and the data can still be unusable if the items were poorly worded, the response options were mis-specified, or the questionnaire mixed independent items with a multi-item scale without knowing the difference. This guide covers the construction decisions inside a research questionnaire that determine whether it produces analyzable data — wording, response-option design, scale construction, ordering effects, sensitive-question handling, and pretesting — as distinct from the broader sampling and administration decisions covered under survey research methods.
Question wording: the three failure patterns that silently corrupt data
Most bad survey data is not caused by a wrong statistical test. It is caused by an item that respondents could not answer the way the researcher assumed they would. Three wording patterns account for most of it.
Double-barrelled questions
A double-barrelled question asks about two things at once but allows only one answer, so a respondent who agrees with one half and disagrees with the other has no valid response available. “The lab’s equipment is modern and well-maintained” forces a single rating onto two separate constructs (equipment age and maintenance quality) that can diverge. The fix is to split it into two items and, if they are meant to jointly represent one underlying construct, treat them as separate items within a scale rather than as one question.
Leading questions
A leading question embeds a cue toward a particular answer in its wording. “Don’t you agree that the new onboarding process is an improvement?” primes agreement before the respondent has evaluated anything. The corrected form states the object neutrally and lets the response scale carry the valence: “How would you rate the new onboarding process compared to the previous one?”
Loaded questions
A loaded question carries an unstated, often emotionally charged assumption the respondent is not given room to reject. “How much has the department’s failure to communicate affected your work?” presupposes that a failure to communicate occurred; a respondent who disagrees with that premise cannot answer honestly. The fix separates the presupposition from the measurement: first ask whether communication has been a problem at all, then, conditionally, ask about its effect.
Annotated before/after: five items rewritten
The comparisons below use one small, made-up department-survey scenario purely to illustrate the wording patterns above and the fixes described in this guide — it is not drawn from any real survey, organization, or dataset.
| Problem | Poorly worded item | Corrected item | What changed |
|---|---|---|---|
| Double-barrelled | “The training was well-organized and useful.” | Item A: “The training was well-organized.” Item B: “The training was useful to my work.” | Split into two independently answerable items. |
| Leading | “Don’t you think the new scheduling system saves time?” | “How, if at all, has the new scheduling system changed the time you spend on scheduling tasks?” | Removed the embedded agreement cue; neutral stem, response scale carries direction. |
| Loaded | “How frustrating is it that IT support is so slow to respond?” | “How would you rate the response time of IT support?” (followed conditionally by an open item on impact) | Removed the unstated negative premise from the stem. |
| Vague response option | “How often do you use the shared equipment?” — Never / Sometimes / Often / Always | “In the past 30 days, how many times did you use the shared equipment?” — 0 / 1–2 / 3–5 / 6–10 / More than 10 | Replaced subjective frequency labels (no shared meaning across respondents) with a defined recall period and numeric bands. |
| Priming via order | Satisfaction item placed immediately after a block of items listing specific complaints | Overall satisfaction item moved to the start of the questionnaire, before any topic-specific items | Removed the priming effect of preceding content on a general evaluative judgment. |
Response-option design
Response options are not a cosmetic layer on top of the question — they are part of the measurement. Three design choices matter most:
- Balance. Response options should offer an equal number of positive and negative categories around any midpoint, so the scale does not push responses in one direction before the respondent has answered anything.
- Exhaustiveness and mutual exclusivity. Every plausible respondent state needs exactly one applicable option; categorical items (e.g., employment status, discipline) need an “other” or “not applicable” escape valve or item non-response rises.
- Defined, not vague, frequency/quantity labels. “Sometimes” and “often” mean different things to different respondents. Where possible, anchor frequency items to a defined recall period (“in the past 30 days”) and, ideally, numeric bands rather than adverbs.
Likert scale construction
A Likert item asks a respondent to rate agreement, frequency, or intensity along an ordered scale; a Likert scale, properly speaking, is a set of such items summed or averaged to measure one underlying construct (see the scale-vs-items distinction below). Three construction decisions recur on every project.
How many points?
Five- and seven-point scales are the most commonly used formats in the literature, and no single point count is correct for every context. More points can capture finer gradations and typically raise reported reliability slightly, but past roughly seven to nine points respondents struggle to reliably distinguish adjacent categories, and gains flatten out.
Include a midpoint?
This is a genuine, context-dependent trade-off rather than a settled rule. Including a neutral midpoint gives genuinely neutral or undecided respondents a legitimate, honest option instead of forcing them toward an arbitrary side; omitting it forces a directional choice, which is sometimes the researcher’s actual intent. The commonly cited guidance from the survey-methodology literature (Krosnick & Fabrigar) is to offer a midpoint on obscure or low-salience topics, where many respondents genuinely have no basis for an opinion, and to consider omitting it on topics where social-desirability pressure could push respondents to hide behind a “neutral” answer instead of committing to a real, if uncomfortable, position.
Label every point, or only the endpoints?
Fully labelling every scale point (e.g., “Strongly disagree / Disagree / Neither agree nor disagree / Agree / Strongly agree”) gives every respondent the same verbal anchor for each numeric position, which improves comparability across respondents and is the generally preferred approach for scales intended to be summed into a composite score. Labelling only the endpoints (“1 = Strongly disagree” … “5 = Strongly agree”, with 2–4 unlabelled) leaves the intermediate points open to individual interpretation, which introduces more measurement noise across respondents but can be defensible when the underlying construct is a continuous magnitude judgment rather than a discrete set of agreement levels.
Scale vs. independent items — and why this decides whether Cronbach’s alpha even applies
This distinction is the single most consequential decision in questionnaire construction, because it determines which analysis is valid afterward.
- A scale is a set of multiple items deliberately written to measure one underlying, usually unobservable, construct (e.g., job satisfaction, research self-efficacy, perceived usability). Because the items are theorized to be different observable manifestations of the same latent construct, they are expected to correlate with each other, and it is meaningful to sum or average them into a single composite score and to report an internal-consistency statistic such as Cronbach’s alpha for that composite.
- A set of independent items asks about genuinely separate constructs that happen to sit in the same questionnaire — age, department, preferred communication channel, and a satisfaction rating are four different things, not four indicators of one thing. These items should never be summed into a composite, and computing Cronbach’s alpha across them is a methodological error: alpha assumes the items share a common underlying factor, and applying it to unrelated items produces a number that looks like a reliability statistic but measures nothing coherent.
The practical test: before writing a single Likert item, decide explicitly whether it belongs to a scale (and which one) or stands alone. If it belongs to a scale, every other item in that scale should plausibly correlate with it because they share a construct. If you cannot articulate what construct two items would jointly measure, they are independent items, not a scale — treat and report them as such.
Question order and priming
Item order is not neutral. Two order effects recur across survey-methodology research:
- Priming. Answering earlier items can activate specific considerations that carry over into how a respondent interprets and answers a later, more general item. A general “overall satisfaction” item placed after a block of items about specific complaints will tend to be primed by those complaints; placing general/overall items before topic-specific detail items avoids this.
- Fatigue and order-position effects. Items late in a long questionnaire receive less careful attention (satisficing) than items near the start; place the items most central to the research question earlier, and place demographic/classification items at the end unless they are needed to route respondents through skip logic.
Sensitive-question handling
Questions touching income, health status, substance use, discrimination, or other sensitive topics are especially exposed to social-desirability bias — the tendency to answer in the way one believes is more socially acceptable rather than truthfully. Common mitigations used in the survey-methodology literature: placing sensitive items later in the questionnaire after rapport/context has been established; using self-administered rather than interviewer-administered modes for the most sensitive items; offering broad response bands instead of exact figures (e.g., income ranges rather than an exact dollar amount); and, where feasible for highly sensitive topics, using indirect estimation techniques designed specifically to reduce social-desirability distortion. Always pair sensitive items with a clear statement of confidentiality/anonymity and, for identifiable data, the study’s actual data-handling and retention practice — not just a general assurance.
Pretesting and cognitive interviewing
A questionnaire should never go to its full sample without being pretested first. Two complementary methods are standard:
- Cognitive interviewing. A small number of respondents (often 5–15, done in rounds) complete the questionnaire while thinking aloud or while being probed afterward about how they interpreted specific items, what they considered when forming an answer, and why they chose the response option they did. This is the method of choice for catching wording ambiguity, double-barrelled items, and response options respondents interpret differently than the researcher intended — problems a standard pilot survey’s quantitative results will not reveal on their own, because a low or inconsistent response to an item looks the same whether it is caused by bad wording or a genuinely mixed population.
- Pilot testing. A small-scale trial administration of the near-final questionnaire to a sample similar to the target population, used to check overall completion time, item non-response rates, floor/ceiling effects in response distributions, and (for multi-item scales) preliminary internal-consistency statistics before committing to the full-scale administration.
Content-validity review by subject-matter experts is typically done earlier in the sequence, alongside or before cognitive interviewing, to confirm the item pool adequately represents the construct before wording refinement begins.
Response-rate management
Low response rates threaten a survey’s external validity by raising the risk that respondents differ systematically from non-respondents (nonresponse bias) — a risk that exists independently of whether the questionnaire itself is well constructed. Practices commonly used to manage response rate: keeping the questionnaire as short as the research question allows (length is one of the most consistently documented predictors of both response rate and per-item data quality); sending advance notice before the survey launches; using multiple, spaced follow-up reminders to non-respondents; offering multiple response modes where feasible; and, for known populations, comparing early vs. late respondents or respondents vs. known population characteristics as a check for nonresponse bias when the response rate is low. A well-designed questionnaire and an adequately managed response rate are separate requirements — a beautifully worded instrument returned by an unrepresentative fraction of the sample still yields data that cannot be generalized to the intended population, a concern connected to the sample-size and power questions covered in power analysis and sample size calculation.
Frequently asked questions
What is the difference between a questionnaire and a survey?
The questionnaire is the instrument — the fixed set of items and response formats. The survey is the full research method built around it: defining the sampling frame, choosing an administration mode, fielding the instrument, and reporting a response rate. See research questionnaire and survey research methods for the full definitions.
Should every questionnaire use Likert scales?
No. Likert-type items are well suited to measuring attitudes, agreement, and perceived intensity, but factual/classification items (age, role, frequency of a countable behavior) are usually better captured with categorical, numeric, or defined-band response options rather than an agreement scale.
How many items does a scale need before Cronbach’s alpha is meaningful?
There is no fixed minimum in principle, but a scale with only one or two items rarely produces a stable or interpretable alpha; most published scales use at least three to five items per construct. See Cronbach’s alpha for interpretation thresholds and when to use McDonald’s omega instead.
Is a 5-point or 7-point Likert scale better?
Neither is universally correct. Both are widely used; more points can capture finer distinctions and modestly raise reported reliability, but responses become harder for respondents to reliably differentiate past roughly seven to nine points. Choose based on the construct’s expected granularity and consistency with any existing validated instrument you are adapting.
Do I need IRB or ethics review to pilot-test a questionnaire?
That depends on your institution’s determination process and whether the pilot data will be used or reported as part of the research, not just as an internal wording check. Confirm with your local IRB/ethics office rather than assuming pretesting is automatically exempt.
Related concepts
- Research Questionnaire — the instrument-level definition this guide builds on.
- Survey Research Methods — the broader methodology (sampling, administration, response rate).
- Cronbach’s Alpha — the reliability statistic that applies only to genuine multi-item scales, not independent items.
- Power Analysis and Sample Size Calculation — determining how many respondents a survey needs before response-rate attrition.
- Research Methods & Statistics — the full cluster hub.







