Social desirability bias is the systematic tendency for respondents to answer in a way that will be viewed favorably by others, rather than in a way that accurately reports their behavior, attitude, or belief. It is not a sampling problem — the right people are in the study, and even a perfect sampling frame does not fix it, because the distortion happens inside the instrument, at the moment a respondent decides how to answer a question they find sensitive. That is why it is filed as a measurement problem rather than a design or sampling one, and why the fix is instrument design, not a bigger or more representative sample.
What Social Desirability Bias Is, and What It Isn’t
Operationally, a response shows social desirability bias when it is displaced toward the culturally approved answer rather than the respondent’s actual behavior or belief, and the size of that displacement is correlated with how sensitive or normatively loaded the topic is. Self-reported exercise frequency, for example, is reliably higher than accelerometer-measured activity; self-reported caloric intake is reliably lower than doubly-labeled-water measurements. Both are consistent with social desirability bias because the direction of the error tracks the social norm (be active, don’t overeat), not random noise.
It is worth distinguishing from three neighboring concepts it commonly gets collapsed into:
- Sampling bias is a problem with who answers, not how they answer — see sampling bias for the distinction. A survey can have a perfectly representative sample and still be badly distorted by social desirability effects within every respondent who answers.
- Acquiescence bias (the tendency to agree with statements regardless of content) is a response style driven by the format of the question, not its content; social desirability bias is driven specifically by how normatively loaded the content is.
- Non-response bias occurs when certain respondents decline to answer at all; social desirability bias occurs among respondents who do answer, but answer inaccurately.
The Two Mechanisms: Self-Deception and Impression Management
The influential framework here, developed by Delroy Paulhus, splits socially desirable responding into two distinct components rather than treating it as one phenomenon:
- Self-deceptive enhancement — the respondent honestly believes the favorable self-description; there is no intent to deceive the researcher, because the respondent isn’t aware of the distortion in their own self-assessment.
- Impression management — the respondent knows the honest answer and deliberately reports something more favorable, calibrated to who they believe is reading the response.
The distinction matters for countermeasures. Impression management responds strongly to anonymity and to removing an interviewer from the room, because it is a conscious, audience-directed distortion. Self-deceptive enhancement responds much less to anonymity, because the respondent isn’t consciously withholding anything — they are reporting what they believe is true. A researcher who only anonymizes a survey has addressed impression management and left self-deceptive enhancement largely untouched.
The most widely used instrument for measuring an individual’s general tendency toward socially desirable responding is the Marlowe-Crowne Social Desirability Scale, a 33-item true/false inventory; Paulhus’s own Balanced Inventory of Desirable Responding (BIDR) separates the two components into 20-item self-deceptive-enhancement and impression-management subscales. Both are used less as end products in themselves and more as covariates — a researcher administers one alongside the substantive survey and statistically controls for the score, or screens out respondents who score above a threshold on validity checks.
Where It Bites Hardest
Social desirability bias is not evenly distributed across topics. It concentrates wherever a real social or legal norm exists for the “right” answer:
- Health behaviors — self-reported smoking, alcohol consumption, diet, and physical activity are consistently biased toward the healthier answer relative to biomarker or device-measured ground truth.
- Income and socioeconomic status — both over- and under-reporting occur depending on context (over-reporting to appear successful in some settings, under-reporting to qualify for a benefit or appear modest in others).
- Prejudice, political attitudes, and stigmatized beliefs — direct questions about racial attitudes, discriminatory beliefs, or unpopular political positions are among the most heavily studied applications of the countermeasures below, precisely because direct self-report is known to understate the true prevalence.
- Sexual behavior, substance use, and other illegal or stigmatized conduct — self-reported drug use, risky sexual behavior, and criminal conduct are all reliably under-reported in direct-question formats.
- Workplace and subordinate feedback — employees rating a supervisor, or subordinates rating conditions to someone who could plausibly see the results, show the same displacement even without a formal sensitivity label on the topic.
A useful diagnostic before designing any instrument: ask whether there is a socially “correct” answer to the question as posed. If most people would guess the same direction of bias before seeing any data, the topic needs one of the countermeasures below rather than a direct question.
How to Detect It in Data You Already Have
Detection usually relies on one of three approaches, roughly in order of rigor:
- Comparison against an objective record. Where a biomarker, administrative record, or device measurement exists (accelerometer data for activity, tax records for income, pharmacy records for medication adherence), the gap between self-report and the objective measure is a direct estimate of the bias, though it also captures ordinary recall error, so the two are not perfectly separable.
- Correlation with a social desirability scale. Administering the Marlowe-Crowne scale or BIDR alongside the substantive measure and checking whether the substantive item correlates with social-desirability score is standard practice in psychometric validation; a significant correlation is evidence (not proof) that the item is vulnerable.
- Mode comparison. Running the same question through an interviewer-administered mode and a self-administered mode (see below) and comparing responses is a practical, lower-cost signal: if the two modes diverge on a sensitive item and agree on a neutral one, the divergence on the sensitive item is attributable to the mode’s effect on candor, not to random measurement error.
Countermeasures: Designing It Out of the Instrument
Because the distortion happens at the point of response, the fix has to happen in the instrument, not in the sample or the analysis. The methods below trade off respondent-level protection against statistical usability — the strongest privacy protections (randomized response, list experiments) sacrifice the ability to know any individual respondent’s true answer, in exchange for an unbiased group-level estimate.
| Method | How it works | What it protects against | Main limitation |
|---|---|---|---|
| Self-administration (paper, online, ACASI) | Removes the interviewer from the moment of response; respondent enters answers privately | Impression management directed at an interviewer | Does not address self-deceptive enhancement; requires literacy or audio-assisted delivery |
| Indirect / third-person questioning | Asks about “people in your situation” or “your colleagues” rather than the respondent directly | Reluctance to admit a norm-violating belief as one’s own | Projective inference — assumes the respondent’s answer for others approximates their own view, which is not guaranteed |
| Randomized response technique (Warner, 1965) | Respondent uses a private randomization device (e.g., a coin flip) to decide whether to answer the sensitive question truthfully or give a forced answer; only the respondent knows which rule applied | Both impression management and any residual interviewer effect, at the individual level | No individual response can ever be interpreted; requires a larger sample and more complex analysis to recover the group-level prevalence |
| List experiment (item count technique) | Control group gets a list of non-sensitive items and reports only the count endorsed; treatment group gets the same list plus the sensitive item, also reporting only a count. The difference in mean count between groups estimates the sensitive item’s prevalence | Direct attribution of a sensitive belief or behavior to any individual | Requires two large randomized groups; loses precision at the individual level entirely; the “no design effect” assumption (the sensitive item doesn’t change how respondents count the others) can fail |
| Forced-choice / ipsative formats | Respondent ranks or picks between statement pairs matched for desirability, rather than rating each on its own desirability-laden scale | Uniform inflation across all items (ceiling effects from desirability, not the underlying trait) | Produces ipsative (relative) rather than normative scores, which complicates some statistical comparisons across respondents |
| Anonymity and confidentiality assurance | Explicit, credible assurance that no individual response can be linked back to the respondent | Impression management, particularly where the audience (employer, funder, institution) is identifiable to the respondent | Assurance has to be credible to be effective; a stated but implausible anonymity claim (e.g., a small workgroup survey) does not reduce bias |
In practice, most instruments combine two or three of these rather than relying on one. A common, lower-overhead pattern for a general survey is self-administration plus an explicit, specific anonymity statement plus, for the single most sensitive item, an indirect or third-person phrasing; a list experiment or randomized response is reserved for cases where even indirect phrasing is judged too risky for the respondent (illegal behavior, stigmatized health status) and where the researcher only needs the group-level prevalence, not an individual-level predictor.
What Social Desirability Bias Does Not Fix
None of the countermeasures above substitute for good questionnaire design more broadly. A double-barrelled or leading question is still a double-barrelled or leading question even inside a list experiment. Randomized response and list experiments also do not help with reliability problems unrelated to social desirability (e.g., ambiguous wording, unstable constructs), and they should not be reached for as a default — they carry a real statistical-power and complexity cost, and are justified specifically when the topic is sensitive enough that direct self-administered questioning would still produce a biased answer.
Frequently Asked Questions
What is an example of social desirability bias?
A common textbook example is self-reported voting: in post-election surveys, the number of respondents who say they voted routinely exceeds the number of ballots actually cast, because voting is viewed as a civic good and non-voters under-report not voting. Self-reported physical activity, exceeding device-measured activity, is another frequently cited example.
How do you control for social desirability bias?
Either statistically, by administering a social desirability scale (Marlowe-Crowne or BIDR) alongside the substantive measure and controlling for the score in analysis, or by design, using one of the countermeasures in the table above — self-administration, indirect questioning, randomized response, a list experiment, or a forced-choice format — chosen according to how sensitive the topic is and whether an individual-level or only a group-level estimate is needed.
Does an anonymous survey eliminate social desirability bias?
No. Anonymity reduces impression management, the conscious, audience-directed component, but does little for self-deceptive enhancement, where the respondent isn’t aware their self-assessment is inflated. It also has to be credible to the respondent to have any effect; a nominally anonymous survey within a small, identifiable group often does not function as anonymous in the respondent’s judgment.
Is social desirability bias the same as response bias?
Social desirability bias is one specific type of response bias — the broader category covering any systematic distortion introduced by how respondents answer (as opposed to who is sampled). Acquiescence bias and extreme-response style are other response biases that are not driven by the normative content of the question the way social desirability bias is.







