Skip to main content
v2026.11,610 entries · CC-BY 4.0
Dictionary termTrack CStablev2026.1

Observer Bias

Observer bias is a systematic distortion introduced by the person doing the observing, recording or interpreting, rather than by the participants or the sampling. The Catalogue of Bias defines it as a 'systematic difference between a true value and the value actually observed due to observer variation'. It arises wherever measurement requires judgement - reading an image, scoring a scale, coding an interview transcript, grading a lesion, deciding whether an adverse event is related - and it is strongest where the observer knows which group a participant belongs to and has an expectation about what they should find. It is a form of detection or measurement bias and belongs to the recording end of a study, not the sampling end: it is distinct from selection bias, which distorts who enters the study, and from confounding, which distorts the comparison itself. It is also distinct from the observer effect, where the participant changes behaviour because they know they are being watched; in observer bias the participant behaves normally and it is the observer's record of that behaviour that is wrong. Observer bias is controlled at the design and procedure level - by blinding outcome assessors, separating the people who know exposure status from the people who record outcomes, pre-specifying operational definitions and coding rules, using independent duplicate coding, and quantifying agreement between observers - rather than by any statistical adjustment after the fact.

ByCASRAI Editorial Board
· Last updated 23 Aug 2026

Ask about Observer Bias

Answers are drawn from this dictionary entry and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Examples

Worked examples

  • Is an instance

    An unblinded assessor reading a borderline blood pressure and resolving it toward the value they expect for that treatment arm.

  • Is an instance

    A physiotherapy trial where the treating clinician also scores the outcome scale, knowing which participants received the intervention.

  • Is an instance

    A qualitative researcher committed to a theoretical framing coding ambiguous interview passages as evidence for it.

  • Is an instance

    An investigator assessing adverse event causality who believes the study drug is safe and grades borderline events as unrelated.

  • Is an instance

    A histopathologist grading slides while aware of the clinical history and the expected diagnosis.

Counter-examples

Looks similar, but isn't

  • Not an instance

    Participants working harder because they know they are being observed - that is the observer effect, and the classic instance is the Hawthorne effect.

  • Not an instance

    A sample that over-represents volunteers who are healthier than the target population - that is selection bias, affecting who was studied rather than how they were recorded.

  • Not an instance

    An apparent association driven by an unmeasured third variable - that is confounding, and blinding the assessor does not touch it.

  • Not an instance

    Random disagreement between two equally unbiased raters - that is observer variation without systematic direction, and it lowers reliability rather than shifting the estimate.

Editorial commentary

What observer bias is

Observer bias is error introduced by the researcher, not by the participant. The Catalogue of Bias
defines it as a systematic difference between a true value and the value actually observed due to
observer variation
, and classifies it as a form of detection bias affecting both observational and
interventional studies. It appears wherever recording a measurement involves an act of judgement, and it is
worst where two conditions coincide: the measurement is subjective, and the observer knows which arm or
group the participant is in.

The mechanism is not usually dishonesty. It is the ordinary human tendency to resolve ambiguity in the
direction of what you expect. The Catalogue’s own worked example is blood pressure measurement, where
clinicians reading a mercury sphygmomanometer round to whole numbers, and where an observer holding a prior
expectation may adjust a borderline reading toward what they believe the value ought to be. Nothing about
that requires bad faith; it requires only a judgement call and an expectation.

Observer bias is not the observer effect

These two are routinely confused and they are opposites in mechanism. The distinction is worth stating
precisely because the remedies do not overlap.

  • Observer bias is a defect in the record. The participant behaves exactly as
    they would have anyway; the researcher’s expectations distort what gets written down or how it is
    interpreted. The person whose behaviour changed is the researcher.
  • The observer effect – see the observer effect in
    research
    – is a defect in the behaviour being measured. The participant changes what they do
    because they know they are being watched, so the record is faithful but the thing recorded is no longer
    representative. The person whose behaviour changed is the participant. The
    Hawthorne effect is the best-known named instance.

The practical consequence: blinding the outcome assessor fixes observer bias and does nothing about the
observer effect, because the participant still knows they are being observed. Conversely, unobtrusive
measurement, habituation periods or
naturalistic observation reduce the observer effect while
leaving observer bias completely untouched. A study can suffer from both at once, and the two need separate
mitigations in the protocol.

What the evidence shows about its size

Observer bias is one of the few methodological problems with direct empirical estimates of its magnitude,
from a series of systematic reviews that compared blinded and non-blinded assessors within the same
trials
– a design that isolates the assessor’s contribution rather than inferring it:

  • Hrobjartsson A, Thomsen ASS, Emanuelsson F, Tendal B, et al. Observer bias in randomised clinical
    trials with binary outcomes: systematic review of trials with both blinded and non-blinded outcome
    assessors.
    BMJ 2012;344:e1119.
  • Hrobjartsson A, Thomsen ASS, Emanuelsson F, Tendal B, et al. Observer bias in randomized clinical
    trials with measurement scale outcomes: a systematic review of trials with both blinded and nonblinded
    assessors.
    CMAJ 2013;185(4):E201-E211.
  • Hrobjartsson A, Thomsen ASS, Emanuelsson F, Tendal B, et al. Observer bias in randomized clinical
    trials with time-to-event outcomes: systematic review of trials with both blinded and non-blinded outcome
    assessors.
    International Journal of Epidemiology 2014;43(3):937-948.

The Catalogue of Bias summarises this body of work as showing non-blinded assessors exaggerating effects
by roughly 36% for binary outcomes, 68% for measurement-scale outcomes and 27% for time-to-event outcomes.
The ordering is the instructive part: the more room for judgement the outcome leaves, the larger the
distortion. Scale-based outcomes are the most vulnerable; hard, objectively-recorded endpoints such as death
from a registry are the least.

How it is prevented

All of the effective controls are procedural and are put in place before data collection begins. None of
them is a statistical correction, because there is no post-hoc adjustment for an observer who recorded the
wrong number.

Blind the outcome assessor

The single most effective control, and the one the empirical evidence above is built on. Blinding the
person who records or adjudicates the outcome is separable from blinding participants or treating
clinicians: in an unblindable intervention – surgery, physiotherapy, a behavioural programme – the
participant necessarily knows their allocation, but a central assessor reading a de-identified image or
transcript need not. See
blinding and masking in clinical trials.
Where full blinding is impossible, document why, and blind whatever component can be blinded.

Separate exposure knowledge from outcome recording

Where an assessor cannot be blinded to the intervention, they can often be kept away from the exposure
data, the baseline covariates, or the earlier assessments in a longitudinal series. Physically separating
the two datasets, rather than relying on the assessor to disregard what they have seen, is the reliable
implementation.

Pre-specify operational definitions and coding rules

Ambiguity is the raw material of observer bias. A coding frame or outcome definition written before the
data are seen, with explicit inclusion and exclusion rules and worked boundary cases, removes most of the
discretion that expectation would otherwise fill. In qualitative work this means a codebook fixed and
documented before coding, with any subsequent changes logged and applied retrospectively to already-coded
material.

Use independent duplicate coding, then measure agreement

Two observers coding independently, with disagreements resolved by a documented adjudication procedure
rather than by discussion until one yields, both reduces and exposes the problem. Agreement should then be
quantified and reported, not asserted – see
the intraclass correlation coefficient for
continuous measures, and test-retest vs
inter-rater reliability
for which form of reliability answers which question. Note the limitation: high
inter-rater agreement rules out random observer variation but does not rule out observer bias, because two
observers who share the same expectation can agree with each other and both be systematically wrong.

Train observers, including on their own priors

Training on the instrument, the recording procedure and the time windows is standard. The Catalogue of
Bias additionally recommends training observers to recognise their own predispositions and habits, and is
candid that complete elimination is unlikely – which is a reason to report the residual risk rather than to
claim it away.

Report it honestly

Where blinding of outcome assessment was not possible, say so in the limitations, name the outcomes
affected, and state what was done instead. Reporting guidelines for trials and observational studies expect
this, and an explicit statement is more credible than silence.

Where it appears outside clinical trials

Observer bias is not a trials-only problem. It occurs in structured observation and ethnographic
fieldwork, where the fieldworker’s framing shapes what is recorded as significant; in qualitative coding,
where a researcher invested in a theme finds it; in imaging, histopathology and any grading task; in
systematic review screening and data extraction, which is why duplicate independent extraction is standard;
and in adverse event causality assessment, where the assessor’s belief about the drug influences the
relatedness judgement. See
qualitative research methods and
data collection methods for the design context.

Related concepts and how they differ

  • Selection bias – distorts who is in the sample, not what was
    recorded about them.
  • Confounding – a third variable distorts the
    association; the measurement itself may be perfectly accurate.
  • Publication bias – operates on which studies become
    visible, after the observation is complete.
  • Construct validity and
    types of validity – observer bias is a threat to
    measurement validity specifically.
  • Reliability – observer variation is the
    component of unreliability that inter-rater statistics are designed to detect.

Sources

Catalogue of Bias Collaboration. Mahtani K, Spencer EA, Brassey J. Observer bias. In: Catalogue Of
Bias, 2017. Plus the three Hrobjartsson et al. systematic reviews cited above, whose bibliographic details
were verified against Crossref.

Also known as

observer expectancy in recording · assessor bias · detection bias (observer) · experimenter bias in measurement

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="Observer Bias"
      vocab-term-identifier="https://casrai.org/dictionary/term/observer-bias" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/observer-bias",
  "name": "Observer Bias",
  "identifier": "https://casrai.org/dictionary/term/observer-bias",
  "description": "Observer bias is a systematic distortion introduced by the person doing the observing, recording or interpreting, rather than by the participants or the sampling. The Catalogue of Bias defines it as a 'systematic difference between a true value and the value actually observed due to observer variation'. It arises wherever measurement requires judgement - reading an image, scoring a scale, coding an interview transcript, grading a lesion, deciding whether an adverse event is related - and it is strongest where the observer knows which group a participant belongs to and has an expectation about what they should find. It is a form of detection or measurement bias and belongs to the recording end of a study, not the sampling end: it is distinct from selection bias, which distorts who enters the study, and from confounding, which distorts the comparison itself. It is also distinct from the observer effect, where the participant changes behaviour because they know they are being watched; in observer bias the participant behaves normally and it is the observer's record of that behaviour that is wrong. Observer bias is controlled at the design and procedure level - by blinding outcome assessors, separating the people who know exposure status from the people who record outcomes, pre-specifying operational definitions and coding rules, using independent duplicate coding, and quantifying agreement between observers - rather than by any statistical adjustment after the fact.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/reproducibility#set",
  "url": "https://casrai.org/dictionary/term/observer-bias",
  "sameAs": [
    "observer expectancy in recording",
    "assessor bias",
    "detection bias (observer)",
    "experimenter bias in measurement"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "dateModified": "2026-08-23T08:12:03",
  "inLanguage": "en"
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →