Skip to main content
v2026.11,610 entries · CC-BY 4.0

The Hawthorne Effect: Why Being Watched Changes the Data

The Hawthorne effect is the tendency for people to change behavior because they know they are being studied. Learn the Western Electric study history, what the Levitt & List and Jones re-analyses actually found, and how to detect and design against reactivity.

Ask about The Hawthorne Effect: Why Being Watched Changes the Data

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

The Hawthorne effect is the tendency for people to change their behavior — usually improving performance — simply because they know they are being observed or studied, independent of whatever variable the researcher is actually manipulating. It is named after a set of productivity studies conducted at Western Electric’s Hawthorne Works plant near Chicago between 1924 and 1932. The name has become the default shorthand for reactivity in research design, but the study that gave it that name is a much weaker piece of evidence than its ubiquity in textbooks suggests: the two serious modern re-analyses of the original data find little support for the dramatic “any change increased output” version of the story that is still widely repeated. That gap between reputation and evidence is itself the most useful thing to understand about the Hawthorne effect before you try to design around it.

This page covers the original studies, what happened when economists and sociologists went back and re-analyzed the raw data decades later, and — because the label is invoked constantly in study design and grant review — a practical section on detecting and designing against reactivity in your own work. For the broader family of related phenomena (demand characteristics, social desirability bias, observer/experimenter bias, evaluation apprehension), see the observer effect guide, which the Hawthorne effect is one specific, named instance of.

The Western Electric studies: what was actually done

The Hawthorne Works was a Western Electric manufacturing plant employing around 30,000 workers. Between 1924 and 1932 it was the site of a series of separate, sequential studies commissioned first by the plant’s own engineers and the National Research Council, and later taken over and reinterpreted by a Harvard Business School research team led by Elton Mayo, with Fritz Roethlisberger and William Dickson doing much of the on-site data collection and later co-authoring the studies’ canonical published account. The work is usually described in four phases:

  • The illumination experiments (1924–1927). Engineers varied lighting levels in selected production departments and compared output against unchanged control departments, trying to find a lighting level that maximized productivity. This is the phase the popular “Hawthorne effect” story is actually about, and the phase both modern re-analyses (below) went back and re-examined.
  • The Relay Assembly Test Room (1927–1932). A small group of workers assembling telephone relays was moved to a separate room and studied under a sequence of changes to rest breaks, working hours, and incentive pay, with researchers present and workers aware of the changes.
  • The Bank Wiring Observation Room (1931–1932). A different small group was simply observed, without experimental manipulation, to study informal group norms around output restriction — this phase is the source of much of Mayo’s later writing on informal social organization in the workplace, rather than reactivity specifically.
  • The interviewing program (roughly 1928–1930). Researchers conducted tens of thousands of employee interviews, originally structured and later unstructured, that became an early influence on non-directive interviewing technique in organizational research.

The term “Hawthorne effect” itself is not from the original 1930s studies at all. It was coined later — the label is usually credited to sociologist Henry A. Landsberger, who used it in his 1958 book reassessing the studies, building on an informal use by John R. P. French a few years earlier. The researchers who ran the original experiments did not describe what they observed with that name.

What the re-analyses actually showed

For decades the illumination studies were assumed to be exactly as described in secondary and textbook accounts, in part because the original raw data were thought to be lost. Two later re-analyses changed that picture substantially.

Stephen R. G. Jones (1992) examined the relay-assembly-test-room records in a paper published in the American Journal of Sociology and found that once the introduction of a new incentive pay scheme and ordinary learning-curve effects over the course of the study were accounted for, there was very little productivity variation left to attribute to “being observed” as a mechanism in its own right.

Steven Levitt and John List (2011) located the original illumination-study data — the records long assumed destroyed — and published a formal re-analysis, “Was There Really a Hawthorne Effect at the Hawthorne Plant? An Analysis of the Original Illumination Experiments,” in the American Economic Journal: Applied Economics (an earlier version circulated as NBER Working Paper No. 15016). Their central finding: output did not track the lighting manipulations in the clean, dramatic pattern usually described. Instead, productivity tracked the day of the week (dips on Sundays and Mondays, rises toward the end of the work week) and the pay-period cycle far more closely than it tracked when lighting was changed. They did find weaker, more circumscribed evidence consistent with some reactivity — output was somewhat higher while a manipulation was actively in progress than during gaps between manipulations — but nothing resembling the “productivity rose after every single change, including a return to baseline” story most textbooks still repeat.

Neither re-analysis argues that reactivity isn’t real. It plainly is, and it is documented far more cleanly in better-controlled modern studies designed specifically to isolate it (again, see the observer effect guide for that literature). What the re-analyses undercut is specifically the idea that the Hawthorne Works illumination study is strong evidence for it. The honest position, and the one worth taking on a study proposal or in a methods section, is to use “Hawthorne effect” as the familiar name for a real phenomenon while being explicit that its namesake study is weaker evidence for it than its fame implies.

Why this matters for measurement and study design

Whatever you call it, the underlying threat is a specific and well-defined one: the act of measuring can change the value being measured, and it changes it in a way that is correlated with the fact of measurement itself rather than with your independent variable. That makes it a threat to two things at once:

  • Construct validity — the behavior you record under observation may not be the behavior that would occur unobserved, so you may not be measuring the construct you think you’re measuring. See types of validity in research for how this sits alongside content, criterion, and other validity threats.
  • Internal validity — if awareness of being studied differs systematically between your treatment and control conditions (for example, a treatment group that gets more attention, more visits, or more contact with researchers than the control group), reactivity becomes a confound that can masquerade as a treatment effect. See experimental design for how validity threats are generally categorized.

Detecting and designing against it

None of the strategies below eliminate reactivity outright; each trades off feasibility, cost, and what it can and can’t rule out. Most well-designed studies combine two or more of them rather than relying on a single fix.

Strategy What it does Best used for Limitation
Blinding / masking Keeps participants, staff, or assessors unaware of condition assignment, so awareness of being studied doesn’t correlate with which condition someone is in. Trials and experiments where an active intervention can plausibly be masked (placebo pill, sham procedure, blinded outcome assessment). Often impossible for behavioral, educational, or organizational interventions where the intervention itself is visible to participants. See blinding and masking for the fuller mechanics.
Equal-attention / attention-control arm Gives the control group contact, visits, or researcher attention matched to the treatment group, so any reactivity effect is shared across arms rather than concentrated in one. Any design where the treatment group would otherwise get more researcher contact than the control group. Designing a credible attention-matched control can be as much work as designing the intervention itself.
Unobtrusive / archival measures Uses data that already exists independent of the study — administrative records, transaction logs, sensor data, existing performance metrics — rather than measures collected by a visibly present observer. Outcomes with a clean administrative trace: output records, attendance, sales, usage logs, prescribing data. You’re limited to whatever was already being recorded; can’t add new measures without reintroducing observation.
Run-in / habituation period Introduces the observation apparatus (monitoring, recording, researcher presence) before the intervention or data-collection period actually begins, so participants habituate to being watched before the measurement that counts. Field studies and workplace/organizational research where instrumentation itself (cameras, loggers, an embedded observer) is the reactivity trigger, not the treatment. Adds time and cost; habituation is rarely complete, so it reduces rather than eliminates the effect.
Objective vs. self-report outcomes Prefers outcomes that are harder to consciously adjust (biomarkers, administrative records, third-party ratings) over self-reported behavior or attitudes. Any study where a self-report measure and a behavioral or administrative proxy for the same construct both exist. Objective proxies aren’t always available or don’t always capture the construct of interest.
Statistical control for time trends Explicitly models day-of-week, pay-period, seasonal, or learning-curve effects as covariates, rather than attributing all pre/post change to the intervention. Any before-after or interrupted time-series design — directly motivated by what the Levitt & List and Jones re-analyses found was actually driving the original Hawthorne data. Requires enough repeated observations to separate a time trend from a treatment effect; underpowered short studies can’t do this reliably.
Randomization with a genuine control group Ensures that whatever reactivity exists is, on average, distributed equally across arms, so it cancels out of the treatment-versus-control comparison even if it can’t be removed. Any design where random assignment to condition is feasible. Only controls for reactivity that’s unrelated to assignment; doesn’t help if the treatment itself is what triggers extra attention or awareness.

A useful design habit: ask, for any effect you’re about to report, whether it could instead be explained by when the measurement happened (day of week, point in a pay or academic cycle, position in a study timeline) rather than by what was manipulated. That is precisely the check that the original Hawthorne illumination data failed to survive, and it’s a cheap, mechanical thing to build into an analysis plan before data collection starts — see control group design for how a well-specified control group helps rule this out. Where an intervention involves human judgment or self-report, pairing it with an established reliable instrument (see Cronbach’s alpha for internal-consistency reliability) helps distinguish real change from noise in how a construct is being measured, which is a related but separate concern from reactivity itself — see accuracy vs. precision in measurement for that distinction.

Frequently asked questions

Is the Hawthorne effect the same thing as the observer effect?

Not exactly. “Observer effect” (or reactivity) is the broader umbrella term for any behavior change caused by awareness of being studied. “Hawthorne effect” is the specific, historically named instance of it, tied to the Western Electric studies. In practice the two terms are used almost interchangeably, but “Hawthorne effect” carries the specific (and now contested) historical claim about what happened at that plant, while “observer effect” or “reactivity” doesn’t commit you to that history. See the observer effect guide for the full family of related terms, including demand characteristics and social desirability bias.

Was the Hawthorne effect debunked?

Not entirely, but its evidentiary basis was substantially weakened. The phenomenon it names — people changing behavior because they know they’re being studied — is real and well documented elsewhere. What the Jones (1992) and Levitt & List (2011) re-analyses undercut is the specific claim that the original Hawthorne illumination experiments demonstrate it cleanly; the re-analyzed data track day-of-week and pay-period patterns better than they track the lighting changes themselves.

Can randomization alone control for the Hawthorne effect?

Randomization controls for it only to the extent that reactivity is distributed equally across arms. If the treatment itself is what triggers extra attention, visits, or awareness — which is common in behavioral, educational, and workplace interventions — randomization doesn’t remove the confound between “receiving the treatment” and “receiving more attention.” An attention-matched control arm (see the table above) is usually needed alongside randomization in those cases.

Does the Hawthorne effect apply outside factory or workplace settings?

Yes. The name comes from a workplace study, but the underlying reactivity mechanism applies anywhere a person knows they’re being measured, monitored, or observed — clinical trials, classroom research, usability testing, wearable-device studies, and survey research all show documented reactivity effects, generally studied and reported under the “reactivity,” “demand characteristics,” or “observer effect” labels rather than “Hawthorne effect” specifically.

Related reading

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →