Skip to main content
v2026.11,772 entries · CC-BY 4.0

Elevated Plus Maze Testing: IACUC Requirements and Confounds

What an IACUC protocol must specify for elevated plus maze testing — apparatus design and session parameters, the lighting, time-of-day, and test-order confounds specific to this assay, and ARRIVE 2.0 blinding/randomization.

Written and maintained by CASRAI Editorial Board

Last updated

The elevated plus maze (EPM) is one of the most widely used unconditioned behavioral assays for anxiety-like behavior in rodents, built on a rat’s or mouse’s natural conflict between exploring novel space and avoiding open, elevated areas. Because the procedure itself looks trivial — a few minutes of free exploration on a simple apparatus, no injections, no surgery — it is easy to under-specify in an Institutional Animal Care and Use Committee (IACUC) protocol. The EPM literature is unusually candid about the opposite problem: this is one of the more procedurally fragile assays in the standard behavioral battery, and a protocol that does not pin down apparatus lighting, testing time, handling history, and scoring method is not fully specifying the experiment. This guide covers what a protocol needs to state, the confounds most likely to invalidate the dataset, and the blinding and randomization practices ARRIVE 2.0 expects to see documented.

What the Test Measures, and Why Its Simplicity Is Deceptive

In the standard task, first validated for rats by Pellow, Chopin, File and Briley in 1985, an animal is placed at the center of a plus-shaped apparatus elevated roughly 50 cm above the floor, with two opposing arms enclosed by walls (typically around 40 cm high) and two opposing arms left open, and allowed to explore freely for a single trial, conventionally five minutes. Anxiety-like behavior is read from the balance of exploration between arm types: time spent in the open arms and the number of entries into them, each usually expressed as a proportion of total arm time or total entries so that general locomotor activity — indexed separately by closed-arm entries or total entries — does not get conflated with the anxiety-like measure itself. A lower proportion of open-arm time and entries is interpreted as more anxiety-like behavior; classical anxiolytics such as benzodiazepines reliably increase it, which is the pharmacological validation the assay rests on.

The apparent simplicity is misleading. Because the entire readout depends on how aversive the open arms feel relative to the closed arms on that specific day, in that specific room, under that specific light, the EPM is unusually sensitive to procedural variables that carry little weight in a test like rotarod or grip strength. That sensitivity is exactly why the confounds section below is not optional background reading for an investigator or an IACUC reviewer — it is the difference between a replicable dataset and noise dressed up as a finding.

What an IACUC Protocol Must Specify

Scientific Justification and the 3Rs

Every EPM protocol needs an explicit 3Rs justification. Replacement is generally unavailable for this endpoint — anxiety-like behavior in an intact, freely moving animal has no validated in vitro or computational substitute — though the protocol should still document that alternatives were considered. Reduction depends on an a priori power calculation rather than a group size inherited from a prior paper, and on controlling the confounds below well enough that data are not later discarded to unexplained variability. Refinement is where most of the protocol’s welfare content lives: apparatus fall protection, session duration limits, monitoring during the trial, and the handling procedure that precedes testing, covered next.

Apparatus Design and Session Parameters

The protocol should state the apparatus’s exact configuration rather than gesture at “a standard elevated plus maze”: arm length and width, closed-arm wall height, elevation above the floor, and whether the open arms carry a small raised ledge (a common refinement that reduces fall risk without materially changing arm aversiveness). It should state trial duration — five minutes is conventional, and any deviation needs a rationale — and whether the design calls for a single trial per animal or repeated testing across days. Repeated EPM testing of the same animal is a substantive methodological choice, not a convenience, for reasons covered under test-order effects below, and a protocol proposing it should justify why against that literature rather than defaulting to it.

Fall Protection and In-Trial Monitoring

Because animals are tested on an elevated, partially unwalled apparatus, the protocol should specify fall-mitigation measures — padding, netting, or a foam floor beneath the maze — and who monitors the animal live or via video during the trial. It should state the criteria for early removal: a fall, prolonged freezing, repeated escape attempts, or overt distress beyond the mild aversive response the task is designed to elicit. Fecal boli count, a common ancillary autonomic measure of stress reactivity, should be scored consistently if reported at all, not noted only when it fits the narrative.

Personnel Qualification

Handlers and scorers need documented training in rodent handling, in recognizing the early-removal criteria above, and in the specific scoring method — live observation, video coding, or automated tracking — used for that protocol. This is the same personnel-competency standard the Guide for the Care and Use of Laboratory Animals applies to every animal procedure, not a lower bar because the task involves no injection or incision.

Confounds That Threaten Validity

These are not abstract methodological concerns; each has been identified repeatedly in the behavioral-pharmacology literature as a cause of failed replication, and each is something an IACUC protocol can and should pin down in writing before data collection starts.

Ambient Lighting Level

Lighting is the single most consequential procedural variable in EPM methodology. The open arms are aversive largely because they are exposed; illumination level directly modulates how aversive that exposure feels. Brighter lighting increases baseline anxiety-like behavior and can produce a floor effect where control animals already avoid the open arms almost completely, leaving no room to detect an anxiogenic manipulation; very dim lighting can flatten the open/closed preference altogether, making the assay insensitive to an anxiolytic treatment it should otherwise detect. A protocol should specify illuminance in lux measured at floor level on both the open and closed arms specifically, not just ambient room lighting — the apparatus’s own geometry means an overhead light typically illuminates the open arms far more than the walled closed arms even when “room lighting” is nominally constant, and that within-apparatus lighting differential is itself part of what the test measures.

Time of Day and Circadian Phase

Rodents are predominantly nocturnal, and baseline arousal and stress reactivity shift across the light/dark cycle. Testing the same cohort at inconsistent times of day introduces variance unrelated to treatment, and testing different groups at systematically different times can produce a group difference that is really a time-of-day effect. The protocol should fix a testing window and hold it constant across the full cohort, and should record actual testing times so a reviewer or reader can assess whether the window was genuinely held constant in practice, not just declared as intended.

Prior Handling and Test-Order Effects: One-Trial Tolerance

File’s 1990 description of “one-trial tolerance” in the plus-maze is a specific, well-documented finding: an animal re-exposed to the EPM after an initial trial shows a blunted response to benzodiazepine anxiolytics on the second exposure, and the nature of open-arm avoidance itself changes with prior maze experience. In practice this means the EPM behaves less like a repeatable outcome measure and more like a single-use one — a protocol proposing repeated testing of the same animals needs to address this directly rather than treat retest data as equivalent to first-exposure data. The same logic extends to handling and test-battery order more broadly: an animal handled roughly, moved through several other procedures, or tested on the EPM after a different stressful assay carries a prior stress or novelty history into the maze that a naive control does not. The protocol should specify pre-test handling (commonly several days of brief, gentle handling to reduce novelty-driven variability) and, where the EPM sits inside a larger behavioral battery, its position in that order and the rationale for it.

Inter-Laboratory Reproducibility

The EPM’s sensitivity to the variables above is not a theoretical worry; it is the practical reason the assay has a documented history of poor reproducibility across laboratories even when investigators believe they are running “the standard protocol.” The best-known demonstration of this problem in mouse behavioral genetics generally — Crabbe, Wahlsten and Dudek’s 1999 study in Science, “Genetics of mouse behavior: interactions with laboratory environment” — tested several inbred strains across multiple independently run laboratories that had gone to considerable lengths to standardize equipment, protocols, and even the animals’ shipment and handling, and still found significant, sometimes strain-dependent, site-specific differences in behavioral outcomes including anxiety-related measures. EPM-specific methodological reviews point to the same under-specified variables — lighting, apparatus dimensions, handling, scoring criteria — as the likely drivers, not some irreducible property of the test itself. The practical implication is the same for a protocol and any eventual methods section: report exact apparatus dimensions, measured lighting, testing time window, handling procedure, and scoring method, on the assumption that “standard EPM” is not specific enough for another lab to actually reproduce.

Blinding and Randomization per ARRIVE 2.0

The ARRIVE 2.0 Essential 10 lists randomization and blinding/masking among the minimum items every animal study should report, and EPM studies are a common place for both to go undocumented. Randomization should cover treatment-group allocation and, given the time-of-day confound above, the order in which groups are tested across the session — testing an entire treatment group consecutively before switching to the next group confounds treatment with time of day, so allocation order should be counterbalanced or randomized across the testing window rather than batched by group. Blinding should extend to whoever scores the trial, whether by live observation or video: the scorer should not know group assignment, and where automated tracking software defines the open- and closed-arm regions of interest, those region boundaries should be set identically across all animals and finalized before group identity is unblinded. See the full ARRIVE 2.0 checklist entry for how these items map onto the rest of a methods write-up.

Documenting the Protocol for IACUC Review

Pulling the above together, a submission-ready EPM protocol section should state: the scientific justification and 3Rs analysis; exact apparatus dimensions, elevation, and ledge presence; measured lighting (lux) at open- and closed-arm floor level; trial duration and single-exposure vs. repeated testing, justified either way; the fixed testing time window and fall-protection/monitoring measures; pre-test handling and, if applicable, the EPM’s position within a larger test battery; personnel and their training record; and the randomization/blinding scheme covering group allocation, testing order, and scoring. This level of specificity is what PHS Policy and, where the species fall within its scope, the Animal Welfare Act expect an IACUC to review as part of a complete protocol document, not a general procedure description filled in after the fact. A behavioral assay with a comparable set of explicit welfare and confound controls to compare against is the Morris water maze, which shares much of the same IACUC review logic for a different, hippocampal-dependent endpoint.

Frequently Asked Questions

Is the elevated plus maze a survival procedure for IACUC classification?

Yes, in the overwhelming majority of protocols it is a survival, non-terminal behavioral procedure — the animal returns to its home cage after the trial. Pain/distress categorization should still reflect the mild aversive stress the task deliberately induces, even though there is no surgery, injection, or euthanasia endpoint built into the test itself.

What lighting level should a protocol specify?

There is no single universal lux value that applies across every strain and species, which is exactly why the protocol needs to state its own value rather than rely on an unstated default. What matters is that illuminance is measured at arm-floor level on both open and closed arms, held constant across the full cohort, and reported — because both very bright and very dim lighting can distort or flatten the open/closed arm preference the test depends on.

Can the same animal be tested on the EPM more than once?

It can, but the “one-trial tolerance” phenomenon means a second exposure is not behaviorally equivalent to the first — prior maze experience blunts anxiolytic drug effects and changes the animal’s baseline response to the apparatus. A protocol proposing repeated testing should justify that choice explicitly rather than treat retest data as interchangeable with first-exposure data.

Why is the elevated plus maze considered to have poor reproducibility across laboratories?

Its outcome depends on how aversive the open arms feel relative to the closed arms on a given day — highly sensitive to lighting, handling history, testing time, and apparatus specifics, variables easy to under-report as “standard protocol.” Landmark multi-laboratory work in mouse behavioral genetics found significant site-specific differences even under deliberately standardized conditions, and EPM-specific reviews point to the same under-specified variables as the likely cause.

Does the elevated plus maze require blinding under ARRIVE 2.0?

Yes. ARRIVE 2.0’s Essential 10 lists blinding/masking as a minimum reporting item for animal studies generally, and it applies directly here: whoever scores open-arm time and entries, live or from video, should not have visible access to group identity while scoring, and automated-tracking region definitions should be finalized before unblinding.

Primary Sources

  • Pellow, S., Chopin, P., File, S.E. & Briley, M. “Validation of open:closed arm entries in an elevated plus-maze as a measure of anxiety in the rat.” Journal of Neuroscience Methods 14(3), 149–167 (1985).
  • Walf, A.A. & Frye, C.A. “The use of the elevated plus maze as an assay of anxiety-related behavior in rodents.” Nature Protocols 2(2), 322–328 (2007).
  • File, S.E. “One-trial tolerance to the anxiolytic effects of chlordiazepoxide in the plus-maze.” Psychopharmacology (1990).
  • Crabbe, J.C., Wahlsten, D. & Dudek, B.C. “Genetics of mouse behavior: interactions with laboratory environment.” Science 284(5420), 1670–1672 (1999).
  • Carobrez, A.P. & Bertoglio, L.J. “Ethological and temporal analyses of anxiety-like behavior: the elevated plus-maze model 20 years on.” Neuroscience & Biobehavioral Reviews 29(8), 1193–1205 (2005).
  • NC3Rs / ARRIVE 2.0 guidelines (Percie du Sert et al., 2020) — Essential 10 reporting items.
  • National Research Council. Guide for the Care and Use of Laboratory Animals, 8th edition.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about Elevated Plus Maze Testing: IACUC Requirements and Confounds

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.