Skip to main content
v2026.11,772 entries · CC-BY 4.0

Mouse Grimace Scale: Scoring, Validation, and Its Role in Humane Endpoint Determination

How the Mouse Grimace Scale scores post-procedural pain via five facial action units, and what an IACUC protocol needs to specify for humane endpoint determination and refinement.

Written and maintained by CASRAI Editorial Board

Last updated

The Mouse Grimace Scale (MGS) is a validated, standardized tool for scoring post-procedural and post-surgical pain in laboratory mice from five facial action units. Developed by Langford et al. and published in Nature Methods in 2010, it gave researchers and animal-care staff a way to assess spontaneous pain without relying on evoked-response tests such as von Frey filaments, which measure sensory threshold rather than the animal’s own experience of pain. For an IACUC-regulated program, MGS matters less as a laboratory technique than as a documented, reproducible refinement measure and a source of objective evidence for humane endpoint decisions — the framing this guide takes throughout.

What the Mouse Grimace Scale measures

MGS scores five facial action units, each coded on a 0–1–2 scale (0 = not present, 1 = moderately present, 2 = obviously present):

  • Orbital tightening — narrowing of the orbital area, a squint-like closing of the eye.
  • Nose bulge — a bulge across the bridge of the nose.
  • Cheek bulge — a rounded bulge of the cheek muscle.
  • Ear position — ears pulled back and/or rotated away from the typical resting orientation.
  • Whisker change — whiskers pulled back against the face, or, alternatively, held forward in a clumped bundle rather than their normal splayed resting position.

A composite MGS score is typically the mean of the five action-unit scores. Higher composite scores correspond to more evident pain-related facial changes. The scale was validated across multiple noxious stimuli and procedures — including laparotomy and intraplantar inflammatory-agent injection — and the original validation work also demonstrated that MGS scores dropped with administration of an appropriate analgesic, supporting its use as an outcome measure for analgesic efficacy, not only as a static severity snapshot.

How MGS is scored in practice

Standard MGS coding is done from still images or short video clips of the unrestrained mouse’s face, captured at a defined interval after a procedure, rather than live in real time. This matters for reproducibility: a still image lets a scorer examine a facial expression at leisure and lets a second, blinded scorer independently re-code the same frame, which is the basis of any reported inter-rater reliability figure. Facilities that score live, at the cage side, are trading some of that reproducibility for speed and lower observer burden — a legitimate operational choice, but one a protocol should state explicitly rather than leave implicit, since it changes what the resulting score can be compared against.

Practical coding conventions worth specifying in an SOP or protocol appendix:

  • Define the observation window (e.g., a fixed number of minutes post-procedure, and at what follow-up intervals) rather than leaving timing to observer discretion.
  • Score each of the five action units independently before summing or averaging, rather than forming a single gestalt “pain impression” first.
  • Have a second, blinded observer re-score a subset of images to establish and document inter-rater agreement for that facility’s own scorers, rather than assuming the original validation study’s reliability transfers automatically to a new team, species substrain, or camera setup.

MGS in IACUC protocols and humane endpoint determination

On a research-administration site, the operationally important question is not how to photograph a mouse’s face but what a protocol needs to say about doing so. An IACUC reviewing a protocol that proposes MGS as a monitoring tool should expect the submission to specify:

  • Justification — why facial-expression scoring is the appropriate welfare measure for this specific procedure, and what it adds beyond body-weight loss, activity level, and other standard clinical scoring-sheet parameters already required under the Guide for the Care and Use of Laboratory Animals.
  • Scoring thresholds tied to action — at what composite MGS score (or pattern of individual action-unit scores) analgesia is administered, escalated, or a humane endpoint is triggered, decided in advance rather than left to real-time judgment calls.
  • Personnel qualification — who is trained to score MGS reliably, how that training was verified, and how new scorers are brought up to an acceptable agreement level before they score independently. Species- and technique-specific training expectations are the same kind documented more generally in IACUC training requirements.
  • Refinement steps taken as a result — what changes to analgesic regimen, monitoring frequency, or procedure timing followed from MGS findings during protocol development or pilot work, if any.

These are the kind of specifics an IACUC protocol gets sent back for omitting, and the same category of detail the IACUC’s broader oversight role exists to enforce — a monitoring tool named in a protocol without a decision rule attached to it is not yet a humane endpoint plan.

Where MGS fits in the 3Rs

MGS is a Refinement tool, not a Replacement or Reduction one. It does not reduce the number of animals used or replace an animal model with a non-animal alternative; it refines the procedure and its post-procedural care by giving investigators and veterinary staff an objective, standardized signal of pain severity that can inform earlier or more targeted analgesic intervention, and by making it possible to identify a genuine humane endpoint before a study’s pre-registered clinical scoring-sheet thresholds are reached on cruder measures alone. The broader 3Rs framework — Replacement, Reduction, Refinement — and how each applies across animal-research protocol review is covered in full in Animal Research Ethics: The 3Rs, IACUC Oversight, and the Law.

A complement to welfare monitoring, not a replacement for it

MGS is designed to sit alongside existing welfare-monitoring parameters, not substitute for them. Body-weight trend, activity and posture, coat condition, food and water intake, and any procedure-specific clinical signs remain part of a complete monitoring plan. Facial-expression scoring adds a dimension those measures do not capture well on their own — a mouse can maintain body weight and normal activity while still showing facial signs consistent with pain, particularly in the acute post-surgical window before weight loss or reduced activity would otherwise register. Conversely, MGS alone should not be the sole trigger for a humane endpoint decision: it is one input into a scoring sheet that should weight multiple independent signals, consistent with routine husbandry and health-monitoring practice covered in Mouse Husbandry: Housing, Care, and Compliance Standards and, for colony-level context, Sentinel Mice: How Rodent Colony Health Monitoring Works.

MGS-informed humane endpoint decisions are also directly relevant to procedure-specific justification. A protocol proposing retro-orbital injection or specifying a particular route of administration can use MGS data to support the stated refinement rationale, and a protocol’s euthanasia method — for example cervical dislocation as a humane endpoint procedure — should be tied to the same decision criteria the monitoring plan defines.

Blinding, randomization, and ARRIVE 2.0

Because MGS scoring involves a degree of subjective judgment, bias-control measures matter as much as they do for any other outcome measure in a preclinical study. The ARRIVE 2.0 guidelines for reporting animal research call for the experimental groups to be allocated randomly and for outcome assessors — including anyone scoring MGS images — to be blinded to treatment group. In practice, this means MGS images or clips should be coded with a non-identifying label, scored by someone who did not perform the procedure and does not know which treatment arm the animal was in, and only unblinded after scoring is complete. A protocol or manuscript that describes MGS scoring without describing how blinding and randomization were implemented has not fully met ARRIVE 2.0’s design-reporting expectations, and reviewers should treat that omission the same way they would treat it for any other subjectively-scored outcome.

Validated uses and known limitations

Strain differences

Baseline facial expression and the amplitude of MGS-scored responses are not uniform across mouse strains and substrains. A scoring threshold or “normal resting face” calibrated on one commonly used strain should not be assumed to transfer directly to a genetically distinct line without local verification, particularly for genetically engineered lines with altered baseline behavior or morphology. Protocols working with less-studied strains should note this as a limitation of any MGS-based threshold they propose, rather than importing published cutoffs unexamined.

Video/photograph scoring versus live scoring

As above, still-image or video-based scoring is the better-characterized, more reproducible method and the one the original validation work relied on; live cage-side scoring trades some of that rigor for lower staff burden and faster turnaround. Neither is categorically wrong, but a protocol or publication should state which was used, since the two are not interchangeable for reliability purposes.

Inter-rater reliability

MGS was developed and validated specifically to be scorable with good agreement between independent, minimally trained observers, and that reliability is part of what distinguishes it from a purely subjective “does this animal look painful” judgment. Reliability is not, however, a fixed property of the scale in the abstract — it depends on image quality, consistent observation timing, and scorer training, which is why establishing local inter-rater agreement (see the scoring-in-practice section above) rather than assuming a published figure applies unchanged is the more defensible practice for a given facility’s own program.

What MGS does not do

MGS scores facial action units associated with pain; it is not a general sickness, distress, or welfare index, and a low MGS score does not by itself rule out non-pain-related suffering (e.g., from illness, isolation, or a non-painful but distressing manipulation). It is validated for spontaneous, ongoing pain assessment, not as a screening tool for evoked/reflexive pain thresholds, which remain the province of tests such as von Frey filament assessment. Treating MGS as a complete welfare-monitoring solution on its own, rather than one component of a multi-parameter scoring sheet, is a misuse of what the tool was validated to do.

Frequently asked questions

Is the Mouse Grimace Scale required by an IACUC?

No single facial-expression scale is federally mandated. What is expected, per the Guide for the Care and Use of Laboratory Animals, is that a protocol define objective, procedure-appropriate criteria for pain monitoring and humane endpoints; MGS is one validated way to meet that expectation for procedures with an expected acute-pain component, not the only acceptable method.

Does MGS replace the need for a clinical scoring sheet?

No. MGS is designed to be one parameter within a broader scoring sheet that also tracks body weight, activity, posture, and procedure-specific clinical signs, not a standalone replacement for that sheet.

Can MGS be automated?

Machine-learning-based automated facial-expression scoring for mice is an active area of methods development, but manual, trained-observer scoring against the original five-action-unit rubric remains the well-characterized, validated reference method a protocol should default to describing unless an automated method has been independently validated for the specific strain, procedure, and imaging setup in use.

Does a rat equivalent exist?

Yes — the same research group subsequently developed and validated a Rat Grimace Scale using an analogous facial action-unit approach, published separately from the original mouse work.

Sources

Langford DJ, Bailey AL, Chanda ML, et al. “Coding of facial expressions of pain in the laboratory mouse.” Nature Methods. 2010;7(6):447–449. doi:10.1038/nmeth.1455.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about Mouse Grimace Scale: Scoring, Validation, and Its Role in Humane Endpoint Determination

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.