Skip to main content
v2026.11,610 entries · CC-BY 4.0

Thurstone Scaling: Equal-Appearing Intervals and Judge Ratings

Thurstone scaling (equal-appearing intervals) uses a judge panel to rate statement favorability; the median becomes the scale value and the semi-interquartile range (Q) flags ambiguous items. It is Likert scaling’s direct historical predecessor — and the judge-panel labor it requires is exactly why Likert’s simpler respondent-rated method displaced it in practice.

Ask about Thurstone Scaling: Equal-Appearing Intervals and Judge Ratings

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

Thurstone scaling — formally the method of equal-appearing intervals — is the technique Louis Thurstone and Ernest Chave laid out in The Measurement of Attitude (1929) to turn a pool of opinion statements into a scale with genuine interval-level properties: a panel of judges rates where each statement sits, and the median of those judge ratings becomes the statement’s scale value. It is the direct historical predecessor of the Likert scale, and understanding why it was largely displaced by Likert’s simpler technique explains a real, still-relevant trade-off between measurement rigor and data-collection cost.

What equal-appearing intervals means

The core idea is to separate two different judgments that most attitude scales blur together: how favorable a statement is, versus how much a given respondent agrees with it. Thurstone’s method calibrates the first judgment once, using a panel of judges, before any respondent ever sees the instrument.

The standard procedure runs in five stages:

  1. Generate a large item pool. Researchers write far more candidate statements than the final scale will need — commonly on the order of 100+ statements spanning the full range of the attitude, from extremely unfavorable to extremely favorable.
  2. Recruit a judge panel. A panel of judges (Thurstone and Chave’s original studies used panels in the hundreds; methods texts describe workable panels well below that, though more judges narrow the standard error of each scale value) sorts every statement — not by whether they personally agree with it, but by where it belongs.
  3. Sort into eleven equal-appearing categories. Each judge places every statement into one of eleven ordered piles, from 1 (“extremely unfavorable”) to 11 (“extremely favorable”), with the categories intended to look equally spaced to the judge — hence “equal-appearing intervals.” A judge who personally holds an extreme attitude is still asked to rate a statement’s position, not their agreement with it.
  4. Compute the scale value and the ambiguity index. For each statement, the median of all judges’ category placements becomes its scale value. Alongside it, the semi-interquartile range — conventionally written Q and computed as (Q3 − Q1)/2 from the same judge ratings — measures how much the judges disagreed about where the statement belongs. A low Q means the judges converged tightly on one position; a high Q flags a statement whose meaning is ambiguous or that plausibly belongs in more than one place on the scale.
  5. Select the final item set. Researchers keep statements with low Q, spread as evenly as possible across the 1–11 range, and discard the rest. Respondents then simply check every statement they personally agree with; a respondent’s attitude score is the median (or mean) of the scale values of the statements they endorsed.

Worked example: median scale value and Q from judge ratings

The table below is an illustrative composite, not data from any real study — invented to show the calculation, not to represent an actual attitude survey. Twenty judges rate two candidate statements for a scale on attitudes toward requiring data management plans (DMPs) in grant applications, sorting each into one of the eleven categories.

Category (1–11) 1 2 3 4 5 6 7 8 9 10 11
Item A — “DMPs should be required for all grant applications.” 0 0 0 0 0 1 2 5 7 4 1
Item B — “DMPs are bureaucratic paperwork with no real benefit.” 3 2 1 1 2 2 1 2 2 2 2

Using the standard grouped-data interpolation formula for the median and quartiles of ordinal category data (each category treated as an interval of width 1, e.g. category 8 spans 7.5–8.5), computed and verified with a standalone script rather than by hand:

  • Item A: median (scale value) = 8.79, Q1 = 7.90, Q3 = 9.50, so Q = 0.80. The 20 judges cluster tightly in the upper end of the scale — a low-Q, unambiguous item that a researcher would keep.
  • Item B: median (scale value) = 6.00, Q1 = 2.50, Q3 = 9.00, so Q = 3.25. Judges scattered across nearly the whole range — despite landing at a “neutral” median of 6, this item is highly ambiguous and would normally be discarded rather than kept just because its median looks central.

This is the mechanism that makes equal-appearing intervals more than a rebranded Likert scale: the median comes from judges rating the statement’s position, and Q gives researchers an explicit, quantified reason to reject an item that a simple face-value read (or a respondent-facing pilot test alone) would not surface.

The labor cost that made Likert scaling win out

Rensis Likert’s 1932 paper, A Technique for the Measurement of Attitudes, introduced what is now called summated ratings: respondents rate their own agreement with each statement directly (typically on a 5-point strongly-disagree-to-strongly-agree scale), and their total or average across items is their attitude score. No separate judge panel is involved at any stage.

That single difference is the specific labor cost that decided which method became the default in survey research. Building a Thurstone scale requires running an entire second study — recruiting a judge panel large enough to produce stable medians and Q values, having every judge sort every candidate statement, and then computing and screening scale values — before a single real respondent is ever surveyed. Likert’s method skips that step entirely: write the items, administer them, analyze the responses. For a research team without a standing item bank or a budget for a judge-rating study, that is the difference between one data collection and two.

Likert also showed, in the same paper, that summated-ratings scales achieved reliability comparable to Thurstone-scaled instruments in practice, which removed the main psychometric argument for paying the extra labor cost. What Likert scaling gives up in return is the interval-level claim: Thurstone scale values are calibrated against judge consensus and behave more defensibly as an interval scale, while Likert totals are ordinal sums that are conventionally analyzed as if they were interval data — a simplification, not a proof. See levels of measurement for what that distinction does and doesn’t license statistically.

Thurstone vs. Likert vs. Guttman: choosing among the classic scaling methods

Method Who rates what Item-development cost Measurement level claimed
Thurstone (equal-appearing intervals) Judge panel rates each statement’s position; respondents endorse statements they agree with High — a full judge-rating study precedes fielding Interval (calibrated via judge medians)
Likert (summated ratings) Respondents rate their own agreement directly Low — no judge panel required Ordinal, conventionally treated as approximately interval
Guttman (scalogram analysis) Respondents endorse items directly; items are checked for a cumulative structure after the fact Low to field, but requires items with a genuine cumulative (hierarchical) structure Ordinal, with a reproducibility check

The three methods answer different questions. A Likert scale asks how strongly someone agrees; Guttman scaling asks whether agreement with a hard item implies agreement with every easier item; Thurstone scaling asks judges to calibrate where each statement sits before anyone’s agreement is measured at all. In current practice, Thurstone scaling survives mainly in methodological pedagogy, in psychophysics-adjacent work using Thurstone’s related law of comparative judgment, and in the occasional applied study where a genuinely interval-level attitude measure is worth the judge-panel cost — not as a default choice for a routine questionnaire.

When Thurstone scaling is still the right tool

  • You need attitude scores that support interval-level statistics (means, standard deviations treated at face value) and can defend that claim rather than assume it, unlike a Likert sum.
  • You have access to a large, appropriately informed judge panel and the time to run a calibration study before fielding.
  • You are building a fixed, reusable item bank (e.g., a standardized attitude inventory intended for repeated use across studies), where the one-time judge-rating cost is amortized over many later administrations.
  • You need the Q ambiguity index specifically — a documented, quantified basis for excluding items that a purely face-valid item pool would not otherwise flag.

For a one-off survey with a normal budget and timeline, a Likert or semantic-differential instrument almost always wins on cost without a meaningful loss of practical validity, which is exactly the trade Likert’s 1932 paper made explicit.

Frequently asked questions

What does “equal-appearing intervals” actually mean?

It describes the judges’ sorting task, not a mathematical proof: judges are instructed to treat the eleven categories as if they were equally spaced, so that a one-category difference near the unfavorable end is assumed comparable in size to a one-category difference near the favorable end. The category boundaries are not independently verified as equal; “equal-appearing” is the method’s name for that assumption.

Why is the median used instead of the mean for the scale value?

The median is far less sensitive to a handful of outlying judges who badly misjudge where a statement belongs, which matters because judge panels are typically much smaller than respondent samples and a single extreme rating can distort a mean noticeably.

What is a “good” Q value?

There is no single universal cutoff in the original method — researchers compare Q across the full item pool and retain the statements with the lowest Q values relative to the rest of the pool, discarding the most ambiguous ones, rather than applying one fixed threshold across every study.

Do respondents ever see the judges’ scale values?

No. Respondents see only the final statements and simply check the ones they agree with; the numeric scale values exist purely for scoring after the fact and are never shown on the instrument itself.

Is Thurstone scaling the same as the law of comparative judgment?

No, though both are Thurstone’s work. Equal-appearing intervals has judges sort each statement independently into a category; the law of comparative judgment instead has judges make pairwise comparisons between statements, which is more labor-intensive still but avoids relying on judges’ absolute category placements. Equal-appearing intervals is the more commonly taught and applied of the two.

See also: Likert scale survey design, Guttman scaling, levels of measurement, questionnaire design, Cronbach’s alpha, and the research methods hub.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.