Skip to main content
v2026.11,610 entries · CC-BY 4.0

Guttman Scaling: Cumulative Items, Scalogram Analysis, and the Coefficient of Reproducibility

How Guttman scaling orders cumulative items, scores a scalogram, and uses the coefficient of reproducibility (and scalability) to test whether a scale is genuinely unidimensional, with a fully worked example.

Ask about Guttman Scaling: Cumulative Items, Scalogram Analysis, and the Coefficient of Reproducibility

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

Guttman scaling asks a narrower question than most survey-analysis techniques: not “how strongly does this respondent agree,” but “is agreement with a hard item evidence that this respondent also agrees with every easier item on the list.” When a set of items genuinely has that cumulative structure, a single number — the respondent’s count of endorsed items — reproduces the entire response pattern. That reproducibility, checked directly rather than assumed, is what the technique is for.

Cumulative items: the logic Guttman scaling depends on

A Guttman scale (Louis Guttman, “A Basis for Scaling Qualitative Data,” American Sociological Review, 1944) is built from items ordered along a single continuum by difficulty or intensity, such that endorsing a harder item logically implies endorsing every item easier than it. The classic worked domain is Emory Bogardus’s 1925 social-distance scale, which orders acceptance of an outgroup from close kinship through marriage down to exclusion from the country entirely — a respondent willing to accept marriage into the group is assumed willing to accept the weaker forms of contact too, and in a clean Guttman scale that assumption holds in the actual data, not just in the item design.

This is what separates a Guttman scale from a Likert scale. A Likert instrument sums or averages independent items that each measure the same construct at roughly the same intensity — there is no logical ordering between “I trust this institution” and “I trust this institution’s leadership,” and no expectation that agreeing with one implies agreeing with the other. A Guttman scale is ordinal by item design: the items themselves are ranked, and a respondent’s scale score is literally the number of items endorsed, because in a perfectly cumulative scale that count identifies exactly which items were endorsed. No averaging, no item weighting — the response pattern is either reproducible from the score or it is not.

An illustrative item set (not a real study)

The following five-item set and ten-respondent dataset are constructed for this page to show the mechanics. They are not drawn from a published study, and the coefficients below are computed directly from this constructed dataset, not asserted as general findings.

Items, framed as willingness statements about research data sharing and ordered from least to most demanding:

  1. A — Willing to write a data management plan for the project.
  2. B — Willing to deposit processed data in the home institution’s repository.
  3. C — Willing to deposit processed data in a public, discipline-specific repository.
  4. D — Willing to share data with named colleagues before publication, on request.
  5. E — Willing to make raw, unprocessed data openly available immediately on collection.

Real item ordering is established before data collection, from pretest endorsement rates or theoretical grounds, not derived after the fact from a single sample — the point of scalogram analysis is to check that ordering against real responses, not to invent it from them.

Scalogram analysis: arranging responses to check the pattern

A scalogram is the response matrix itself, arranged with items in difficulty order (easiest first) and respondents sorted by total score (highest first). In a perfectly cumulative scale, that arrangement produces a clean triangular block of 1s in the upper-left and 0s in the lower-right, with no gaps. Ten respondents against the five items above, sorted this way:

Respondent A B C D E Score
R1 1 1 1 1 1 5
R2 1 1 1 1 0 4
R3 1 1 1 0 0 3
R4 1 1 1 0 0 3
R6 1 1 0 1 0 3
R5 1 1 0 0 0 2
R8 1 1 0 0 0 2
R9 0 1 0 0 0 1
R7 1 0 0 0 0 1
R10 0 0 0 0 0 0

Eight of the ten rows fall into the clean triangular pattern their score predicts. Two do not, and both are bolded above: R6 endorses D (“share pre-publication with named colleagues”) without endorsing the easier C (“public repository deposit”), and R9 endorses B without the easier A. Both are real, substantively plausible response patterns — a researcher can reasonably be more willing to share with a named colleague than with a public repository — and that plausibility is exactly why scalogram analysis scores them as errors rather than discarding them as noise: the deviation is data about where the cumulative assumption breaks down, not a data-entry mistake to explain away.

Scoring the scalogram: the coefficient of reproducibility

Reproducibility is scored respondent by respondent. For a respondent with scale score s, the error-free (“ideal”) pattern is 1 on the s easiest items and 0 on the rest; the error count for that respondent is the number of positions where the actual response differs from that ideal pattern (this is the standard scalogram-scoring method set out in McIver & Carmines, Unidimensional Scaling, Sage, 1981). R6 (score 3, pattern 11010 against ideal 11100) has 2 errors: a 0 where the ideal calls for 1 (item C) and a 1 where it calls for 0 (item D). R9 (score 1, pattern 01000 against ideal 10000) also has 2 errors, by the same logic. Every other respondent matches its ideal pattern exactly, contributing 0 errors.

Guttman’s coefficient of reproducibility is:

CR = 1 − (total errors ÷ total entries), where total entries = number of respondents × number of items.

For this dataset: total errors = 2 + 2 = 4, across 10 respondents × 5 items = 50 entries.

CR = 1 − (4 ÷ 50) = 0.92

Guttman’s original rule of thumb treats CR ≥ 0.90 as the threshold for an acceptable unidimensional scale — this dataset clears it, but CR alone is not sufficient to conclude that, for the reason in the next section.

Why CR alone can mislead: minimum marginal reproducibility and the coefficient of scalability

CR has a well-known defect: an item with a very lopsided marginal — almost everyone endorses it, or almost no one does — is reproducible from the score almost by construction, independent of whether the scale is genuinely cumulative. Item E above (willing to release raw data immediately, endorsed by only 1 of 10 respondents) contributes almost no errors no matter what, simply because predicting “everyone says no” is right 90% of the time by chance. A scale padded with several such lopsided items can post a high CR while telling you almost nothing about cumulative structure.

Minimum marginal reproducibility (MMR) is the reproducibility a rater would achieve by guessing each item’s own majority response for every respondent — a floor, not a result. For item A (8 of 10 respondents endorse it), the majority-response guess is right for the same 8; the minority-response guess would only be right for the other 2, so the larger of the two, 8, is what counts. Doing this for all five items in the dataset above:

Item Endorsing / N Majority-category count
A 8/10 8
B 8/10 8
C 4/10 6
D 3/10 7
E 1/10 9

MMR = Σ(majority-category counts) ÷ total entries = 38 ÷ 50 = 0.76

The coefficient of scalability (CS) uses MMR to ask what fraction of the possible improvement over that chance floor the scale actually achieved:

CS = (CR − MMR) ÷ (1 − MMR) = (0.92 − 0.76) ÷ (1 − 0.76) = 0.1600 ÷ 0.2400 = 0.67

The conventional threshold is CS ≥ 0.60. This dataset clears both thresholds — CR = 0.92 and CS = 0.67 — which is the correct joint reading: CR alone would have accepted a scale that padding could inflate; CS confirms the reproducibility is doing real work above the chance floor, not just riding a lopsided item.

Reading item E’s role correctly

Item E’s near-unanimous “no” is not automatically a flaw. It can genuinely be the hardest, most discriminating item on the continuum — that is exactly what a low endorsement rate on the intended top item should look like if the scale is well constructed. The diagnostic move is to check E’s individual contribution to error, not its marginal alone: an item with a lopsided marginal and a disproportionate share of the scale’s errors is a candidate for removal or rewording; a lopsided item contributing no errors is doing its job as the top of the scale. In this dataset E contributes zero errors, so its extreme marginal is not, by itself, a problem.

Building and checking a Guttman scale in practice

  1. Generate a candidate item pool for a construct you have reason to believe is genuinely cumulative — not every attitude or behavior is; forcing a non-cumulative construct into this format just produces a scale that fails its own coefficients.
  2. Order items by expected difficulty or intensity before collecting the scoring sample, from pretest endorsement rates or from the substance of the construct itself.
  3. Collect responses and build the scalogram: sort items by observed marginal (descending endorsement) and respondents by total score, and inspect the resulting matrix visually before computing anything — a scale that is badly non-cumulative is often visible in the arrangement itself.
  4. Compute CR and CS using the ideal-pattern error count shown above. Report both; CR alone is not an adequate quality statement for the reason above.
  5. Diagnose failing items by their individual contribution to total errors, not by marginal alone. An item responsible for a disproportionate share of the errors is the one to revise or drop, then re-score the reduced item set — removing items changes both CR and CS, so this is iterative, not a one-pass check.
  6. Report sample size and item count alongside the coefficients. Both CR and CS are sample statistics computed on a specific N and n, not fixed properties of the item wording; a scale that reproduces well in one sample is not guaranteed to in another.

Where Guttman scaling fits, and where it doesn’t

Guttman scaling is the right tool when the underlying construct plausibly has a true cumulative order — stages of adoption, escalating commitment, graduated exposure or acceptance — and when the research question is specifically whether that order holds empirically. It is the wrong tool for constructs that are multidimensional, or where items are meant to be independent indicators of the same underlying intensity rather than steps on a ladder; that is Likert-scale territory, not Guttman’s. See levels of measurement for where an ordinal Guttman scale sits relative to interval-level instruments, and questionnaire design for the item-writing groundwork — expert review and cognitive interviewing during item development — that a defensible cumulative item set depends on before any scalogram analysis runs.

The technique’s real limitation is the strictness of the deterministic assumption itself: real response data is rarely perfectly cumulative, which is exactly why CR and CS exist as tolerance checks rather than pass/fail tests on a triangle of 1s and 0s. Later probabilistic scaling models relax that strict determinism into a probability of endorsement that increases with respondent standing, at the cost of needing a larger sample and more assumptions to estimate — a tradeoff worth knowing about even on a page focused on the classical, deterministic version of cumulative scaling.

Frequently asked questions

What is Guttman scaling used for?

It tests whether a set of items has a true cumulative structure — whether endorsing a harder item on a construct reliably implies endorsing every easier item — and, if so, reduces the response pattern to a single reproducible scale score. It is used where the construct itself is plausibly a ladder of increasing intensity, difficulty, or commitment, not for generic attitude measurement.

What counts as a good coefficient of reproducibility?

Guttman’s original convention is CR ≥ 0.90. On its own that threshold is necessary but not sufficient — a scale with one or more very lopsided items can clear 0.90 almost by construction, which is why the coefficient of scalability (CS ≥ 0.60, conventionally) is checked alongside it, not instead of it.

How is Guttman scaling different from a Likert scale?

A Likert scale sums independent items that each indicate the same construct at similar intensity; there is no expectation that agreeing with one item implies agreeing with another. A Guttman scale orders items along a single continuum by design, such that a respondent’s total score should, in principle, reproduce exactly which items were endorsed.

What is a scalogram?

The response matrix arranged with items ordered by difficulty and respondents ordered by score, used to visually and statistically check whether responses form the triangular pattern a truly cumulative scale predicts.

What is the coefficient of scalability?

CS = (CR − MMR) ÷ (1 − MMR), where MMR is the reproducibility achievable by simply guessing each item’s own majority response for every respondent. It measures how much of the possible improvement over that chance floor the scale’s actual reproducibility achieves, correcting for the fact that lopsided item marginals inflate CR on their own.

For related survey-instrument material, see the research methods pillar, Likert scale construction and analysis, questionnaire design, levels of measurement, and Cronbach’s alpha for reliability once items are finalized.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.