Written and maintained by CASRAI Editorial Board
Last updated
Test equating is the statistical process that lets a score on one form of a test mean the same thing as a score on a different form of the same test — so a researcher who revises a survey instrument, rotates alternate forms across testing sessions, or runs the same measure across years can report scores on one comparable scale instead of two incompatible ones. This guide covers what actually distinguishes equating from the broader term linking, the three classic equating designs and what makes an anchor test defensible, how equipercentile equating works as a nonlinear alternative to a simple mean/SD adjustment, and how item response theory reframes the same problem as a scale-transformation exercise using characteristic-curve methods like Stocking-Lord and Haebara.
Equating vs. Linking: Not the Same Claim
The two terms get used interchangeably in casual writing, but the distinction is load-bearing for what you can honestly claim about the result. Equating is reserved for forms built to the same content and statistical specifications, intended to measure the same construct at the same difficulty level — alternate forms of a certification exam, or successive waves of the same standardized instrument built from a shared item bank. Because the forms are interchangeable by design, equated scores can be treated as fungible: a 72 on Form A means exactly what a 72 on Form B means, and either score can substitute for the other in any downstream decision.
Linking is the broader term for connecting scores across instruments that are not interchangeable — different tests measuring related but distinct constructs, or the same construct measured at different points on a developmental scale. A linked score relationship is real and useful, but it carries more assumptions and less claim to strict fungibility than an equated one. Vertical linking (sometimes called vertical scaling) connects tests administered to groups of different ability, such as a reading assessment given across several grade levels — the forms are deliberately unequal in difficulty because the underlying trait itself is expected to grow. Horizontal equating, by contrast, connects forms of similar difficulty given to groups of similar ability, such as two annual administrations of the same certification exam.
Reporting an equated score as if the underlying relationship were only linked (or vice versa) overstates or understates what the analysis actually supports — the label belongs in the methods section, not just the appendix.
Three Designs, and Why the Data-Collection Plan Comes First
Equating is a data-collection design decision before it is a statistical one; which method is even available depends on how the forms were administered.
- Single-group design. The same examinees take both forms, usually with order counterbalanced to separate a real form difference from a practice or fatigue effect. Statistically the cleanest design (one group means no need to adjust for ability differences between samples), but rarely feasible outside small-scale research contexts — most operational testing programs cannot ask the same people to sit two full-length exams.
- Equivalent-groups (random-groups) design. Examinees are randomly assigned to take one form or the other. Random assignment makes the two groups equivalent in ability by design, so any difference in the score distributions can be attributed to the forms themselves rather than to who took which one. This is the design large-scale testing programs use when spiraling alternate forms within a single administration.
- Common-item nonequivalent-groups design (the anchor-test / NEAT design). Different, non-equivalent groups take different forms, but every form embeds a common set of anchor items. Because the same items were answered by both groups, the anchor score is used to statistically separate how much of the total-score difference between groups reflects a real ability difference between the groups versus a real difficulty difference between the forms. This is the design used whenever forms are administered on different occasions to different, uncontrolled populations — the most common real-world situation, and the reason anchor-test design gets disproportionate attention in the equating literature.
What Makes an Anchor Test Defensible
An anchor test only does its job if it is a miniature version of the full test it is embedded in, not just any subset of shared items. The anchor should mirror the content coverage and difficulty distribution of the full-length forms, rather than skewing toward easier or harder items or a single content area — a skewed anchor systematically misestimates the group ability difference it exists to correct for. Anchor items are typically embedded unnumbered or unscored within each form so that examinees respond to them under the same conditions as the rest of the test, and their position and surrounding item context should stay consistent across administrations, since item position and context effects (an item performing differently depending on what precedes it) can otherwise masquerade as a form-difficulty difference. Because the anchor is doing real statistical work, not just providing a content sample, its psychometric properties are checked directly — item-level stability across administrations is itself part of the equating quality-control process, not an assumption taken on faith.
Equipercentile Equating
The simplest equating method, linear equating, assumes the two forms’ score distributions differ only in mean and standard deviation and applies a single linear transformation to convert one form’s raw scores to the other’s scale. Equipercentile equating drops that assumption: it defines equivalent scores as whichever raw scores on each form correspond to the same percentile rank in their respective distributions, and connects those points with a curve rather than a straight line. The practical difference shows up whenever the two forms’ distributions differ in shape as well as location — skew, kurtosis, or a ceiling/floor effect on one form but not the other — situations where a linear transformation misestimates the equated score at the tails of the distribution even if it is accurate near the mean. Because raw sample-based percentile curves are jagged from sampling noise, equipercentile equating is typically applied after smoothing the score distributions (either presmoothing the raw frequency distributions or postsmoothing the resulting equating function), trading a small amount of bias for a meaningful reduction in the sampling variability of the equated scores.
IRT-Based Linking: Putting Two Scales on the Same Metric
Item response theory reframes equating as a scale-transformation problem rather than a raw-score-distribution problem. Because IRT item parameters are estimated on an arbitrary scale set by the calibration sample, item and person parameters from two separate calibration runs are not directly comparable until that arbitrary scale difference is removed — which is what IRT linking does.
Two calibration strategies get to that result differently. Concurrent calibration estimates all item parameters from both forms in a single calibration run, with the shared anchor items constrained to have identical parameters across forms — the linking is built directly into the estimation, at the cost of assuming the anchor items truly behave identically in both administrations. Separate calibration estimates each form’s item parameters independently and then finds the linear transformation (a scale factor and an intercept, often labeled A and B) that best aligns the two scales afterward, which is more robust to a shift in an individual anchor item’s behavior but requires an explicit linking step.
The dominant separate-calibration methods are characteristic-curve methods — the Stocking-Lord and Haebara procedures — which choose the transformation constants that minimize the difference between the anchor items’ test or item characteristic curves under the two forms’ parameter estimates, evaluated across the ability range rather than at a single summary statistic. Stocking-Lord works at the test-characteristic-curve level (the anchor items’ curves summed together); Haebara works item by item and sums the squared differences across items. Both are standard, widely implemented alternatives to older mean/mean and mean/sigma linking methods that used only the anchor items’ parameter means, and both require the same anchor-test design discipline described above — an anchor that misrepresents the full test’s content or difficulty distorts the linking transformation regardless of which estimation method is used to compute it.
Where This Matters Beyond Large Testing Programs
Equating and linking are usually associated with large standardized-testing operations, but the same problem shows up any time a research program revises an instrument mid-study or compares scores across a longitudinal series: a patient-reported outcome measure shortened for a follow-up wave, a survey rotated in alternate forms to reduce respondent burden, or a legacy instrument being replaced by a newer version built from the same item bank. Differential item functioning analysis is the companion check worth running alongside any anchor-test design — DIF asks whether an individual anchor item behaves differently across the specific groups being compared, which is exactly the failure mode that silently corrupts an equating or linking result if it goes unchecked. Item banks built under a common IRT scale (the PROMIS system referenced in CASRAI’s item response theory guide is a working example) are equating infrastructure by design: every item in the bank is pre-linked to the same underlying scale, which is what makes computerized adaptive administration and cross-study score comparability possible without a fresh equating study every time a form changes.
Frequently Asked Questions
Is equating the same as norming?
No. Norming converts a raw score to a percentile or standard score relative to a reference population’s score distribution; it describes where a score falls within one distribution. Equating converts a raw score on one form to the corresponding raw (or scale) score on a different form so the two forms can be used interchangeably. A test can be normed without ever being equated to an alternate form, and equating says nothing on its own about where a given score falls in a population.
Can you equate two tests that measure different constructs?
Not defensibly under the label “equating.” If the forms measure meaningfully different constructs or are built to different specifications, the appropriate term is linking, and the resulting score correspondence should be presented with the weaker claim that implies — a linked score relationship is not a license to treat the two scores as interchangeable in a way that specifically requires equivalence, such as a pass/fail cut score set on one form and applied unchanged to the other.
How many anchor items does a common-item design need?
There is no single number that applies across every testing context — it depends on the anchor’s reliability relative to the full test and how large a group-ability difference the design needs to detect. What is consistent across the equating literature is the qualitative requirement: the anchor needs to be long and reliable enough to estimate the group difference precisely, and its content/difficulty mix needs to mirror the full test, not just meet a minimum item count.
Does equating remove the need to check for differential item functioning?
No, and treating it that way is a common error. Equating corrects for an assumed uniform difference between forms or groups; it does not detect or correct for an individual item behaving differently across specific subgroups. A DIF check on the anchor items is a separate, complementary analysis, not a step equating makes redundant.








