Skip to main content
v2026.11,610 entries · CC-BY 4.0

Item Response Theory for Research Scales

What item response theory is, how the item-characteristic curve and its difficulty and discrimination parameters work, what the 1PL, 2PL, and 3PL models add, and why IRT (unlike classical test theory) enables adaptive testing.

Ask about Item Response Theory for Research Scales

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

Item response theory (IRT) models the probability that a person responds to a scale item in a particular way as a function of that person’s underlying trait level and the item’s own statistical properties — not, as classical scoring does, by simply summing responses and treating every item as interchangeable. For researchers building or selecting a scale — a patient-reported outcome measure, a personality inventory, an attitude survey — IRT is the framework behind item banks, computerized adaptive testing, and cross-sample score comparability. This guide covers the item-characteristic-curve logic IRT is built on, what the difficulty and discrimination parameters actually mean, how the 1PL, 2PL, and 3PL models differ, why IRT (unlike classical test theory) enables adaptive testing, and what a researcher building a scale needs to know before choosing one framework over the other.

From Classical Test Theory to Item Response Theory

Classical test theory (CTT) models an observed score as a true score plus random error, and it scores a scale by summing (or averaging) item responses — every item counts equally toward the total, regardless of how well that specific item distinguishes high-trait from low-trait respondents. CTT’s item statistics — the proportion endorsing an item, the item-total correlation — describe how an item behaved in one specific sample. Change the sample’s trait distribution and those statistics change with it, even though nothing about the item itself changed.

IRT starts from a different unit of analysis: the single item response, modeled as a probability that depends on the respondent’s position on a latent trait (denoted θ, theta) and on parameters that describe the item itself. Because the model estimates person trait level and item parameters on the same scale, and because (under reasonable model fit) an item’s parameters hold constant across different samples drawn from the same population, IRT item statistics travel across studies and administrations in a way CTT’s sample-bound statistics do not. That single property — sample-independent item calibration — is what makes item banking, test equating across different forms, and adaptive administration possible in the first place; none of them work if an item’s difficulty is really just an artifact of who happened to take the test.

The Item Characteristic Curve: What It Actually Models

The core object in IRT is the item characteristic curve (ICC), sometimes called the item response function. For a single item, it plots the probability of a given response (for a binary item: endorsing it, or answering it correctly) against θ on the horizontal axis. The result is an S-shaped (logistic) curve: at very low trait levels the probability of endorsement sits near zero, at very high trait levels it approaches one, and in between it rises through an inflection point. Every respondent has some position on that same theta axis, and the ICC tells you, for any position, how likely that person is to endorse the item.

The general form of the model (the two-parameter logistic, explained fully below) is:

P(θ) = 1 / (1 + e−a(θ−b))

Two things about that curve carry all the information IRT extracts from an item: where it crosses the midpoint of the probability range, and how steeply it rises when it does. Those two features are exactly what the difficulty and discrimination parameters describe.

Difficulty and Discrimination: What the Parameters Actually Mean

Despite the name (inherited from ability testing), difficulty and discrimination apply just as directly to an attitude, personality, or health-outcome item as to an exam question — “difficulty” just means how much of the underlying trait it takes before a respondent is likely to endorse the item.

  • Difficulty (b) — the location of the item on the theta scale: the trait level at which a respondent has a 50% chance of endorsing the item (ignoring guessing). A low-b item is easy to endorse — most respondents, even those low on the trait, say yes to it. A high-b item requires a high trait level before endorsement becomes likely. On a depression scale, for example, “I sometimes feel tired” would sit at low difficulty (most respondents endorse it regardless of depression severity) while “I have thoughts that life isn’t worth living” would sit at high difficulty (only respondents with substantial depressive symptoms endorse it).
  • Discrimination (a) — the steepness of the curve at its midpoint, which determines how sharply the item distinguishes respondents just below its difficulty level from respondents just above it. A high-discrimination item’s probability of endorsement changes quickly over a narrow range of theta — it’s diagnostic right around its difficulty level, but not very informative far from it. A low-discrimination item’s curve is shallow: endorsement probability creeps up gradually across a wide range of theta, so the item doesn’t sharply separate anyone. An item with near-zero or negative discrimination is one CTT would flag through a weak or reversed item-total correlation — IRT recovers the same warning as a parameter on the curve itself.

A well-designed scale spreads its items’ difficulty parameters across the range of theta the instrument is meant to measure, with reasonably high discrimination throughout, so that no matter where a respondent sits on the trait, some items on the scale are actually informative about that respondent.

1PL, 2PL, and 3PL: What Each Model Adds

“IRT” is a family of models, not one equation, and the three most common variants differ in exactly which item parameters they let vary:

  • 1PL (the Rasch model) — developed by the Danish mathematician Georg Rasch and published in 1960, the 1PL model estimates only difficulty; discrimination is fixed at the same value for every item (equivalently, treated as not varying), and there is no guessing parameter. This is a meaningfully stricter model than it sounds: because every item is assumed equally discriminating, an item that doesn’t fit that assumption in the real data is flagged as a misfitting item rather than accommodated by letting its discrimination float. Rasch measurement is used deliberately for that property — strict, testable model fit — particularly in fields (education measurement, some patient-reported outcome work) where the Rasch tradition is well established as its own methodology, distinct from broader IRT.
  • 2PL — extends the model to let discrimination vary by item, alongside difficulty. This logistic formulation is generally credited to Allan Birnbaum’s chapters in Lord and Novick’s 1968 Statistical Theories of Mental Test Scores, and it’s the most commonly used IRT model for attitude, personality, and other Likert-type research scales, since it lets items differ in how informative they are without adding a guessing parameter that rarely applies to non-cognitive items.
  • 3PL — adds a pseudo-guessing parameter (c), the lower asymptote of the curve: the probability that a respondent very low on the trait still endorses (or gets right) the item by chance. This matters for multiple-choice cognitive/ability items, where a low-ability test-taker can still guess correctly, but it has little natural interpretation for most attitude or symptom-report scales — there’s no equivalent of “guessing” whether you agree with a statement about your own mood. 3PL is standard in large-scale ability testing (see the “Adaptive Testing” section below) and rare in research-scale development outside that context.

For most research-scale work — a new patient-reported outcome measure, an attitude or belief inventory — the real choice is between 1PL/Rasch (favored for its strict, diagnostic fit-testing and smaller sample requirements) and 2PL (favored when items are expected to genuinely differ in how discriminating they are, which is the norm for most researcher-written item pools).

Why IRT Enables Adaptive Testing

A fixed-form test or survey gives every respondent the same items, whether or not those items are informative for that respondent’s trait level — a high-difficulty item wastes a low-trait respondent’s time telling you almost nothing you didn’t already know, and vice versa. Because each item’s discrimination and difficulty are estimated on the same theta scale, IRT lets you calculate an item information function for every item: how much statistical information that item provides at each point on theta. An item is maximally informative near its own difficulty level and contributes very little far from it.

Computerized adaptive testing (CAT) uses that property directly. Instead of administering a fixed item set, a CAT algorithm: (1) starts with a provisional estimate of the respondent’s theta, often from a moderate-difficulty starting item; (2) selects the next item from the bank as the one that provides maximum information at the current theta estimate; (3) updates the theta estimate after each response; and (4) repeats, item by item, until a stopping rule is met — typically a target standard error of measurement, or a maximum item count. Because every item administered is chosen to be maximally informative for that specific respondent, CAT reaches a given measurement precision in substantially fewer items than a fixed-length form covering the same trait range would need. This is the direct payoff of sample-independent item calibration: none of it works unless an item’s difficulty and discrimination parameters mean the same thing regardless of who has taken the test so far, which is exactly the property CTT’s sample-bound statistics don’t have.

This isn’t a theoretical example. Large-scale testing programs (the GRE’s adaptive sections, for one) and, closer to most research-scale work, the NIH’s PROMIS item banks for patient-reported outcomes are built on exactly this logic — large, IRT-calibrated item pools that can be administered adaptively (a short CAT form) or as fixed short-forms drawn from the same calibrated bank, with scores that stay comparable across administration modes because the underlying item parameters are the same either way.

IRT vs. Classical Test Theory: When Each Is the Right Tool

Dimension Classical Test Theory Item Response Theory
Unit of analysis The total (summed) score The individual item response
Item statistics Sample-dependent (proportion endorsing, item-total correlation) Approximately sample-independent, given model fit
Enables item banking / equating across forms Not directly — requires separate equating methods Yes, by design
Enables adaptive testing No Yes
Typical sample size Smaller samples workable Larger — commonly cited guidance runs from roughly 100–200 for Rasch/1PL up toward several hundred for 2PL, and more again for 3PL, though the precise minimum depends on item count and model complexity
Common software Any general stats package (SPSS, R base, Excel) Specialized: R’s mirt and ltm packages, eRm for Rasch models, or standalone programs like Winsteps
Best fit for Smaller studies, simpler scoring needs, quick reliability checks Item banks, adaptive administration, cross-study or cross-form score comparability

Neither framework strictly supersedes the other. A well-conducted CTT analysis (internal consistency, item-total correlations, exploratory factor structure) is still the right starting point for most new scales, especially with modest sample sizes, and CASRAI’s guide to reliability in research measurement covers that ground in depth. IRT becomes the better tool once the scale needs to travel — across samples, across shortened or adapted forms, or into an adaptive administration — because that’s specifically the problem CTT’s sample-bound item statistics can’t solve.

Practical Considerations for Researchers Building a Scale

  • Start with construct and content work before modeling. IRT tells you how items behave statistically; it says nothing about whether the item pool actually covers the intended construct. Establish content validity and draft a broad initial item pool before any IRT calibration.
  • Budget a larger sample than a CTT-only analysis would need. The comparison table above gives the general pattern — smaller for 1PL/Rasch, larger for 2PL, larger still for 3PL — and the exact number to plan for depends on the specific software’s estimation method and the number of items being calibrated simultaneously; consult a psychometrician or the documentation for your chosen software before finalizing a target sample size.
  • Check model fit before trusting the parameters. Both item-level fit statistics (e.g., infit/outfit for Rasch models) and overall model comparisons (1PL vs. 2PL vs. 3PL, via likelihood-ratio tests or information criteria) matter — a parameter estimated from a poorly fitting model isn’t more trustworthy just because it came from a more sophisticated framework than CTT.
  • Match the model to the item type and use case. Rasch/1PL for programs that need strict, diagnostic item fit; 2PL for most Likert-type research scales; 3PL essentially only for multiple-choice ability/achievement testing where guessing is a real phenomenon.
  • Don’t reach for IRT if you don’t need what it provides. If the scale will be used once, in one sample, without adaptive administration or cross-form equating, a well-executed CTT analysis is simpler, needs a smaller sample, and answers the questions a single study actually has.

Frequently Asked Questions

Is IRT always better than classical test theory?

No. IRT solves specific problems — sample-independent item calibration, item banking, adaptive testing, cross-form score comparability — that a single, one-off study often doesn’t have. For a modest-sample study using a scale once, a properly conducted CTT analysis (reliability, item-total correlations, factor structure) is frequently the more practical and adequately powered choice.

Can I use IRT with a small sample?

The Rasch (1PL) model is the most sample-efficient of the common IRT models and is sometimes used with samples in the low hundreds, but 2PL and especially 3PL models estimate more parameters per item and generally need correspondingly larger samples to produce stable estimates. There’s no single universal minimum — it depends on item count, model, and estimation method — so this is worth confirming against your specific software’s guidance or a psychometric consultant rather than a single rule of thumb.

Do I need IRT to compute a reliability estimate?

No. Standard reliability statistics like Cronbach’s alpha come from classical test theory and don’t require an IRT model. IRT has its own, related reliability concept — the test information function, which shows how much measurement precision the instrument provides across the theta range — but it’s a different (and more granular) computation, not a prerequisite for ordinary reliability reporting.

What’s the difference between IRT and factor analysis?

They’re closely related mathematically — a 2PL IRT model for binary items is equivalent to a particular item-factor-analysis model — but they’re typically used for different purposes. Factor analysis is usually used to establish or confirm a scale’s underlying structure (how many dimensions, which items load on which); IRT is usually applied after that structure is settled, to calibrate items on a single dimension for scoring, banking, or adaptive use.

Related CASRAI Resources

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.