A Likert scale is the most widely used format for measuring attitudes, opinions, and self-reported experience in survey research: a respondent is presented with a statement and asked to indicate their level of agreement, frequency, importance, satisfaction, or likelihood along an ordered set of response options. It was introduced by psychologist Rensis Likert in his 1932 paper “A Technique for the Measurement of Attitudes,” and the format has changed remarkably little since.
Two adjacent CASRAI guides cover the surrounding territory: Survey Question Types catalogues the full range of question formats a Likert scale belongs to, and Questionnaire Design covers assembling those formats into a coherent instrument. This guide is the deep dive on the Likert format itself — how to build one correctly, the unipolar/bipolar and point-count decisions that determine whether the data is even interpretable, a working bank of example items, and the long-running debate over how Likert data may legitimately be analysed.
Likert item vs. Likert scale: a distinction worth getting right
Most usage in the wild — including a lot of published research — conflates two different things:
- A Likert item is a single statement paired with an ordered response scale (e.g., “The training materials were easy to follow” — Strongly Disagree to Strongly Agree). Answered alone, it produces one ordinal data point.
- A Likert scale, properly, is the sum or mean of several Likert items that are designed to measure the same underlying construct (e.g., five items that together operationalise “job satisfaction”). The scale score is the composite, not any one item’s response.
This distinction is not pedantry. It is the hinge the entire analysis debate below turns on: treating a single item’s 1–5 response as an interval-level number is a much harder claim to defend than treating the summed score of a validated multi-item scale that way. If your instrument uses only one item per construct, you have Likert items, not a Likert scale, and your analysis options are more limited — see the analysis section below.
Anatomy of a Likert item
A well-formed Likert item has three parts:
- The stem — a single, unambiguous declarative statement or question. It should assert exactly one idea; a stem that bundles two claims (“The onboarding process was fast and well-organised”) is double-barrelled and makes the response uninterpretable, because agreement or disagreement could reflect either half.
- The response options — an ordered set of categories (typically 4 to 7) representing degrees of the underlying dimension.
- Anchor labels — the words attached to each response option.
On anchor labelling specifically: label every point, not just the endpoints. Labelling only “Strongly Disagree” and “Strongly Agree” and leaving the middle points as bare numbers forces respondents to infer what “3” or “4” means relative to the labelled ends, which introduces inconsistency across respondents and undermines the assumption that the intervals are even roughly comparable. Fully labelled scales (e.g., Strongly Disagree / Disagree / Neither Agree nor Disagree / Agree / Strongly Agree) produce more consistent, more interpretable data.
Unipolar vs. bipolar: the decision that determines whether your data means anything
This is a genuinely consequential design choice, not a stylistic one, and it is worth making deliberately for every item.
- Bipolar scales run between two opposite poles through a genuine neutral midpoint: Strongly Disagree ↔ Strongly Agree, or Extremely Dissatisfied ↔ Extremely Satisfied. Bipolar items measure direction and intensity at once — a respondent can be negative, neutral, or positive, and how strongly so.
- Unipolar scales run from an absence of the attribute up to its maximum, with no opposite pole: Not at all → Extremely, or Never → Always. Unipolar items measure intensity or frequency of a single attribute — there is no “opposite” of frequency to anchor a true zero-crossing midpoint against.
The failure mode is mixing the two logics inside one item. Asking “How often do you agree with this statement?” on a disagree–agree scale conflates frequency (unipolar) with attitude direction (bipolar) and produces a response nobody can interpret consistently. As a rule of thumb: use bipolar wording and a labelled midpoint when you genuinely expect respondents to sit on either side of neutral (attitudes, satisfaction, agreement); use unipolar wording when you are measuring the degree of one thing that doesn’t have a natural opposite (frequency, intensity, importance). Unipolar scales often work well with fewer points — five is frequently sufficient — because there is no need to carve out symmetric gradations on both sides of a centre.
How many response points?
4, 5, 7, and 10-point versions are all in active use, and the honest answer is that there is a real trade-off rather than a single correct number:
- More points (5–7) generally give respondents finer discrimination and can modestly improve reliability and the ability to detect differences between groups, but the gains flatten out past roughly seven points — beyond that, respondents typically cannot reliably distinguish adjacent categories, and added precision is illusory.
- Fewer points (4 or an even-numbered scale generally) are faster to complete and reduce respondent fatigue on long instruments, but coarser categories lose information and can compress genuine variation.
- 10-point scales are common in satisfaction/NPS-style contexts by convention, but they exceed the range most respondents can meaningfully differentiate on a bipolar attitude construct, and their endpoints are prone to skew.
There is no consensus number, and CASRAI is not going to pretend otherwise — choose based on how finely the construct can genuinely be discriminated, how long the instrument already is, and consistency with any existing validated scale you are adapting rather than building from scratch.
The midpoint decision
Whether to include a labelled neutral/midpoint category is its own trade-off:
- Include a midpoint and you give respondents who are genuinely neutral, ambivalent, or lack an opinion a legitimate place to sit — but some respondents who do lean mildly one way will default to the midpoint anyway to avoid effort or exposure (a form of satisficing), which biases attitude estimates toward the centre.
- Omit the midpoint (a “forced-choice” even-point scale) and you eliminate that neutral-dumping, but you force genuinely undecided or neutral respondents into a directional answer that misrepresents their actual position, inflating apparent polarisation.
Neither choice is free of distortion. The decision should be driven by whether “no opinion” or “genuinely neutral” is a real, expected position for your population on this construct (include a midpoint) or whether you specifically need to prevent non-response hiding inside a comfortable middle category (omit it).
Worked example items
The items below are illustrative examples written for this guide, not drawn from any specific published instrument — adapt wording and anchors to your own construct, and always pilot-test before fielding.
Agreement (bipolar, 5-point)
Stem: “The onboarding documentation was clear and easy to follow.”
Anchors: Strongly Disagree / Disagree / Neither Agree nor Disagree / Agree / Strongly Agree
Frequency (unipolar, 5-point)
Stem: “In the past month, how often did you consult the shared protocol repository before starting a new procedure?”
Anchors: Never / Rarely / Sometimes / Often / Always
Importance (unipolar, 5-point)
Stem: “How important is real-time inventory visibility to your day-to-day lab work?”
Anchors: Not at all Important / Slightly Important / Moderately Important / Very Important / Extremely Important
Satisfaction (bipolar, 7-point)
Stem: “Overall, how satisfied are you with the turnaround time on equipment repair requests?”
Anchors: Extremely Dissatisfied / Dissatisfied / Somewhat Dissatisfied / Neither Satisfied nor Dissatisfied / Somewhat Satisfied / Satisfied / Extremely Satisfied
Likelihood (unipolar, 5-point)
Stem: “How likely are you to recommend this training module to a colleague?”
Anchors: Not at all Likely / Slightly Likely / Moderately Likely / Very Likely / Extremely Likely
Mismatched anchors — what to avoid
Stem: “How often do you agree that the lab’s safety briefings are useful?”
This conflates a frequency question (“how often”) with an agreement judgment, and no single anchor set can honestly label the response options. Rewrite as either a pure frequency item (“How often do you attend the lab’s safety briefings?”) or a pure agreement item (“The lab’s safety briefings are useful.”).
The analysis controversy: can you take a mean of Likert data?
This is the most consequential methodological question a Likert scale raises, and it is a genuine, long-running debate in the measurement literature rather than a settled fact either direction on the internet tends to present it as.
The core issue is level of measurement. A Likert item is ordinal: response categories have a defined order, but the psychological distance between “Agree” and “Strongly Agree” is not established to equal the distance between “Neither Agree nor Disagree” and “Agree.” Arithmetic operations like the mean assume equal intervals between values — an assumption ordinal data does not guarantee.
The strict position
Applied strictly, single Likert items should be summarised with median and mode and reported as frequency distributions, and compared using non-parametric tests appropriate to ordinal data: the Mann-Whitney U test for two independent groups, Kruskal-Wallis for three or more, and chi-square tests of association for categorical comparisons.
The applied convention
In practice, a large share of published survey research treats summed or averaged multi-item Likert scales (not single items) as approximately interval and analyses them with means, t-tests, and ANOVA. This is a materially more defensible move than doing the same to a single item: summing several items measuring one construct tends to produce a composite score with a wider range and a more continuous, often approximately normal distribution, which weakens the practical impact of any one item’s interval-assumption violation. See ANOVA and t-test for what those parametric options assume and require.
Where the actual consensus sits
Stated plainly: treating the mean of a single Likert item as a precise, interval-scaled quantity is hard to justify and should be done cautiously, with the ordinal caveat stated explicitly. Treating the mean of a validated, multi-item Likert scale as approximately interval is broadly accepted practice in applied social-science and health-services research, provided the scale’s reliability has actually been established (see below) — not assumed. This is exactly why the item/scale distinction at the top of this guide matters in practice, not just in terminology. The same ordinal-coding tension shows up in qualitative work too; see Qualitative Data for the parallel argument against treating coded categorical responses as more precise than they are.
Reliability and validity
Reliability — whether the instrument produces consistent results — is assessed for a multi-item Likert scale using Cronbach’s alpha, which measures internal consistency: the degree to which items intended to measure the same construct actually correlate with one another. See Cronbach’s Alpha for interpretation thresholds and when to use omega instead. Alpha requires multiple items measuring the same construct to compute at all — it is meaningless, and cannot be calculated, for a single Likert item.
Validity — whether the scale actually measures the construct it claims to — is a separate question from reliability and is not established by a high alpha alone. See Reliability in Research for how reliability and validity relate and differ.
Response biases specific to rating scales
- Acquiescence bias — a tendency to agree with statements regardless of content. Mitigation: include reverse-worded items so agreement isn’t always the “same-direction” answer — though reverse-coding introduces its own problem, since poorly worded reversed items (double negatives, awkward phrasing) confuse respondents and can reduce rather than improve data quality. Use reverse items sparingly and pilot-test them.
- Central tendency bias — a tendency to avoid extreme response categories and cluster around the midpoint, compressing genuine variation. More common with unfamiliar or emotionally loaded topics.
- Extreme responding — the opposite tendency, disproportionately selecting endpoint categories; varies by population and can distort cross-group or cross-cultural comparisons if not accounted for.
- Social desirability bias — responding in the way the respondent believes is expected or favourable, rather than their actual position, especially on sensitive topics. Anonymity and self-administration reduce but do not eliminate it.
- Straightlining / satisficing — selecting the same response option down an entire grid of items with minimal engagement, especially on long batteries. Mitigation: vary item direction, keep grids short, and screen for identical-response patterns during data cleaning.
Visualising Likert data
A single mean value discards most of the information in a Likert response distribution and can actively mislead when responses are bimodal (clustered at both ends with few in the middle) — a mean near the midpoint would suggest consensus around “neutral” when no respondent actually chose it. Prefer:
- Stacked bar charts, showing the full frequency distribution across response categories for one or more items.
- Diverging (bipolar) stacked bar charts, which anchor the neutral category at a shared centre line and extend disagreement to the left and agreement to the right — the standard, and generally clearest, way to visualise a bank of bipolar Likert items side by side.
Common errors checklist
- Double-barrelled stems that bundle two claims into one item.
- Anchors that don’t logically match the stem (e.g., frequency wording on an agreement scale).
- Leaving midpoints and interior points unlabelled while only labelling the endpoints.
- Reporting the mean of a single Likert item without acknowledging its ordinal nature.
- Running parametric tests (t-test, ANOVA, Pearson correlation) on Likert data without stating and defending the interval assumption.
- Treating a 5-point (or any) Likert item as though it had ratio properties — i.e., assuming a “4” reflects twice the attitude of a “2.” Likert data, even treated as approximately interval, is never ratio-scaled: there is no true, meaningful zero.
- Calling a single item a “scale” and reporting Cronbach’s alpha for it, which is not computable for one item.
Frequently asked questions
Is a Likert scale ordinal or interval data?
Strictly, it is ordinal: categories are ordered but the distance between adjacent categories is not established to be equal. Many researchers treat the summed score of a validated multi-item scale as approximately interval for analysis purposes; this is a more defensible move for a multi-item composite than for a single item.
Can you calculate a mean for Likert data?
For a single item, doing so is contested and should be reported alongside median/mode with the ordinal caveat stated. For a summed or averaged multi-item scale with established reliability, reporting means is widely accepted practice.
What’s the difference between a Likert item and a Likert scale?
A Likert item is one statement with an ordered response set. A Likert scale is the composite (sum or mean) of several items measuring the same underlying construct. The terms are frequently used interchangeably in casual usage, but the distinction matters for what analysis is defensible.
Should a Likert scale have an even or odd number of points?
Odd-numbered scales include a labelled midpoint/neutral option; even-numbered “forced-choice” scales omit it. Each has a documented distortion risk — midpoint-dumping by mildly-opinionated respondents versus forcing genuinely neutral respondents into a direction — and the right choice depends on whether “no opinion” is a real, expected position for your construct and population.
How many points should a Likert scale have?
5 and 7 are the most common choices in practice. Reliability gains from adding points taper off past roughly seven, since respondents generally cannot reliably discriminate finer categories beyond that.
What is a unipolar Likert scale?
A scale running from the absence of an attribute to its maximum (e.g., Not at all → Extremely), used for intensity or frequency constructs that don’t have a natural opposite pole. Contrast with bipolar scales, which run between two opposite poles through a genuine neutral midpoint.







