Skip to main content
v2026.11,610 entries · CC-BY 4.0

Semantic Differential Scale: Construction, Scoring & When to Use It

Constructing valid bipolar adjective pairs, Osgood’s evaluation/potency/activity structure, scoring conventions, and when a semantic differential beats a Likert item.

Ask about Semantic Differential Scale: Construction, Scoring & When to Use It

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

A semantic differential scale asks a respondent to rate a concept — a brand, an institution, a policy, a piece of technology, even an abstract idea — on a series of continuous scales anchored at each end by a pair of polar-opposite adjectives (for example, Weak — Strong or Passive — Active), rather than by an agreement statement. The respondent marks a point between the two poles that best reflects where the concept sits on that dimension. It was developed by psychologist Charles E. Osgood and colleagues and published in their 1957 book The Measurement of Meaning (Osgood, Suci, and Tannenbaum), as a way to measure the connotative meaning a concept carries for a respondent — not what it denotes literally, but the affective, evaluative charge attached to it.

This guide covers how to build bipolar adjective pairs that actually measure a single dimension (the most common way these scales go wrong), the scoring convention Osgood used and why it matters for analysis, and the specific measurement goals — attitude, image, and connotative meaning — where a semantic differential outperforms a Likert scale. For the broader survey-instrument context, see CASRAI’s guides on survey question types and questionnaire design.

Osgood’s three dimensions: evaluation, potency, activity

Osgood did not set out to build a single-purpose attitude scale. He was trying to answer a broader question in psycholinguistics: does the emotional or connotative meaning people attach to concepts have a consistent underlying structure, regardless of the concept being rated? To test this, he collected ratings of many different concepts across dozens of bipolar adjective pairs, then ran exploratory factor analysis on the results.

Three factors consistently accounted for most of the variance across concepts and cultures:

  • Evaluation — the good/bad dimension: good–bad, pleasant–unpleasant, kind–cruel. This factor is almost always the strongest and is the one most survey applications actually care about — it is the closest analogue to a general attitude score.
  • Potency — the strong/weak dimension: strong–weak, hard–soft, heavy–light. Captures perceived power or magnitude, independent of whether that power is viewed favorably.
  • Activity — the active/passive dimension: active–passive, fast–slow, excitable–calm. Captures perceived energy or dynamism.

The practical implication for building a scale today: if your goal is a general attitude or connotation measure, evaluation-dimension pairs are almost always where your item pool should concentrate. Potency and activity pairs are worth including when the concept genuinely varies on those dimensions for your respondents (brand-personality and organizational-image research uses all three routinely), but padding an evaluation-focused instrument with off-dimension pairs just to hit an item count adds noise, not signal.

Constructing valid bipolar adjective pairs

The single most common error in building a semantic differential is treating “two adjectives that feel related” as interchangeable with “a genuine bipolar pair.” They are not the same thing, and the difference determines whether the resulting scale measures one dimension or silently blends several.

The non-antonym-pair error

A valid pair must be true polar opposites on one specific continuum — not merely two words that could both plausibly describe the concept, and not a synonym dressed up as a contrast. Two failure patterns show up repeatedly in poorly built instruments:

  • Cross-dimension pairing. Honest — Aggressive looks like a contrast because both words carry connotation, but “honest” sits on the evaluation dimension and “aggressive” sits closer to potency/activity. A respondent who thinks the concept is honest but also aggressive has no coherent way to place a single mark between poles that aren’t actually opposite ends of the same idea — the item collapses two constructs into one scale and the resulting score is uninterpretable.
  • Near-synonym pairing. Reliable — Untrustworthy reads as a contrast but the poles are not symmetric opposites of the same word sense — “untrustworthy” implies an active judgment (“has been shown unreliable”) that “reliable” doesn’t symmetrically negate. The fix is a true antonym pair drawn from the same word family: reliable–unreliable, or, if the connotative charge of “untrustworthy” is specifically what you want to capture, trustworthy–untrustworthy as its own pair, not mixed with “reliable.”

Two checks catch most instances of this before piloting:

  1. The antonym-dictionary check. Every candidate pair should be each other’s listed antonym in a standard thesaurus, not merely a pair the item writer intuited as opposite. Osgood’s own item pool was built this way — adjective pairs were selected and refined using antonym-frequency data, not free association.
  2. The relevance check. A pair can be a perfectly valid antonym pair in general English and still be a bad fit for a specific concept. Hot–cold is a clean antonym pair, but asking respondents to rate a research questionnaire‘s hot–cold makes no sense as a concept-rating task; the same pair is exactly right for rating a cup of coffee. Pilot each candidate pair against the actual concept being rated (not a generic list reused from an unrelated study) and drop any pair respondents report as “doesn’t apply” or leave blank at a noticeably higher rate than the rest of the set — that is a direct behavioral signal the pair doesn’t fit the concept, not just a theoretical concern.

Balance the direction of the “positive” pole

A further construction detail that is easy to skip and cheap to fix: don’t put every socially-desirable pole on the same side (e.g., always the right-hand end). Respondents develop a position habit — marking the same side repeatedly without reading each pair — when polarity is constant across the whole instrument. Randomize which pole (positive or negative) appears on the left versus the right across items, then reverse-code the flipped items before scoring. This is the semantic-differential equivalent of mixing reverse-worded items into a Likert battery, and it exists for the same reason: to catch straight-lining and reduce acquiescence-style response sets.

Scoring conventions

Osgood’s original format used a 7-point scale, most commonly scored -3 to +3 with 0 as the neutral midpoint, though a 1-to-7 equivalent (with 4 as the midpoint) is functionally identical and more common in modern survey software, since some platforms handle a signed scale awkwardly. Either convention is acceptable as long as it’s applied consistently and documented in the methods section — what matters for downstream analysis is the number of scale points and which end is which, not whether the reported numbers happen to include a minus sign.

Two scoring decisions to make explicit before fielding, not after data collection:

  • Item-level vs. composite score. A single pair (e.g., just good–bad) produces one number per respondent per concept — useful for a quick directional read, but thin as a measurement of a broader construct. Where the goal is a construct like “attitude toward X” rather than a single adjective judgment, average or sum several evaluation-dimension items into a composite, the same logic that turns individual Likert items into a Likert scale — a composite pools out item-specific noise and gives a more stable, more reliable score than any one pair alone.
  • Reverse-coding before aggregating. Any item where the positive pole was randomized to the low-numbered end must be reverse-coded (subtract the raw score from the scale’s maximum plus one, or flip sign on a -3 to +3 scale) before it’s combined with items scored in the standard direction. Aggregating raw, unreversed scores silently cancels out real variance and is one of the more common analysis errors in published work that uses this format.

Analysis: profile view and composite/factor scores

Semantic differential data supports two complementary analyses that a single Likert item doesn’t naturally offer:

  • Semantic profile. Plotting the mean score for every adjective pair, in order, produces a visual “profile” of how a concept is perceived across the full item set — a signature shape that can be compared concept-to-concept (how does Brand A’s profile compare to Brand B’s?) or group-to-group (do novice and expert respondents produce different profiles for the same concept?). This visual, multi-dimensional comparison is the semantic differential’s most distinctive analytic use, and it has no direct Likert equivalent, since a Likert battery is normally collapsed to a single composite score rather than read as a shape.
  • Dimension scores. Where the item pool spans evaluation, potency, and activity pairs, factor analysis (or, for a pre-specified structure, confirmatory factor analysis) confirms whether the pairs actually load on the intended three dimensions for your sample and concept, then produces a separate composite score per dimension rather than one flattened score. Treat this as a construct-validity check, not an optional step — an item pool that doesn’t factor as expected is telling you some pairs are measuring something other than what you intended, which is exactly the non-antonym-pair and cross-dimension problems described above resurfacing statistically.

As with Likert data, whether a semantic differential’s ordinal response points can be treated as interval-level for parametric statistics (means, t-tests, ANOVA) is a live methodological debate rather than a settled rule; the semantic differential’s case for interval treatment is generally considered somewhat stronger than a single Likert item’s, because the equal-appearing-intervals assumption between two clearly labeled poles is easier to defend than between category labels like “agree” and “strongly agree” — but it is still an assumption to state and defend in a methods section, not treat as automatic.

Semantic differential vs. Likert: when to prefer which

Both formats measure attitudes, and both are legitimate, well-established choices — the decision is about what, specifically, is being measured, not which instrument is generically “better.”

Use a semantic differential when Use a Likert scale when
The goal is the overall connotation or image a concept holds — how a brand, institution, product, or idea “feels” across several qualities at once. The goal is agreement/disagreement with a specific, explicit statement or claim.
A multi-dimensional profile comparing several concepts (or the same concept across groups/time) is genuinely useful to the research question. A single composite construct score is the actual deliverable — no profile comparison is needed.
Respondents can meaningfully react to single adjectives without needing a full sentence of context (works best for concrete, familiar concepts: brands, products, well-known institutions). The construct needs precise verbal framing that a bare adjective pair can’t carry (e.g., a specific behavioral claim, a policy statement, a satisfaction judgment tied to a defined experience).
Cross-cultural or cross-language comparability of connotative meaning is a study goal — this is literally what Osgood built the method to test. The instrument needs to slot into an existing validated Likert battery for comparability with prior research using that instrument.

In practice: a satisfaction survey asking “The support team resolved my issue quickly” (agree–disagree) is a Likert use case — it’s testing agreement with a specific claim. Asking the same respondent to rate “this company” on Unhelpful–Helpful, Slow–Fast, Impersonal–Personal is a semantic differential use case — it’s building a connotative profile of the company itself, not testing agreement with any one statement. Many strong instruments use both: a semantic differential battery for overall image/attitude, paired with Likert items for specific, actionable claims a team can act on directly.

Worked example: rating a research data repository

An example item set for a semantic differential measuring researcher attitudes toward a specific data repository, mixing dimensions deliberately (evaluation-heavy, one potency and one activity pair included):

Adjective pair Dimension
Untrustworthy — Trustworthy Evaluation
Confusing — Clear Evaluation
Outdated — Modern Evaluation
Weak — Robust Potency
Slow — Fast Activity

Each pair is presented on a 7-point scale (1–7, or -3 to +3), with the positive pole randomized left/right per item and reverse-coded during scoring. The three evaluation items are averaged into an overall attitude composite; potency and activity are reported separately since only one item represents each and a single item cannot be meaningfully averaged into a multi-item composite.

Frequently asked questions

How many adjective pairs does a semantic differential need?

There’s no fixed minimum, but a single pair only ever produces a single-item measurement with all the reliability limitations that implies — the same limitation a single Likert item has. For a composite attitude score, plan on at least three to five evaluation-dimension pairs measuring the same underlying construct, more if the concept is complex enough to warrant separate potency and activity composites as well.

Can a semantic differential use more or fewer than 7 points?

Yes — 5-point and 9-point variants both appear in published research, and the choice follows the same general trade-off as Likert response-point counts: fewer points are faster to complete and easier for respondents to use consistently; more points allow finer discrimination but raise the risk of respondents not reliably distinguishing adjacent points. Seven is Osgood’s original convention and remains the most common default absent a specific reason to deviate.

Do the adjective pairs have to be single words?

Osgood’s original pairs were single adjectives, which is what keeps the format fast to complete and easy to scan visually as a profile. Short phrase pairs are sometimes used in applied market-research instruments (e.g., Easy to use — Difficult to use), but every phrase pair should still pass the same antonym and relevance checks described above — a phrase pair is not exempt from needing to be a genuine bipolar opposite.

Is a semantic differential score interval or ordinal data?

Treat it as ordinal by default and justify any interval-level treatment explicitly, the same standard applied to Likert data. The argument for interval treatment is somewhat stronger here because the two clearly-labeled poles anchor the scale more concretely than category labels do, but it remains an assumption to state, not a default to assume silently.

For the surrounding instrument-design decisions — how many response points, unipolar vs. bipolar framing, and building a full multi-item scale — see CASRAI’s guides on the Likert scale and questionnaire design, and the broader research methods hub.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.