Skip to main content
v2026.11,772 entries · CC-BY 4.0

Q Methodology: Q-Sorts, By-Person Factor Analysis, and Reading the Factor Arrays

Q methodology transposes the factor-analytic matrix so each person’s whole Q-sort becomes a variable. That inversion sets the loading threshold from the number of statements, not participants — and drives the Q-set, P-set, rotation and factor-retention decisions.

Ask CASRAI · included with Regulatory Radar

Ask about Q Methodology: Q-Sorts, By-Person Factor Analysis, and Reading the Factor Arrays

Ask CASRAI answers research-administration questions about this guide and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Written and maintained by CASRAI Editorial Board

Last updated

The single thing that decides whether a Q study is done correctly is which way round the data matrix sits. In ordinary (R) factor analysis, the items are the variables and the people are the cases. Q methodology transposes that: each completed Q-sort — one person’s whole rank-ordering — becomes a variable, and the statements become the cases. Everything downstream follows from that inversion, and most of the errors that get published follow from forgetting it.

The most consequential consequence is the one people miss. Because statements are the cases, the standard error used to judge factor loadings is computed from the number of statements, not the number of participants. Brown’s primer gives it as SE = 1/√N, where N is the number of statements, and works it for a 20-statement Q sample: 1/√20 = 1/4.47 = 0.22. Recruiting more participants does not shrink that standard error. If you want more precision on which sorts define a factor, you need a longer Q-set — adding people does not help.

What Q methodology actually answers

Q establishes which coherent viewpoints exist on a contested topic and what each one holds together. It does not establish how common those viewpoints are. Brown is explicit about the boundary: after locating three distinct perspectives in his worked example he notes that “we do not know the proportions of factor Ia, Ib, or II types which exist in the general population (a matter of nose-counting best left to surveys).”

That is the correct division of labour. Q tells you the shape of the discourse; a probability survey tells you its distribution. A Q study that reports percentages of the population holding each viewpoint has overstepped, and a reviewer should catch it. Conversely, a Likert survey that averages across respondents will happily produce a mean nobody actually holds — which is exactly the failure Q was designed around, because each sort is ipsative: statements are ranked against each other within a person, not scored independently.

The five decisions that determine the study

1. Sampling the concourse

The concourse is the full universe of things that can be sensibly said about the topic — Brown takes the term from the Latin concursus, “a running together.” It is not restricted to text: published Q studies have used newspaper clippings, cartoons, audio recordings and country-music selections as sortable items.

In the 289 healthcare Q studies surveyed by Churruca and colleagues, the concourse was most often built from academic literature (n = 186, 64.4%), interviews (n = 138, 47.8%), expert input (n = 59, 20.4%), focus groups (n = 45, 15.6%) and grey literature (n = 44, 15.2%). Thirty-three studies (11.4%) reused a Q-set from a previous study. Twelve studies (4.2%) did not report how the concourse was assembled at all — which makes the resulting factors uninterpretable, because a factor is only ever a structure within the statements you chose to include.

2. Reducing to the Q-set

The Q-set is a structured miniature of the concourse, not a random draw from it. Brown recommends Fisherian experimental-design logic: categorise the concourse along one or more dimensions, then sample across cells so the final set spans the argument space. The reasoning is close to maximum variation sampling — you are deliberately covering the range, not estimating a population.

Empirically, final Q-sets in the healthcare review ranged from 16 to 275 statements, with a median of 42. Two studies did not report the number at all. Aim for coverage and readability rather than a target count, but note the statistical consequence above: a 16-statement Q-set gives SE = 0.25, so almost nothing will load significantly.

3. The P-set

The person sample is small and purposive. Brown observes that P-sets “rarely exceed 50” even in public-opinion work, and the healthcare review reports a typical target of around 40–60 participants. Actual practice is wide: P-sets ranged from 5 to 299 for healthcare consumers (median 33) and from 4 to 710 for providers (median 39).

You recruit people expected to hold different views, not a representative cross-section — the objective is that every viewpoint likely to exist has at least a couple of people who can express it. Because the goal is coverage of perspectives rather than statistical power, the justification you write resembles the logic in data saturation and information power more than a power calculation.

4. The sorting grid

Participants rank the statements onto a grid, usually quasi-normal, under a clear condition of instruction. Brown’s procedure has the participant read everything first, then split the pack into three rough piles (agree / disagree / neither) before placing cards on the grid — a step that materially improves sort quality and is frequently skipped.

Grid width in the healthcare review: −4 to +4 was most common (n = 109, 37.7%), followed by −5 to +5 (n = 90). Twelve studies (4.2%) used no negative values at all, for example +1 to +9 — legitimate, but it changes how participants read the extremes, so it must be reported.

Whether the distribution should be forced (fixed number of cards per column) or free remains contested. The practical argument for forcing is that it obliges genuine trade-offs and makes sorts directly comparable; the argument against is that it can misrepresent people with flat or extreme views. Either way, follow the sort with an interview. Brown is emphatic on this: the sort tells you which statements to ask about, and the ±3 items are the obvious starting point because they are demonstrably the most salient — though items placed at 0 can be revealing precisely because they lack salience.

5. Extraction and rotation

Every sort is correlated with every other, producing a person-by-person correlation matrix, which is then factored. Two live choices follow.

Extraction. Centroid extraction is the historical Q convention (its rotational indeterminacy is the point — it leaves room for judgmental rotation); principal components is the mainstream alternative. In the healthcare review, PCA was used in 110 studies (38.1%), centroid in 91 (31.5%), and in 87 studies (30.1%) the extraction method was not reported or was unclear. Note that centroid factor analysis is mathematically an approximation of principal axis factoring, so the practical difference between the two families is smaller than the rhetoric around it suggests.

Rotation. This choice is not cosmetic. Akhtar-Danesh compared no-rotation against Varimax, Equamax and Quartimax on two real datasets (40 sorts / 19 statements on marijuana legalisation; 33 sorts / 42 statements on childhood obesity). Matching each rotated factor to its unrotated counterpart by largest absolute correlation, he found only 3 distinguishing statements in common between Factor 1 unrotated and its matched Varimax factor in Dataset 1, with the factor scores on even those three “quite different.” The count of Q-sorts loading on that factor went from 22 unrotated to 13 (Varimax), 10 (Equamax) and 20 (Quartimax). His conclusion: “factors can change substantially from one rotation to another.”

Varimax is currently the accepted default, and it dominates practice — 206 of 289 healthcare studies (71.3%), against just four (1.4%) using by-hand rotation and two (0.7%) using both. But 73 studies (25.3%) did not clearly report what rotation, if any, was used. Given the size of the effect above, that omission alone makes a paper hard to appraise. Judgmental (by-hand) rotation is defensible when you have a theoretical reason to look at the data from a particular vantage point — it is Q’s abductive step — but it must be declared and justified.

Reading the output

Significant loadings and defining sorts

The conventional threshold at p < .01 is 2.58 × (1/√N), N again being the number of statements. A published worked example with a 40-statement Q-set gives 2.58 × 1/√40 = 0.4079: any sort loading above that is significant. Brown’s primer offers a looser rule of thumb — roughly 2 to 2.5 times the standard error — so with 20 statements he treats loadings between 0.44 and 0.56 as the grey zone. The two conventions disagree by design; state which you used.

A significant loading is not the same thing as a defining sort. To define a factor, a sort must load significantly on only that factor and account for the majority of its common variance — that is, be more associated with that one factor than with all the others combined. Only defining sorts are merged into the factor array. Sorts that load on two factors, or that load on none, are confounded and non-significant respectively, and both are informative: a large confounded group usually means you have over-extracted.

Factor arrays, distinguishing and consensus statements

Each factor is expressed as a factor array: a single idealised Q-sort computed as a weighted average of its defining sorts, converted to z-scores and mapped back onto the original grid. This is the crucial interpretive difference from R-mode work. In R-mode analysis, factor scores are per-person estimates with well-known indeterminacy problems — see factor scores and when not to use them. A Q factor array is not a person’s score; it is a reconstructed viewpoint, readable as a whole sort.

Distinguishing statements are those placed significantly differently by one factor than by every other; consensus statements are placed non-significantly differently across all factors. Brown’s rough guide is that a difference of 2 grid positions between factor arrays can be treated as significant. Some analysts instead use an effect-size criterion — Akhtar-Danesh used a Cohen’s d of 0.80. Consensus statements are routinely under-reported and are often the most policy-relevant output, because they show where a fractured debate actually agrees.

How many factors to keep

This is where Q studies most often go wrong. The commonly cited criteria are eigenvalue > 1.0, Humphrey’s rule (the cross-product of a factor’s two highest loadings exceeds twice the standard error — sometimes applied loosely as merely exceeding the standard error), and Watts and Stenner’s practical floor of at least two sorts loading significantly and exclusively on the factor. Interpretability is the fourth and final gate.

The eigenvalue > 1 rule is the weakest of these and is criticised in general factor-analytic practice for the same reason it is criticised here — it over-retains. The same caution applies in mainstream component analysis; see the retention rules discussed in principal component analysis in SPSS.

A recent Q study of cognitive errors in clinical decision-making illustrates the risk transparently: from 48 participants the authors retained an eight-factor solution explaining 66% of variance, with a standard error of 0.14, but several of those factors had only two or three sorts loading significantly — a point the authors themselves flag. Factors defined by two sorts are fragile: they can be an artefact of rotation choice as easily as a genuine shared viewpoint. Number of factors identified across the 289 healthcare studies ranged from 0 to 21, and the most common solution was four factors (n = 100, 34.6%).

Variance explained

Total variance explained across those studies ranged from 20.0% to 90.8% (mean 53.4%, SD 11.6) — and 90 studies (31.1%) did not report it at all. There is no threshold to pass. A Q solution is not trying to account for all the variance; it is trying to account for the shared, patterned part of it. Report the figure and let the reader judge.

Software

PQMethod remains the field standard — 186 of 289 healthcare studies (64.4%) — followed by PCQUANL/QUANL (n = 36, 12.5%) and PCQ (n = 21, 7.3%). General-purpose tools appeared rarely: SPSS in 9 studies (3.1%), Q-Assessor in 5 (1.7%), QMethod in 4 (1.4%), and one each for Qanalyze, SAS and Stata. There is also a CRAN R package, qmethod (“Analysis of Subjective Perspectives Using Q Methodology”), and a Stata command, qfactor, for analysts who want the analysis inside an existing scripted pipeline. For online administration, FlashQ was the most common platform among the 33 studies collecting data online. Feature sets and licensing change; check current documentation before committing a study to a tool.

What to report

Churruca and colleagues published a 14-item reporting checklist precisely because so much of this goes unreported. The items that most often go missing, and that a reviewer should insist on:

  • How concourse items were collected, and how they were refined and reduced to the final Q-set
  • Piloting procedure and what changed as a result
  • The condition of instruction — the exact prompt participants sorted against
  • Grid shape and range, and whether the distribution was forced
  • Extraction method and rotation method, both named explicitly
  • The software and version used
  • The criteria used to decide how many factors to extract, rotate and interpret
  • The loading-significance threshold and the formula it came from
  • Variance explained by the retained solution
  • A rich narrative for each factor, supported by statement numbers, array positions and participant quotes

Publishing the full factor loading table — every sort against every retained factor, with defining sorts marked — costs half a page and lets a reader re-run your retention decision. It is the single most useful thing a Q paper can include.

Frequently asked questions

Is Q methodology qualitative or quantitative?

Both, and the label matters less than the inference it licenses. The data are collected qualitatively (a person modelling their own viewpoint) and analysed with standard factor-analytic mathematics. Brown’s framing is that the statistics are “in the background” and that Q is properly a method for the systematic study of subjectivity. Journals and reviewers generally treat it as a mixed method.

How many participants do I need for a Q study?

Fewer than you would expect, and the number is not chosen for statistical power. Around 40–60 is the commonly cited working range; published healthcare studies had medians of 33 (consumers) and 39 (providers). The binding requirement is that each viewpoint you expect to find has enough people to define a factor — a minimum of two sorts loading significantly and exclusively. More participants do not tighten your loading threshold, because that depends on the number of statements.

Why is the standard error based on statements rather than people?

Because of the transposition. In Q the statements are the cases over which each pair of sorts is correlated, so the reliability of the correlation between two people’s sorts depends on how many statements they ranked. Brown gives SE = 1/√N with N as the statement count.

Should I use Varimax or by-hand rotation?

Varimax if you are exploring without a prior hypothesis — it is the accepted default and dominates published practice. By-hand (judgmental) rotation if you have a substantive theoretical reason to view the factor space from a particular angle, which you state in advance and justify in the paper. What you must not do is leave the choice unreported: rotation demonstrably changes which sorts load and which statements distinguish factors.

Can Q methodology results be generalised to a population?

No, and claiming otherwise is the most common overreach in Q papers. Q demonstrates that particular viewpoints exist and describes their internal structure. It says nothing about prevalence. If you need prevalence, the standard follow-on is to convert the factor arrays into a Q-sort-derived instrument and field it on a probability sample.

What is the difference between a distinguishing statement and a consensus statement?

A distinguishing statement is placed significantly differently in one factor’s array than in every other factor’s — it is what makes that viewpoint distinctive. A consensus statement is placed with no significant difference across all retained factors — it marks common ground. Both should be reported; consensus statements frequently carry the practical implications.

What is a confounded Q-sort?

A sort that loads significantly on more than one factor, so it cannot be assigned as a defining sort to any of them. A handful is normal. A large proportion usually signals that you have retained too many factors, or that your Q-set does not cleanly separate the positions in the concourse.

References

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.