The Delphi method is a structured technique for building expert consensus on a question that resists direct measurement — a forecast, a priority ranking, a clinical or policy judgment — by putting the same question to a panel of experts across two or more anonymous, iterated rounds, feeding back the group’s aggregated responses between rounds, and letting individual judgments converge (or clearly fail to converge) as a result. It was developed at the RAND Corporation in the early 1950s by Olaf Helmer, Norman Dalkey and Nicholas Rescher, originally for military and technological forecasting, and has since become a standard tool in health services research, clinical guideline development, research-priority setting, and technology forecasting more broadly.
Three features distinguish it from an ordinary panel discussion or committee vote: anonymity (panelists respond independently and never see who said what), controlled feedback (a facilitator summarizes the group’s response statistically or thematically between rounds), and iteration (panelists see the group’s aggregated position and are given the chance to revise their own answer, or explain why they haven’t). The goal is to let expert judgment converge on its own merits, without the status effects, groupthink pressure, or loudest-voice-wins dynamics that face-to-face panels are prone to.
How the Delphi Method Works: The Round Structure
A classical Delphi study runs through a fixed sequence of steps, repeated across rounds until a stopping rule (see below) is met:
- Define the question and select the panel. The facilitator drafts a precise question or set of items (e.g., “rate the priority of each of these 12 research infrastructure investments”) and recruits a panel of subject-matter experts.
- Round 1: open or structured elicitation. Panelists respond independently — in a classical Delphi, Round 1 is often open-ended, generating the initial list of items; in a modified Delphi, the facilitator supplies a pre-defined item list drawn from a literature review, so Round 1 is already a rating round.
- Aggregate and feed back. The facilitator compiles all responses, typically as a summary statistic (median, mean, or percentage distribution) for each item, and returns this summary to every panelist along with their own prior response — never with individual attributions.
- Round 2 (and further rounds): re-rate with feedback. Panelists see where their own answer sat relative to the group and are asked to re-rate each item, with the option to leave written justification if they remain an outlier.
- Check for consensus or stability. After each round, the facilitator applies a pre-specified consensus definition (see below). If it isn’t met and a stopping rule hasn’t been triggered, another round runs.
- Report results. The final round’s aggregated ratings, along with the level of agreement reached and any items that failed to reach consensus, are reported — ideally alongside enough procedural detail (panel composition, response rates per round, exact consensus definition) that the study can be appraised or replicated.
Worked Example: Ranking Research Infrastructure Priorities
The following is an illustrative composite, not a real study, provided to show how the numbers move round to round.
Suppose a research office runs a three-round Delphi with 16 panelists to rank five candidate investments on a 1-9 priority scale (9 = highest priority), with consensus pre-defined as an interquartile range (IQR) of 2 or less on the 9-point scale.
| Item | Round 1 median (IQR) | Round 2 median (IQR) | Round 3 median (IQR) | Consensus reached? |
|---|---|---|---|---|
| Electronic lab notebook rollout | 7 (4) | 8 (2) | 8 (1) | Yes, Round 3 |
| Shared instrumentation core facility | 8 (3) | 8 (2) | 8 (2) | Yes, Round 2 |
| Institutional data repository | 6 (5) | 6 (4) | 6 (4) | No — stopped at Round 3 |
| Statistical consulting service | 5 (3) | 6 (2) | 6 (2) | Yes, Round 2 |
| Open-access publication fund | 7 (5) | 7 (3) | 7 (2) | Yes, Round 3 |
Two items reached the IQR-2 threshold by Round 2 and were dropped from further rounds; two more converged by Round 3. The data repository item never converged — its Round 3 IQR of 4 is reported as-is, alongside a note that panelists were split, rather than forced toward an artificial single number. That non-convergence is itself a valid and often useful finding: it tells the commissioning office the panel genuinely disagrees, rather than that consensus simply hasn’t been reached yet.
Selecting and Sizing the Expert Panel
There is no single mandated panel size, and unlike a probability sample, Delphi panels are deliberately non-random: panelists are selected purposively for relevant expertise, not to be statistically representative of a population. Reported panel sizes vary widely by field and question scope, but two practical constraints tend to set the range:
- Small, homogeneous expert panels (roughly 8-15 people) are common in clinical and technical Delphi studies where the pool of genuinely qualified experts is itself small, and where fewer rounds may be needed to reach a stable answer.
- Larger, more heterogeneous panels (30 or more) are used when the goal is to capture a broader range of stakeholder perspectives — patients and clinicians, or researchers and administrators, on the same panel — and these larger panels typically need three or more rounds to stabilize.
Panel composition should be justified and reported, not just its size: who was eligible, how they were identified (professional societies, citation searches, snowball referral), and what the response rate was at each round. Attrition across rounds is a genuine methodological risk — a panel that shrinks substantially between Round 1 and the final round should prompt scrutiny of whether the final “consensus” reflects only the most persistent respondents.
Defining and Measuring Consensus
“Consensus” has to be defined numerically before the study starts, not judged impressionistically after the last round — otherwise a facilitator can unconsciously declare consensus whenever the results look tidy. A 2014 systematic review of Delphi consensus definitions (Diamond et al., published in the Journal of Clinical Epidemiology) found wide variation in practice across published studies, but identified the following as the most common approaches:
| Consensus definition | How it works | Typical threshold |
|---|---|---|
| Percent agreement | Share of panelists rating an item within a defined band (e.g., 7-9 on a 9-point scale) | 70-80% is common; the review found 75% was the median threshold used |
| Interquartile range (IQR) | Spread of the middle 50% of ratings around the median | IQR of 1-2 on a 9-point scale is common |
| Coefficient of variation | Standard deviation divided by the mean, tracked round to round | Declining CV across rounds used as evidence of convergence rather than a fixed cutoff |
| Stability | Change in an individual panelist’s rating between two consecutive rounds falls below a set threshold | Used alongside, not instead of, a group-level agreement measure |
Whichever definition is used, it should be specified in the study protocol before Round 1, applied consistently, and reported alongside the final results — including any items that failed to reach it.
Stopping Rules: When to End the Rounds
Delphi studies stop on one of two broad grounds, and the choice should be decided in advance:
- Consensus achieved. The pre-defined threshold (percent agreement, IQR, or stability criterion) is met for all or a specified proportion of items, and the study ends there even if that happens in Round 2.
- A fixed round limit is reached. Many published studies simply cap the process at a pre-specified number of rounds — commonly two to four — regardless of whether full consensus is reached on every item, both to control panelist burden and attrition and because response quality and engagement tend to fall off after three or four rounds.
A minimum of two rounds is generally required for the method to do what it’s designed to do — a single round with no feedback loop is not a Delphi study, it’s just a survey. Beyond that minimum, more rounds buy diminishing returns: most of the movement in panelist ratings happens between Rounds 1 and 2, with Round 3 and beyond producing progressively smaller shifts. That pattern is part of why many studies pre-commit to stopping once ratings are stable across two consecutive rounds, rather than chasing full agreement indefinitely.
Modified and Real-Time Delphi Variants
| Variant | What’s different from classical Delphi | Typical use |
|---|---|---|
| Modified (or “Ranking-Type”) Delphi | Round 1 uses a pre-supplied item list (from a literature review or prior work) instead of an open-ended elicitation, skipping straight to rating | Saves a round when the item space is already reasonably well defined |
| RAND/UCLA Appropriateness Method | A modified Delphi combining two anonymous rating rounds with a moderated in-person or virtual discussion between them | Widely used in clinical appropriateness and guideline panels |
| Policy Delphi | Explicitly seeks to structure and expose disagreement and its underlying arguments, rather than converge on a single number | Contested policy questions where the range of positions is itself the useful output |
| Real-time (or e-Delphi) | Panelists see aggregated group feedback and can revise their own rating within a single continuous online session rather than waiting for discrete rounds | Reduces the weeks-to-months timeline of a multi-round postal or email Delphi to hours or days |
Strengths and Limitations
The Delphi method’s core strength is structural: anonymity removes dominance and status effects that distort face-to-face panels, and controlled feedback lets genuine expert judgment shift in response to the group’s collective view without social pressure to simply defer to the most senior voice in the room. It is well suited to questions where evidence is incomplete or mixed and expert judgment has to fill the gap — exactly the situation where an unstructured committee vote is most vulnerable to groupthink.
Its limitations are equally structural. Panel selection is subjective and can bias results toward whichever expert community the facilitator draws from; there is no agreed standard for panel size or number of rounds, which makes cross-study comparison difficult; attrition across rounds can silently shift who the “consensus” actually represents; and the process measures convergence of expert opinion, not ground truth — a well-run Delphi can produce a stable, well-documented consensus that is nonetheless wrong. It should be treated as a method for structuring and documenting expert judgment under uncertainty, not as a substitute for empirical evidence where empirical evidence is actually available.
Frequently Asked Questions
How many rounds does a Delphi study need?
A minimum of two, since the feedback-and-revision step is what defines the method. Most published studies run two to four rounds; three is common for larger or more heterogeneous panels. Studies typically stop either when a pre-defined consensus threshold is met or when a pre-specified round limit is reached, whichever comes first.
What’s a good Delphi panel size?
There is no single correct number. Small, expertise-constrained panels (roughly 8-15) are common in clinical and technical studies; larger, more heterogeneous stakeholder panels (30+) are used when breadth of perspective matters more than depth of specialization in any one panelist. The panel’s selection criteria and response rate per round matter more for credibility than the raw headcount.
How is consensus actually measured?
It has to be defined numerically in advance — most commonly as a percent-agreement threshold (often 70-80%), an interquartile range cutoff on a rating scale, or a round-to-round stability criterion — and applied consistently, with any items that don’t reach it reported rather than dropped silently.
What’s the difference between the Delphi method and a focus group?
A focus group is a single, face-to-face, moderated group discussion where participants hear and react to each other in real time. A Delphi panel never meets: rounds are anonymous and asynchronous, and participants respond to a statistical or thematic summary of the group rather than to each other directly. Delphi trades the richness of live group interaction for freedom from status effects and groupthink.
What’s the difference between a classical and a modified Delphi?
A classical Delphi’s first round is open-ended, generating the item list from the panel itself. A modified Delphi starts Round 1 with a pre-defined item list — usually built from a literature review — so the panel begins rating rather than generating items, which shortens the overall process by a round.
Related CASRAI Resources
The Delphi method sits alongside other techniques for generating qualitative and semi-structured data from people rather than instruments. See Designing Semi-Structured Interviews for one-on-one elicitation, Coding Qualitative Interview Data and Thematic Analysis: A Step-by-Step Guide for analyzing what a Delphi round or interview produces, and Grounded Theory for a related inductive approach to building explanatory frameworks from qualitative data. For panel recruitment logic, Sampling Methods: Probability and Non-Probability Types covers the purposive and snowball approaches most Delphi panels rely on, and Mixed Methods Research covers combining a Delphi round with quantitative survey data. Qualitative Research provides the broader methodological context this technique sits within.







