Skip to main content
v2026.11,772 entries · CC-BY 4.0

SUCRA Treatment Rankings: What the Number Does (and Doesn’t) Tell You

SUCRA compresses a treatment’s entire ranking distribution into one 0-100% number. Here’s what it represents, how it differs from probability-of-being-best, and the mean-vs-distribution limitation methodologists emphasize.

Written and maintained by CASRAI Editorial Board

Last updated

SUCRA (Surface Under the Cumulative Ranking curve) compresses an entire treatment’s ranking distribution — every probability of finishing 1st, 2nd, 3rd, and so on across a whole network meta-analysis — into a single number between 0% and 100%. That compression is exactly what makes SUCRA useful for a quick comparison table, and exactly why it can mislead a reader who treats it as a stand-in for the actual effect estimates. This guide covers what the number represents, how it is built from the underlying rank probabilities, and the limitation methodologists most consistently emphasize: SUCRA reports the mean of a distribution, and a mean can look identical for a precisely estimated effect and a highly uncertain one.

What SUCRA actually summarizes

In a network meta-analysis comparing three or more treatments, each treatment doesn’t get a single rank — it gets a full probability distribution across every possible rank. A treatment might have a 40% chance of being ranked 1st, a 35% chance of being ranked 2nd, a 20% chance of 3rd, and a 5% chance of being last. Presenting that distribution treatment-by-treatment (often as a “rankogram,” a bar chart of rank probabilities) is complete but unwieldy once a network has more than four or five treatments and a reader wants to compare all of them at a glance.

SUCRA was introduced specifically to solve that presentation problem. It was defined by Salanti, Ades, and Ioannidis in “Graphical methods and numerical summaries for presenting results from multiple-treatment meta-analysis: an overview and tutorial” (Journal of Clinical Epidemiology, 2011), a paper written explicitly to give reviewers a small set of graphical and numerical tools — rankograms, cumulative ranking curves, and SUCRA among them — for communicating ranking results from a network meta-analysis without forcing every reader to parse a full rank-probability matrix.

The construction is the cumulative version of the rankogram: for each treatment, sum the probability of being ranked 1st, then 1st-or-2nd, then 1st-or-2nd-or-3rd, and so on through every rank. Plotted against rank position, that cumulative probability traces a curve. SUCRA is the area under that curve, normalized so a treatment that is certain to be the single best option scores 100%, and a treatment certain to be the single worst scores 0%. A treatment sitting in the unremarkable middle of a network, with its rank probability spread fairly evenly across positions, lands somewhere near 50%.

SUCRA is not “probability of being best”

The single most common misreading is treating SUCRA as though it were the probability that a treatment is the single best option in the network — sometimes written as P(best) or, in a Bayesian analysis run via MCMC, the proportion of simulation iterations in which a treatment ranked first. That number exists and is reported alongside SUCRA in most network meta-analyses, but it is a different quantity.

P(best) only looks at the top rank. SUCRA looks at the whole distribution: a treatment that is very rarely ranked 1st but consistently ranked 2nd or 3rd, and almost never ranked near the bottom, can carry a high SUCRA despite a low or even negligible P(best). Conversely, a treatment with a genuinely high chance of being best but a long tail of occasionally ranking last — a “high variance” treatment — can score a lower SUCRA than its P(best) alone would suggest. SUCRA is designed to reward consistency across the whole ranking, not just a shot at the top spot, which is exactly why it became the more commonly reported summary once analysts wanted one number rather than a full table of rank probabilities.

Bayesian and frequentist versions compute it differently

SUCRA was originally a Bayesian quantity: it depends on simulated rank probabilities, typically produced via Markov chain Monte Carlo sampling from the posterior distribution of relative treatment effects, following the modeling framework set out in NICE DSU Technical Support Document 2. Rücker and Schwarzer later (BMC Medical Research Methodology, 2015) proposed the P-score as a frequentist analogue that answers the same question — how does a treatment’s whole ranking distribution compare to the rest of the network — but computes it analytically from the point estimates and standard errors of a normal approximation, without simulation. In a well-behaved network the two numbers are close; the choice between them tracks the choice between a Bayesian and a frequentist approach to the network meta-analysis itself, not a separate methodological decision.

The limitation that actually matters: SUCRA is a mean, not a distribution

SUCRA (like P(best) and P-score) is a summary computed from the mean of each treatment’s rank distribution. That is precisely what makes it compact, and precisely what it discards: the width of the distribution the mean was computed from. Two treatments can land at nearly identical SUCRA values — say, both around 75% — for entirely different underlying reasons. One might have a tightly clustered rank distribution because its treatment effect is estimated with real precision: a large evidence base, a well-connected part of the network, a narrow confidence or credible interval around its effect estimate. The other might land at the same SUCRA purely because its rank probabilities are spread out roughly symmetrically around the same average position — a treatment supported by one small, imprecise trial, with a confidence interval wide enough to plausibly cover almost any true effect, best or worst included.

Nothing in the SUCRA value itself distinguishes those two cases. A ranking table sorted by SUCRA alone will place them side by side as though they were equally good options, when one is a genuinely well-supported result and the other is a coin flip that happened to average out to the same rank. This is the caveat methodologists raise most consistently about ranking metrics in network meta-analysis: a treatment ranking should never be read as a substitute for looking at the actual effect estimates and their uncertainty. The heterogeneity and imprecision that a forest plot or a table of pairwise effect estimates makes visible at a glance is exactly what a SUCRA-only summary hides.

This is also why the guidance from the Cochrane Handbook and NICE DSU on presenting network meta-analysis results treats ranking metrics as a complement to, never a replacement for, the full set of relative effect estimates with their confidence or credible intervals. A rankogram or cumulative ranking curve is a genuine improvement over SUCRA alone for exactly this reason: it at least shows the shape of the distribution the single number was extracted from, even if it still doesn’t show you the effect-size uncertainty directly.

How to read a SUCRA table without being misled by it

  • Always pull up the effect estimates behind the ranking. Before treating a SUCRA difference as meaningful, check the actual pairwise or network effect estimates (odds ratios, mean differences, hazard ratios) and their intervals for the treatments in question.
  • Check interval width, not just the point estimate. A treatment with a wide confidence or credible interval that happens to be centered near the top of the network can post a respectable SUCRA while carrying far less certainty than a treatment with a narrower interval and a similar or even slightly lower SUCRA.
  • Look at the rankogram or cumulative ranking curve, not the summary number alone. A flat, spread-out cumulative curve behind a mid-range SUCRA tells a different story than a steep curve landing at the same value.
  • Treat small SUCRA differences as noise. Given that SUCRA is itself derived from uncertain effect estimates, a few percentage points of separation between two treatments is rarely a basis for a clinical or policy recommendation on its own.
  • Check the underlying evidence base per treatment. A treatment’s rank distribution is often wide simply because it has thin direct or indirect evidence in the network — a small number of contributing studies, or comparisons made mostly through indirect evidence rather than head-to-head trials.

Where SUCRA fits in reporting a network meta-analysis

The PRISMA extension for network meta-analysis asks authors to report how treatments were ranked and to interpret ranking results in the context of the effect estimates and the certainty of the evidence, rather than as a standalone finding. In practice that means a well-reported network meta-analysis presents a SUCRA (or P-score) ranking alongside — never instead of — a forest plot or league table of pairwise effects, a discussion of heterogeneity and inconsistency, and, ideally, a GRADE-style certainty-of-evidence assessment per comparison. Choosing between the fixed-effect and random-effects assumptions, and between a Bayesian or frequentist estimation framework, both affect the effect estimates that SUCRA is ultimately summarizing — so those modeling choices belong in the same methods section as the ranking table, not off to one side of it.

For readers assembling or critically appraising a systematic review that includes a network meta-analysis, the Cochrane Handbook‘s chapter on comparing multiple interventions and the underlying meta-regression and inconsistency-checking literature are the right next stops for the statistical detail this guide doesn’t cover — SUCRA’s transitivity assumption in particular is worth checking before trusting any ranking the network produces.

Frequently asked questions

Is a higher SUCRA always better?

Higher means the treatment’s whole rank distribution sits closer to “best” across the network, so yes, all else equal, a higher SUCRA reflects a more favorable ranking position. The caution is in “all else equal” — a higher SUCRA built on a wide, uncertain rank distribution is not automatically a stronger recommendation than a lower SUCRA built on a tight, well-supported one. Check the effect estimates before treating the ranking as decisive.

What SUCRA value counts as “good”?

There’s no fixed threshold in the methodology — SUCRA is a relative, within-network ranking, not an absolute quality score. A treatment scoring 80% in a network of three options and one scoring 80% in a network of ten are not directly comparable claims about the same thing; SUCRA only orders treatments against the others actually included in that specific network.

Can two treatments have the same SUCRA but different clinical conclusions?

Yes, and this is the central limitation this guide covers. Two treatments can post nearly identical SUCRA values while one has a precisely estimated effect (narrow interval, strong evidence base) and the other has a highly uncertain one (wide interval, thin or indirect evidence). The SUCRA value alone doesn’t distinguish them — the effect estimates and their intervals do.

Is SUCRA the same thing as a Bayesian credible interval?

No. A credible interval (or, in a frequentist network meta-analysis, a confidence interval) describes the uncertainty around a specific effect estimate for one comparison. SUCRA describes where a treatment tends to fall across the entire ranking of all treatments in the network. They answer related but different questions, and a complete report includes both rather than substituting one for the other.

Does SUCRA apply outside network meta-analysis?

SUCRA as originally defined is specific to network (multiple-treatment) meta-analysis, where ranking more than two options simultaneously is the point. A standard two-arm, pairwise meta-analysis doesn’t produce a ranking distribution in the same sense, so SUCRA isn’t a metric that applies there.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about SUCRA Treatment Rankings: What the Number Does (and Doesn’t) Tell You

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.