Skip to main content
v2026.11,610 entries · CC-BY 4.0

Levels of Evidence: The OCEBM 2011 Table, GRADE, and Where the Pyramid Breaks

A levels-of-evidence hierarchy ranks study design; GRADE rates certainty in one effect estimate for one outcome. This guide sets out what the OCEBM 2011 table actually says, what changed from the 2009 version, the published thresholds at which GRADE upgrades observational evidence or downgrades randomised trials, and what the methodological critique of the evidence pyramid actually argues.

Ask about Levels of Evidence: The OCEBM 2011 Table, GRADE, and Where the Pyramid Breaks

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

Short answer: a levels-of-evidence hierarchy and GRADE are not competing versions of the same tool, and treating them as interchangeable is the single most common error in the way evidence gets labelled. A hierarchy such as the Oxford Centre for Evidence-Based Medicine (OCEBM) 2011 Levels of Evidence ranks study designs as a search shortcut: it tells you where the likely best answer to a question is probably sitting. GRADE rates certainty in a specific effect estimate for a specific outcome after you have found and appraised the studies. The two answer different questions, they operate at different units of analysis, and they can — legitimately, by design — reach opposite conclusions about the same body of evidence.

This page sets out what the OCEBM 2011 table actually says (including what changed from the widely-reproduced 2009 version), where a design hierarchy and GRADE diverge and by how much, and what the published methodological critique of the evidence pyramid actually argues. For how GRADE itself works — the four certainty levels and the five downgrade domains — see the CASRAI dictionary entry on GRADE (evidence certainty rating); this page does not repeat it.

The distinction in one line: design versus estimate

The unit of analysis is where the two schemes part company.

  • A levels-of-evidence hierarchy assigns a level to a study or a body of studies, based mainly on design. Its input is “what kind of study is this?” Its output is a rank. It is fast, it can be applied by someone who has only read the abstract, and it is deliberately built to be usable in the minutes a clinician has at the point of care.
  • GRADE assigns a certainty rating to one estimate of effect, for one outcome, in one review or guideline. Its input is the body of evidence plus a judgement about risk of bias, inconsistency, indirectness, imprecision and publication bias. Its output is high, moderate, low or very low certainty — and the same review will typically carry a different rating for mortality than for quality of life, from the very same trials.

That difference has a practical consequence that catches people out constantly: a single study does not have one GRADE rating. A trial can support high-certainty evidence for one outcome and low-certainty evidence for another. A “Level 1 study” has one level regardless of which outcome you are reading it for. Any sentence of the form “this is Level 1 evidence, therefore high-certainty” is a category error, and the OCEBM working group says so itself.

What the OCEBM 2011 Levels of Evidence table actually is

The 2011 table (formally the Oxford Levels of Evidence 2, current published version 2.1) is a grid. Its rows are seven clinical questions; its columns are five “steps”, corresponding to Levels 1 through 5. You start at Step 1 for your question and move rightwards only if nothing is there.

The seven questions it covers are:

  1. How common is the problem? (prevalence)
  2. Is this diagnostic or monitoring test accurate? (diagnosis)
  3. What will happen if we do not add a therapy? (prognosis)
  4. Does this intervention help? (treatment benefits)
  5. What are the common harms?
  6. What are the rare harms?
  7. Is this early-detection test worthwhile? (screening)

Because the levels are defined per question rather than globally, the design ordering is not the same across rows. Three consequences of that are worth stating explicitly, because they are the parts most secondary summaries drop.

For prevalence, a systematic review is not Level 1

For “how common is the problem?”, the OCEBM 2011 table puts a local and current random-sample survey or census at Level 1, and a systematic review of surveys that allow matching to local circumstances at Level 2. The CEBM introductory document is explicit about why: for questions of local prevalence, current local surveys are the ideal, and this is the stated exception to the general preference for reviews over single studies.

This is the clearest available demonstration that a hierarchy is question-relative rather than a universal ranking of designs. A generic pyramid with systematic reviews painted at the apex gets this case wrong.

An observational study can sit at Level 1 or 2 in the table itself

For treatment benefits, Level 2 in the 2011 table is a randomised trial or an observational study with a dramatic effect. For common harms, Level 1 explicitly includes an observational study with a dramatic effect alongside systematic reviews of randomised trials and of nested case-control studies. The hierarchy is therefore not a pure design ranking even on its own terms — effect magnitude is already baked into it.

Level 5 is mechanism-based reasoning, not expert opinion

In the 2011 table, Level 5 is mechanism-based reasoning, and it exists only for diagnosis, treatment benefits, harms and screening. For prevalence and prognosis, Level 5 is marked not applicable. The 2009 table, by contrast, defined Level 5 as expert opinion without explicit critical appraisal, or reasoning from physiology, bench research or first principles.

The n-of-1 misreading the working group had to correct

CEBM publishes a specific correction on its own Levels page, because Level 1 for treatment benefits is routinely misread. The intended reading is “either n-of-1 randomised trials or systematic reviews of randomised trials”. The wrong reading — which the working group names as wrong — is “either systematic reviews of randomised trials or systematic reviews of n-of-1 trials”. If you cite Level 1 for a treatment benefit, an individual n-of-1 trial in the patient you are asking about qualifies; a systematic review of other people’s n-of-1 trials is not what is meant.

What changed between the 2009 and the 2011 tables

Many pages describing “the Oxford levels of evidence” are still describing the March 2009 table. The two are structurally different documents, and the differences matter if you are citing a level in a manuscript or a guideline.

Feature 2009 (Levels of Evidence 1) 2011 (Levels of Evidence 2, v2.1)
Orientation Rows are levels; columns are five question types Rows are seven clinical questions; columns are five steps
Sub-levels Eleven grades: 1a, 1b, 1c, 2a, 2b, 2c, 3a, 3b, 4, 5 Five levels only; the letter sub-levels are gone
Question coverage Therapy/prevention/aetiology/harm; prognosis; diagnosis; differential diagnosis and symptom prevalence; economic and decision analyses Prevalence; diagnosis and monitoring; prognosis; treatment benefits; common harms; rare harms; screening
Economic analyses A full column of the table Dropped
Harms Bundled with therapy in one column Split into separate common-harms and rare-harms rows
Level 5 Expert opinion without explicit critical appraisal, or reasoning from physiology, bench research or first principles Mechanism-based reasoning; not applicable for prevalence or prognosis
Adjusting a level A minus sign appended to a level to flag an inconclusive answer (wide confidence interval, or a review with troublesome heterogeneity) An explicit grade-down / grade-up footnote (see below)
Attribution Phillips, Ball, Sackett, Badenoch, Straus, Haynes and Dawes from November 1998; updated by Howick, March 2009 OCEBM Levels of Evidence Working Group: Howick, Chalmers, Glasziou, Greenhalgh, Heneghan, Liberati, Moschetti, Phillips, Thornton, Goddard, Hodgkinson

If you cite “Oxford Level 1b”, you are citing the 2009 table, because 2011 has no letter sub-levels. That is a legitimate thing to do — the 2009 document is still published — but say which version you mean.

The footnote that makes the hierarchy stop being a hierarchy

The 2011 table carries a footnote attached to every level. In the working group’s own words, a level “may be graded down on the basis of study quality, imprecision, indirectness (study PICO does not match questions PICO), because of inconsistency between studies, or because the absolute effect size is very small; Level may be graded up if there is a large or very large effect size.”

Read that list against GRADE’s five domains. Study quality is GRADE’s risk of bias. Imprecision, indirectness and inconsistency are named identically. A very small absolute effect is an imprecision-adjacent consideration. Large effect size is one of GRADE’s three rate-up criteria. The 2011 OCEBM table has, in a single footnote, imported most of the GRADE machinery into the hierarchy — which is another way of saying that even its authors did not think design rank alone was sufficient.

The introductory document is blunter still. It states that the Levels are not intended to give a definitive judgement about the quality of evidence, and that there will inevitably be cases where lower-level evidence — an observational study with a dramatic effect — provides stronger evidence than a higher-level study, such as a systematic review of few studies producing an inconclusive result. It also states plainly that the Levels will not give you a recommendation, and that unlike GRADE, the scheme explicitly refrains from making definitive recommendations and can be used even when no systematic review exists.

Where the two schemes actually disagree, with the thresholds

The abstract claim “GRADE can downgrade an RCT and upgrade an observational study” is true but useless without the numbers. Here is what the published criteria actually specify.

Rating up: the two-fold and five-fold thresholds

GRADE guidelines 9 (Guyatt GH, Oxman AD, Sultan S, et al., Journal of Clinical Epidemiology 2011;64(12):1311–1316, PMID 21802902) sets out the rate-up criteria. Its stated thresholds: consider rating up one level when methodologically rigorous observational studies show at least a two-fold reduction or increase in risk, and two levels for at least a five-fold reduction or increase. Rating up is also available where a dose-response gradient is present, and where all plausible confounders or biases would have decreased an apparent treatment effect, or would have created a spurious effect where the results suggest none.

An observational study starts at low certainty under GRADE. A five-fold effect from rigorous observational studies can therefore reach high certainty — the same rating a clean randomised trial gets — without any randomisation anywhere in the evidence base.

Rating down: five RCTs are not automatically high certainty

The mirror case is documented with a worked example in Murad MH, Asi N, Alsawas M, Alahdab F, “New evidence pyramid”, Evidence-Based Medicine 2016;21(4):125–127 (PMID 27339128, PMC4975798, open access). A meta-analysis of five randomised trials of intensive glycaemic control in non-critically ill hospitalised patients found a non-significant mortality reduction, relative risk 0.95 (95% CI 0.72 to 1.25). Allocation concealment and blinding were inadequate in most of the trials. That body of evidence is rated down for methodological limitations and for imprecision — the confidence interval spans substantial benefit and substantial harm. The authors’ conclusion is direct: despite there being five RCTs, this evidence should not be rated high in any pyramid.

The same paper gives the upgrade case in one sentence: certainty in the benefit of hip replacement for disabling hip osteoarthritis is rated up despite resting on non-randomised observational studies.

The empirical result behind the disagreement

The strongest evidence that design rank is a weak proxy for effect reliability comes from meta-epidemiology. Toews I, Anglemyer A, Nyirenda JL, et al., “Healthcare outcomes assessed with observational study designs compared with those assessed in randomized trials: a meta-epidemiological study”, Cochrane Database of Systematic Reviews 2024;1:MR000034 (PMID 38174786) — the third version of a review first published in 2014 — pooled 39 systematic reviews and eight overviews, covering 2,869 randomised trials with 3,882,115 participants and 3,924 observational studies with 19,499,970 participants.

The pooled ratio of ratios comparing effect estimates from RCTs with those from observational studies was 1.08 (95% CI 1.01 to 1.15), which the review authors describe as no difference or a very small difference. They rated the certainty of that finding as low, and the review’s own conclusion is that factors other than study design — differences in population, intervention, comparator and outcome, and the level of statistical heterogeneity — need to be considered when RCT and observational results disagree.

Read carefully, that is not a claim that observational studies are as good as trials. It is a claim that knowing only the design label predicts the effect estimate poorly, which is precisely the input a levels-of-evidence hierarchy uses and precisely what GRADE refuses to rely on alone. Note also that the certainty of this very finding is rated low, by GRADE — a useful illustration of the framework applied to itself.

Why the pyramid is not enough: what the published critique argues

The evidence pyramid as usually drawn — bench research and case series at the base, then case-control and cohort studies, then randomised trials, then systematic reviews and meta-analysis at the apex — is a teaching diagram, not a published standard. It has no single authoritative source, which is part of why the critique of it is scattered. The substantive published objections are these.

1. Straight lines imply that design determines certainty

Murad and colleagues (2016) propose replacing the straight horizontal lines between the pyramid’s tiers with wavy lines, precisely to represent GRADE’s rating up and down across the certainty domains. Their argument is that study design alone is insufficient as a surrogate for risk of bias, because methodological limitations, imprecision, inconsistency and indirectness are independent of design and can affect evidence from any design.

2. The apex is the wrong place for systematic reviews

Their second proposed modification is to remove systematic reviews from the top of the pyramid entirely and treat them as a lens through which the other evidence is viewed — appraised and applied — rather than as a tier above randomised trials. The rationale draws on the two-step appraisal framework in the JAMA Users’ Guides: first assess whether the review process itself is credible (comprehensive search, rigorous selection), then assess certainty in the evidence using GRADE. A meta-analysis of well-conducted low-risk-of-bias trials cannot be equated with a meta-analysis of observational studies at higher risk of bias, yet a pyramid puts both in the same top tier.

The paper’s illustration: a meta-analysis of 112 surgical case series in patients with thoracic aortic transection reported mortality of 9% after endovascular repair, 19% after open repair and 46% with non-operative management (p<0.01). It is a meta-analysis. It belongs nowhere near the apex.

3. Level labels do not survive re-evaluation

Murad et al. cite a documented consequence of conflating the two schemes: the American Heart Association treats evidence derived from meta-analyses as level “A”, its highest-confidence label, but when a set of such evidence was re-evaluated using GRADE it turned out to be spread across high, moderate, low and very low quality (Murad MH, Altayar O, Bennett M, et al., Journal of Clinical Epidemiology 2014;67:65–72). A single hierarchy label was concealing a four-way spread in actual certainty.

4. Meta-analytic results depend on analytic choices, not just on inputs

Dechartres A, Altman DG, Trinquart L, et al., “Association between analytic strategy and estimates of treatment outcomes in meta-analyses”, JAMA 2014;312(6):623–630, evaluated 163 meta-analyses and found that estimates of treatment outcomes differed substantially depending on which analytical strategy was used. Berlin JA and Golub RM’s accompanying editorial, “Meta-analysis as evidence: building a better pyramid” (JAMA 2014;312(6):603–605), is where much of this argument was first put to a general medical audience. Heterogeneity and model choice are not eliminated by sitting at the top of a pyramid.

5. The hierarchy was never claimed to do this job

The critique is not an attack on the OCEBM working group, who state the limits themselves. The introductory document opens by saying that no evidence ranking system or decision tool can be used without a healthy dose of judgement and thought, and it describes the Levels as a hierarchy of the likely best evidence and a short-cut for busy clinicians, researchers and patients. The failure mode is not the tool; it is using a search heuristic as a certainty verdict.

Decision rule: which scheme for which job

What you are doing Use Why
Deciding where to start searching for the best available answer OCEBM 2011 Levels Built as a search shortcut, question by question; works even where no systematic review exists
Describing the design of an included study in a review’s characteristics table Plain design terminology, not a level A level compresses design and quality into one number and loses both
Stating how confident readers should be in a pooled result for one outcome GRADE Certainty is per-outcome and depends on the five domains, not the design label
Writing a clinical practice guideline recommendation GRADE (or the scheme the commissioning body mandates) OCEBM explicitly refrains from producing recommendations; see GRADE panels and guideline authorship
Appraising an individual study’s internal validity A design-specific appraisal instrument See choosing a critical appraisal tool and RoB 2 / ROBINS-I
Justifying a claim that observational evidence is strong GRADE rate-up criteria, with the effect magnitude stated The two-fold and five-fold thresholds are the published, checkable criteria

Reporting checklist

If you are going to put a level or a certainty rating in a manuscript, a protocol or a guideline, these are the items that make it auditable:

  • Name the scheme and its version. “OCEBM 2011 Levels of Evidence (v2.1)” or “OCEBM March 2009”, never a bare “Level 1”.
  • Name the question type. A 2011 level is meaningless without saying whether it is a treatment-benefit, harms, prognosis, diagnosis, prevalence or screening level.
  • Say whether you graded up or down, and on what basis. The 2011 footnote permits it; an unqualified level implies you did not consider it.
  • State GRADE certainty per outcome, not per study. One rating per outcome in the summary-of-findings table, with the domains that drove any downgrade named.
  • Where you rate up, give the effect magnitude. A reader cannot check a rate-up against the two-fold or five-fold criterion without it.
  • Never map one scheme onto the other. There is no defensible crosswalk from “Level 1” to “high certainty”; the AHA level “A” re-evaluation above is the documented reason why.
  • Cite the primary documents. The OCEBM table is intended to be read alongside its introductory and background documents; CEBM says so on the table itself.

Common mistakes

  • Describing the 2009 table as “the OCEBM levels”. If your description contains 1a, 1b or 1c, you are describing the 2009 document.
  • Drawing a generic pyramid and calling it OCEBM. The OCEBM levels have never been published as a pyramid. They are a question-by-step grid, and for prevalence the ordering is not the pyramid’s ordering.
  • Treating “systematic review” as automatically top-tier. A systematic review of weak studies inherits their weaknesses; see also MOOSE for meta-analyses of observational studies and AMSTAR 2.
  • Assigning one GRADE rating to a whole review. Ratings are per outcome. A review with six outcomes can carry six different ratings.
  • Using a level to justify a recommendation. OCEBM states explicitly that the Levels will not provide a recommendation.
  • Forgetting publication bias. It is a GRADE downgrade domain with no counterpart in a design hierarchy at all — see publication bias and how to detect it.
  • Ignoring indirectness. The 2011 footnote names it as a PICO mismatch between the study and your question; if your PICO does not match, the level comes down regardless of design.

Frequently asked questions

What is Level 1 evidence?

It depends entirely on the question and the version of the table. In the OCEBM 2011 table, Level 1 for treatment benefits is a systematic review of randomised trials or an n-of-1 randomised trial; for prevalence it is a local and current random-sample survey or census; for prognosis it is a systematic review of inception cohort studies; for diagnostic accuracy it is a systematic review of cross-sectional studies with a consistently applied reference standard and blinding. There is no single design that is “Level 1”.

Is GRADE a levels-of-evidence system?

No. GRADE produces a certainty rating (high, moderate, low, very low) for an effect estimate on a specific outcome, and separately a strength of recommendation. It does not rank study designs, although design determines the starting point before the domains are applied.

Can an observational study ever be high-certainty evidence?

Yes. Under the GRADE rate-up criteria, methodologically rigorous observational studies showing at least a five-fold change in risk can be rated up two levels, which takes low certainty to high. The OCEBM 2011 table reaches a similar place by a different route, placing an observational study with a dramatic effect at Level 2 for treatment benefits and within Level 1 for common harms.

Can a randomised trial ever be low-certainty evidence?

Yes, and this is common rather than exotic. Risk of bias (for example inadequate allocation concealment or blinding), imprecision, indirectness, inconsistency across trials, and suspected publication bias each permit a downgrade. The glycaemic-control example above is a published case of five RCTs that should not be rated high.

Why is the evidence pyramid criticised if it is still taught?

Because it is a good teaching device for one idea (not all evidence is equally reliable) and a poor instrument for the decision it gets used for (how confident should I be in this number). The Murad et al. proposal keeps the pyramid as a teaching tool while changing what it depicts: wavy tier boundaries for GRADE’s rating up and down, and systematic review repositioned as the lens rather than the apex.

Do I have to choose one scheme?

No, and the two are complementary if you keep their jobs separate: use the hierarchy to decide where to look, then use GRADE to say how much the reader should believe what you found. Problems arise only when a search heuristic is reported as a certainty verdict.

Which version of the OCEBM table should I cite?

Version 2.1 of the 2011 table is the current published version, and CEBM asks that the table be cited together with its introductory and background documents rather than read on its own. The March 2009 table remains available and is still the correct citation if your levels use letter sub-levels.

Primary sources

  • OCEBM Levels of Evidence Working Group. The Oxford Levels of Evidence 2 (table, v2.1). Oxford Centre for Evidence-Based Medicine. cebm.ox.ac.uk/resources/levels-of-evidence/ocebm-levels-of-evidence
  • Howick J, Chalmers I, Glasziou P, Greenhalgh T, Heneghan C, Liberati A, Moschetti I, Phillips B, Thornton H. The 2011 Oxford CEBM Levels of Evidence (Introductory Document). Introductory document
  • Oxford Centre for Evidence-Based Medicine. Levels of Evidence (March 2009). 2009 table
  • Guyatt GH, Oxman AD, Sultan S, et al. GRADE guidelines: 9. Rating up the quality of evidence. J Clin Epidemiol 2011;64(12):1311–1316. PMID 21802902.
  • Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ 2008;336(7650):924–926. PMID 18436948.
  • Balshem H, Helfand M, Schünemann HJ, et al. GRADE guidelines: 3. Rating the quality of evidence. J Clin Epidemiol 2011;64(4):401–406. PMID 21208779.
  • Murad MH, Asi N, Alsawas M, Alahdab F. New evidence pyramid. Evid Based Med 2016;21(4):125–127. PMID 27339128; PMC4975798.
  • Toews I, Anglemyer A, Nyirenda JL, et al. Healthcare outcomes assessed with observational study designs compared with those assessed in randomized trials: a meta-epidemiological study. Cochrane Database Syst Rev 2024;1:MR000034. PMID 38174786.
  • Concato J, Shah N, Horwitz RI. Randomized, controlled trials, observational studies, and the hierarchy of research designs. N Engl J Med 2000;342(25):1887–1892. PMID 10861325.
  • Dechartres A, Altman DG, Trinquart L, et al. Association between analytic strategy and estimates of treatment outcomes in meta-analyses. JAMA 2014;312(6):623–630.
  • Berlin JA, Golub RM. Meta-analysis as evidence: building a better pyramid. JAMA 2014;312(6):603–605.

Related CASRAI resources

This guide sits in CASRAI’s publishing and research assessment cluster, alongside the evidence-synthesis material. For the mechanics this page deliberately does not repeat, start with GRADE (evidence certainty rating). For the design vocabulary the hierarchy ranks, see randomised controlled trial, cohort study, case-control study and the comparison of randomised trials versus observational studies. For the synthesis workflow that produces the estimates GRADE rates, see PRISMA and systematic review methodology and meta-analysis. For the judgements that drive downgrades, see confounding, internal versus external validity and effect size.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →