Examples
Worked examples
- Is an instance
A systematic review of randomized trials on a drug's effect starts at High certainty; after the review identifies substantial inconsistency between trial results and a serious risk of bias in several included trials, the panel downgrades the certainty rating to Low, and a clinical practice guideline built on that evidence issues a weak (conditional) rather than strong recommendation.
Counter-examples
Looks similar, but isn't
- Not an instance
A journal's own internal 'level of evidence' scale based only on study design (e.g. 'Level I: randomized trial') is not GRADE: GRADE additionally requires an explicit domain-by-domain assessment of risk of bias, inconsistency, indirectness, imprecision, and publication bias for each outcome, not just a design-based starting tier.
Editorial commentary
The GRADE approach was developed starting in 2000 to address inconsistencies among the many evidence-grading systems then in use across specialties and organizations. It is now the dominant approach used by Cochrane systematic reviews, WHO guidelines, and most major clinical-practice-guideline developers.
The four certainty levels
- High — further research is very unlikely to change confidence in the estimate of effect.
- Moderate — further research is likely to have an important impact on confidence and may change the estimate.
- Low — further research is very likely to have an important impact on confidence and is likely to change the estimate.
- Very Low — any estimate of effect is very uncertain.
Downgrading and upgrading factors
Certainty starts from study design and is adjusted using five downgrading domains — risk of bias, inconsistency, indirectness, imprecision, and publication bias — and, for observational evidence specifically, up to three upgrading factors: a large effect size, a dose-response gradient, and plausible residual confounding that would work against (rather than exaggerate) the observed effect.
From certainty to recommendation strength
GRADE explicitly separates certainty of evidence from the strength of a recommendation: a guideline panel weighs the certainty rating alongside the balance of benefits and harms, patient values and preferences, and resource use to issue a strong or weak/conditional recommendation. See CASRAI’s related content on certainty of evidence and interpretation and the guide to clinical practice guideline authorship and GRADE panel criteria, which covers panel composition and byline conventions specifically.
Frequently Asked Questions
Does GRADE apply only to randomized trial evidence?
No — GRADE assesses evidence from any study design, including observational studies, which start at Low certainty but can be upgraded based on the factors above. Because certainty is rated per outcome rather than read off a design label, GRADE and a design-ranked hierarchy such as the Oxford levels can reach different verdicts on the same body of evidence; see where levels of evidence and GRADE diverge.
Is GRADE the same as the evidence used in a systematic review’s risk-of-bias assessment?
Related but distinct: risk-of-bias tools (e.g. Cochrane’s RoB 2) assess individual studies; GRADE assesses the certainty of the body of evidence for a specific outcome across the included studies as a whole, with risk of bias as one of its five input domains.
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="GRADE (Evidence Certainty Rating)"
vocab-term-identifier="https://casrai.org/dictionary/term/grade-evidence-rating" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/grade-evidence-rating",
"name": "GRADE (Evidence Certainty Rating)",
"identifier": "https://casrai.org/dictionary/term/grade-evidence-rating",
"description": "GRADE (Grading of Recommendations Assessment, Development and Evaluation) is a structured, transparent methodology, developed by the international GRADE Working Group beginning in 2000, for rating the certainty of a body of evidence -- typically from a systematic review -- as High, Moderate, Low, or Very Low, and for connecting that certainty rating to the strength of a resulting clinical or policy recommendation. Certainty starts from study design (randomized evidence starts High, observational evidence starts Low) and is then downgraded for risk of bias, inconsistency, indirectness, imprecision, or publication bias, or upgraded for a large effect, dose-response gradient, or plausible confounding working against the observed effect.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/clinical-research#set",
"url": "https://casrai.org/dictionary/term/grade-evidence-rating",
"sameAs": [],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"dateModified": "2026-08-26T20:11:19",
"inLanguage": "en"
}






