Written and maintained by CASRAI Editorial Board
Last updated
AMSTAR 2 (A MeaSurement Tool to Assess systematic Reviews, version 2) is a 16-item critical-appraisal instrument for evaluating the methodological quality of a systematic review of randomized and/or non-randomized studies of healthcare interventions. It does not appraise the individual studies inside a review — that is what tools like the Risk of Bias 2 (RoB 2) or Newcastle-Ottawa Scale are for. AMSTAR 2 instead appraises the review itself: whether its methods — the search, study selection, data extraction, risk-of-bias assessment, and synthesis — were rigorous enough to trust its conclusions. It was developed by Beverley Shea and colleagues and published in BMJ in 2017 (358:j4008, PMID 28935701), expanding the original 2007 AMSTAR tool (11 items, limited to reviews of randomized trials) to cover reviews that include non-randomized studies as well.
Last verified 2026-08-24 against the original 2017 BMJ publication and the AMSTAR project’s own documentation.
The AMSTAR 2 Confidence Rating
Each of the 16 items is rated Yes, Partial Yes, or No (five items also allow a “No meta-analysis conducted” response where that step doesn’t apply). Seven of the 16 items — items 2, 4, 7, 9, 11, 13, and 15 — are designated critical domains, because a weakness in any one of them can substantially undermine confidence in the review’s conclusions even if every other item is answered well. The overall rating is driven entirely by how many critical-domain weaknesses exist, not by a simple item count:
| Rating | Criteria |
|---|---|
| High | No critical weaknesses and no more than one non-critical weakness. |
| Moderate | No critical weaknesses, but more than one non-critical weakness. |
| Low | One critical weakness, with or without non-critical weaknesses. |
| Critically Low | More than one critical weakness. |
The 7 Critical Items
The seven critical domains are: item 2 (was the review protocol registered before the review began), item 4 (adequacy of the literature search), item 7 (justification for excluding individual studies), item 9 (adequacy of the technique for assessing risk of bias in individual studies), item 11 (appropriateness of the meta-analytical methods, if a meta-analysis was performed), item 13 (whether risk of bias in individual studies was accounted for when interpreting the review’s results), and item 15 (adequacy of the investigation of publication bias). The remaining nine items cover things like whether the review question and inclusion criteria were specified in advance, whether study selection and data extraction were done in duplicate, and whether a list of excluded studies with reasons was provided; a weakness in one of these alone lowers confidence but does not, by itself, cap it at Low.
AMSTAR vs. AMSTAR 2
The original AMSTAR (2007, 11 items) was built for and validated against systematic reviews of randomized controlled trials. AMSTAR 2 (2017) keeps the same underlying logic — appraising the review’s methods, not its findings — but restructures the items, adds explicit critical domains and the four-tier confidence rating, and extends coverage to systematic reviews that include non-randomized studies of interventions, which the original tool did not adequately handle. AMSTAR 2 is the current, actively used version; AMSTAR 1 is largely of historical interest.
For choosing between AMSTAR 2 and the other major appraisal tools — CASP, Newcastle-Ottawa, and RoB 2 — by study design, see CASRAI’s Choosing a Critical Appraisal Tool guide, which places AMSTAR 2 as the recommended tool specifically for appraising a systematic review of interventions.
Frequently Asked Questions
What does AMSTAR 2 stand for?
A MeaSurement Tool to Assess systematic Reviews, version 2. It appraises the methodological quality of a systematic review itself, not the individual studies inside it.
How many items are in the AMSTAR 2 checklist?
16 items, seven of which are designated critical domains (items 2, 4, 7, 9, 11, 13, and 15) that drive the overall confidence rating.
What is the difference between AMSTAR and AMSTAR 2?
The original 2007 AMSTAR had 11 items and was built for reviews of randomized controlled trials only. AMSTAR 2, published in 2017, has 16 items, adds explicit critical domains and a four-tier confidence rating, and extends coverage to reviews that include non-randomized studies of healthcare interventions.
What are the four AMSTAR 2 confidence ratings?
High, Moderate, Low, and Critically Low. The rating depends on the number of critical-domain weaknesses found, not a simple count of all 16 items: one critical weakness caps the rating at Low, and more than one caps it at Critically Low, regardless of how the non-critical items were rated.
Does AMSTAR 2 appraise individual studies or the review as a whole?
The review as a whole — its search strategy, study selection process, data extraction, risk-of-bias assessment methods, and synthesis. Appraising the individual studies included in the review requires a separate tool matched to each study’s design, such as RoB 2 for randomized trials or Newcastle-Ottawa for observational studies.
Related reading: Choosing a Critical Appraisal Tool for Your Study Design — a side-by-side comparison of CASP, AMSTAR 2, Newcastle-Ottawa, and RoB 2, and which one to use for which study design.
See also: Risk of Bias Assessment: RoB 2, ROBINS-I and the Traffic-Light Plot — the tools for appraising the individual studies an AMSTAR 2-rated review is built from.








