Skip to main content
v2026.11,610 entries · CC-BY 4.0
CASRAIRegulatory RadarCompliance intelligence, specialized for research administrationA daily digest of new regulatory and funding items from four official sources, a subscriber dashboard, and 150 questions a day to Ask CASRAI — grounded in cited sources. $49/month.See Regulatory Radar CASRAI · Own product

Choosing a Critical Appraisal Tool for Your Study Design

A decision table matching study design to the right critical appraisal tool (CASP, AMSTAR 2, Newcastle-Ottawa Scale, RoB 2, ROBINS-I, QUADAS-2), plus a step-by-step appraisal procedure and a direct AMSTAR-vs-CASP-vs-Newcastle-Ottawa comparison.

Ask about Choosing a Critical Appraisal Tool for Your Study Design

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

Last verified: August 16, 2026. Critical appraisal is the systematic process of assessing a study’s methods and results for validity, reliability, and applicability before you rely on its findings. The tool you should use depends almost entirely on the study design in front of you — a checklist built for randomized trials will misjudge a cohort study, and vice versa. This guide gives you a direct decision table for matching study design to tool, a step-by-step appraisal procedure, and a side-by-side comparison of the four instruments researchers ask about most: CASP, AMSTAR 2, the Newcastle-Ottawa Scale, and Cochrane’s RoB 2.

Quick answer: which critical appraisal tool for which study design

Use this table as your starting point. Where two tools are listed, the first is the more formal/Cochrane-aligned option; the second is a lighter-weight alternative common in journal clubs and evidence-based-practice teaching.

Study design Recommended tool What it assesses Output
Systematic review of interventions AMSTAR 2 Methodological rigor of the review itself — search strategy, study selection, risk-of-bias assessment, synthesis methods Confidence rating: High / Moderate / Low / Critically low
Systematic review (general appraisal) CASP Systematic Review Checklist Validity, results, and applicability of the review’s conclusions Narrative judgment across ~10 questions
Randomized controlled trial (inside a systematic review) Cochrane RoB 2 Risk of bias across 5 domains: randomization process, deviations from intended interventions, missing outcome data, measurement of the outcome, selection of the reported result Low risk / Some concerns / High risk, per domain and overall
Randomized controlled trial (standalone appraisal) CASP RCT Checklist Validity, results, and relevance of a single trial Narrative judgment across 11 questions
Non-randomized study of interventions ROBINS-I Risk of bias relative to a hypothetical target trial, across 7 domains including confounding Low / Moderate / Serious / Critical risk, or No information
Cohort study Newcastle-Ottawa Scale (cohort version) or CASP Cohort Study Checklist Selection of groups, comparability, and ascertainment of outcome Star rating (max 9) or narrative judgment
Case-control study Newcastle-Ottawa Scale (case-control version) or CASP Case Control Checklist Selection of cases/controls, comparability, and ascertainment of exposure Star rating (max 9) or narrative judgment
Diagnostic accuracy study QUADAS-2 Risk of bias and applicability across patient selection, index test, reference standard, and flow/timing Low / High / Unclear per domain
Qualitative study CASP Qualitative Studies Checklist or JBI Qualitative Checklist Rigor, credibility, and relevance of qualitative methods and findings Narrative judgment across ~10 questions
Cross-sectional study CASP Cross-Sectional Checklist or JBI Checklist for Analytical Cross Sectional Studies Selection, measurement, and confounding Narrative judgment
Economic evaluation CASP Economic Evaluation Checklist Validity of cost-effectiveness methods and reporting Narrative judgment

If you’re conducting a formal PRISMA-based systematic review or a Cochrane review, your protocol and the Cochrane Handbook effectively mandate RoB 2 for randomized trials and ROBINS-I for non-randomized intervention studies — CASP and the Newcastle-Ottawa Scale are not substitutes in that context, even though they assess similar constructs. For a scoping review, formal quality/risk-of-bias appraisal is often optional under PRISMA-ScR, though JBI’s own critical appraisal tools are commonly applied when a scoping review does choose to appraise sources.

How to critically appraise a paper: a step-by-step procedure

The specific questions differ by tool, but the underlying procedure is the same regardless of which checklist you’re using:

  1. Identify the study design first. Read the methods section, not the abstract’s self-description — authors sometimes mislabel their own design. The table above only works once you know what you’re actually holding.
  2. Select the matching tool from the table above, and pre-specify it in your protocol if this appraisal is part of a systematic review (register it, e.g. via PROSPERO, before you start appraising).
  3. Work through the tool’s questions in order rather than skimming for an overall impression. Every major tool (CASP, AMSTAR 2, RoB 2, ROBINS-I) is built around “signalling questions” that feed into a domain-level judgment — skipping to a gut-feel score defeats the purpose.
  4. Distinguish “reported” from “adequately done.” A paper stating “randomization was performed” is not the same as describing an adequate random-sequence-generation method. Score what was actually described, not what was merely claimed.
  5. Reach an overall judgment from the domain-level judgments using the tool’s own algorithm — RoB 2 and ROBINS-I both specify how domain judgments combine into an overall risk-of-bias rating; AMSTAR 2 specifies which of its 16 items are “critical” domains that can sink the overall confidence rating on their own even if other items pass.
  6. Document your reasoning for each judgment, not just the rating. This is what makes an appraisal auditable and reproducible by a second reviewer or a later reader of your review.
  7. Use two independent reviewers where the appraisal feeds a systematic review or meta-analysis, with a documented process for resolving disagreement — this is standard Cochrane Handbook practice and reduces single-reviewer bias in the appraisal itself.
  8. Feed the judgment into your synthesis, rather than discarding it after use. Common next steps: downgrade certainty in a forest plot‘s pooled estimate (e.g. via GRADE), run a sensitivity analysis excluding high-risk-of-bias studies, or add explicit caveats to a narrative synthesis.

AMSTAR 2 vs. CASP vs. Newcastle-Ottawa vs. RoB 2: how the tools differ

These four are the tools researchers most often confuse, largely because their names get used interchangeably in casual conversation even though they appraise different things.

Tool Appraises Structure Output Typical setting
AMSTAR 2 Systematic reviews of randomized and/or non-randomized intervention studies (the review itself, not its included primary studies) 16 items, 7 designated “critical” domains Confidence rating (High/Moderate/Low/Critically low) — not a numeric score Umbrella reviews, health technology assessment, appraising a review before citing its conclusions
CASP Checklists Eight different individual study/review designs (systematic review, RCT, cohort, case-control, qualitative, diagnostic, economic evaluation, cross-sectional) — one checklist per design Roughly 10-11 questions per checklist, varies by design Narrative judgment, no numeric score Journal clubs, evidence-based-practice teaching, appraising a single primary study
Newcastle-Ottawa Scale (NOS) Cohort and case-control studies only — not RCTs, not qualitative studies 8 items across 3 domains: selection, comparability, outcome/exposure ascertainment Star rating, maximum 9 stars Meta-analyses that pool observational studies
Cochrane RoB 2 Individual randomized controlled trials 5 bias domains plus an overall judgment Low risk / Some concerns / High risk, per domain and overall Mandatory for RCTs in a Cochrane systematic review; widely used outside Cochrane too

The practical distinction that trips people up: AMSTAR 2 appraises a review (a synthesis of many studies), while CASP, the Newcastle-Ottawa Scale, and RoB 2 all appraise individual primary studies — they are not competing options for the same job. Within that second group, CASP is design-flexible (one checklist family covering eight designs) and produces a narrative judgment intended for teaching and quick appraisal, while RoB 2 and the Newcastle-Ottawa Scale are the two most commonly specified in formal systematic-review protocols for their respective designs (RCTs and observational studies), because their structured domain judgments are what feed cleanly into a GRADE certainty rating or a risk-of-bias-stratified forest plot.

Common mistakes when choosing or applying a critical appraisal tool

  • Using AMSTAR 2 on a primary study. AMSTAR 2 appraises systematic reviews, not the individual trials or cohort studies a review includes.
  • Applying the Newcastle-Ottawa Scale to a randomized trial. NOS is built for cohort and case-control studies only; it has no randomization domain because it doesn’t assume one exists.
  • Converting a checklist tally into a single numeric “quality score” and ranking studies by it. Tool developers, including Cochrane and JBI, generally discourage summed numeric scores because they weight every domain equally regardless of how much that domain actually matters for a given study’s conclusions — a single fatal flaw in one critical domain should not be averaged away by strong performance elsewhere.
  • Not pre-specifying the appraisal tool in the review protocol. Choosing or switching tools after seeing which studies you’ll be appraising invites selection bias into the appraisal itself.
  • Single-reviewer appraisal with no disagreement-resolution process on a review intended to inform practice or policy — the Cochrane Handbook’s standard recommendation is independent dual appraisal.

Illustrative example: appraising a mixed body of evidence

Illustrative example — not a real, specific review. A reviewer assembling a systematic review on a clinical intervention ends up with 3 RCTs and 5 cohort studies in the included-studies table. The RCTs get appraised with RoB 2 (one assessment per trial, per outcome if outcomes differ in risk of bias). The cohort studies get appraised with the Newcastle-Ottawa Scale, since this is a formal review intended to inform a GRADE certainty rating and the reviewer wants a structured, comparable domain judgment rather than CASP’s narrative format. Both appraisals are done by two independent reviewers, disagreements are resolved by discussion, and the resulting risk-of-bias judgments are reported in a summary table and used to downgrade GRADE certainty for the outcome most affected by high-risk-of-bias studies.

Frequently asked questions

How do I critically appraise a paper?

Identify the study design, select the tool built for that design (see the table above), work through its questions in order rather than skimming for a gut-feel impression, and reach an overall judgment using the tool’s own method for combining domain-level answers — then document your reasoning, not just the final rating.

What’s the difference between AMSTAR, CASP, and Newcastle-Ottawa?

AMSTAR 2 appraises systematic reviews themselves. CASP is a family of eight checklists, one per study design, producing a narrative judgment. The Newcastle-Ottawa Scale appraises only cohort and case-control studies and produces a star rating. They are not interchangeable options for the same task — match the tool to what you’re actually appraising (a review, or a specific type of primary study).

Can I use more than one critical appraisal tool in the same review?

Yes, and for a review that includes multiple study designs this is normal and expected — for example, RoB 2 for the RCTs and the Newcastle-Ottawa Scale or ROBINS-I for the observational studies in the same body of evidence, as in the illustrative example above.

Do critical appraisal tools produce a pass/fail score?

No — the major tools (AMSTAR 2, CASP, RoB 2, ROBINS-I) deliberately avoid a single numeric pass/fail score. AMSTAR 2 and RoB 2/ROBINS-I output categorical confidence or risk-of-bias ratings; CASP produces a narrative judgment. The Newcastle-Ottawa Scale is the exception, producing a star count, though even that is meant to inform judgment rather than serve as a strict cutoff.

Is critical appraisal the same as peer review?

No. Peer review happens before publication and is performed by the journal on behalf of the author. Critical appraisal happens after publication and is performed by a reader or reviewer deciding how much weight to give a paper’s findings — including papers that already passed peer review. See also how to peer review a systematic review or meta-analysis for the distinct process of reviewing a review for publication.

Related CASRAI resources

Sources. Critical Appraisal Skills Programme (CASP UK), casp-uk.net — checklist list confirmed live 2026-08-16. AMSTAR 2 (Shea et al., BMJ 2017;358:j4008), amstar.ca — item count and rating structure confirmed live 2026-08-16. Newcastle-Ottawa Scale, Ottawa Hospital Research Institute (ohri.ca) — study-design scope and domain structure confirmed live 2026-08-16; the 9-star maximum and 4/2/3 domain split are the tool’s well-established, widely-cited convention (Wells et al.). Cochrane RoB 2 and ROBINS-I domain structures per the current Cochrane Handbook for Systematic Reviews of Interventions. QUADAS-2 domain structure is the tool’s standard, widely-taught four-domain framework.

Background: prisma scr — PRISMA-ScR is the PRISMA extension for scoping reviews: a 20-item reporting checklist built around PCC (Population, Concept, Context) rather than PICO, with its own flow-diagram terminology.

See also: forest plot — A step-by-step guide to reading a forest plot: the study rows, confidence-interval whiskers, weights, the diamond, and the line of no effect, walked through with a worked example.

Related reading: Cluster randomized Trials — What the intracluster correlation coefficient (ICC) measures in a cluster randomised trial, how it produces the design effect that inflates required sample size, why the number of clusters matters more than total partici.

Further reading: Risk of Bias traffic light Plot — A practical guide to the RoB 2 tool for randomized trials and ROBINS-I for non-randomized studies: the domains, judgement categories, and how to build a traffic-light plot with robvis.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →