Written and maintained by CASRAI Editorial Board
Last updated
Last verified: August 16, 2026, against the Ottawa Hospital Research Institute (OHRI) Newcastle-Ottawa Scale page (ohri.ca) and the Cochrane Handbook for Systematic Reviews of Interventions. This is a methodology-orientation guide, not a substitute for the official NOS manuals and coding forms, which should be downloaded directly from OHRI before use in a real review.
The Newcastle-Ottawa Scale at a glance
| Domain | Cohort studies | Case-control studies | Max. stars |
|---|---|---|---|
| Selection | Representativeness of the exposed cohort; selection of the non-exposed cohort; ascertainment of exposure; demonstration that the outcome was not present at the start of the study | Adequacy of the case definition; representativeness of the cases; selection of controls; definition of controls | 4 |
| Comparability | Comparability of cohorts on the basis of the design or analysis (control for the most important confounder, plus any additional confounder) | Comparability of cases and controls on the basis of the design or analysis (same two-item structure) | 2 |
| Outcome / Exposure | Assessment of outcome; was follow-up long enough for outcomes to occur; adequacy of follow-up of cohorts | Ascertainment of exposure; same method of ascertainment for cases and controls; non-response rate | 3 |
| Total possible | 9 stars | 9 | |
Each item can receive at most one star (the two Comparability items can together contribute up to two stars), awarded only when a study meets the specific criterion described on the relevant NOS coding form. A study earns a star for meeting a criterion — there is no partial credit within an item.
What the Newcastle-Ottawa Scale is
The Newcastle-Ottawa Scale (NOS) is a quality-assessment tool used to appraise non-randomized studies — specifically cohort studies and case-control studies — for inclusion in systematic reviews and meta-analyses. It was developed by G.A. Wells, B. Shea, D. O’Connell, J. Peterson, V. Welch, M. Losos and P. Tugwell, through a collaboration between the University of Newcastle (Australia) and the University of Ottawa (Canada) — the source of the tool’s name. It is now maintained and published by the Ottawa Hospital Research Institute (OHRI).
Randomized controlled trials have well-established risk-of-bias tools (notably Cochrane’s RoB 2). Observational studies needed an equivalent that could be applied consistently across reviewers, and the NOS was built to fill that gap using a star-rating system rather than a simple checklist, based on three broad perspectives common to case-control and cohort designs: how the study groups were selected, how comparable those groups were, and how exposure or outcome was ascertained.
The scale exists in two parallel versions with the same three-domain structure but different items within Selection and Outcome/Exposure, reflecting the different logic of the two designs — see the Case-Control Study vs. Cohort Study comparison for how the underlying designs themselves differ before scoring either one.
How scoring works, domain by domain
Selection (up to 4 stars)
For a cohort study, Selection asks whether the exposed cohort is truly representative of the average person with that exposure in the community (rather than, say, a selected group of volunteers), whether the non-exposed cohort was drawn from the same community as the exposed cohort, whether exposure was ascertained from a secure record or structured interview rather than self-report alone, and whether the study demonstrated that the outcome of interest was not already present when the cohort was assembled.
For a case-control study, Selection asks whether cases were defined with independent validation (not just self-report or database code alone), whether cases are consecutive or clearly representative rather than a selected series, whether controls were drawn from the community rather than, for example, hospital controls with a different exposure profile, and whether controls had no history of the outcome under study.
Comparability (up to 2 stars)
This is the domain most reviewers get wrong in practice: it is not a general judgment of “did the study control for confounders,” it is scored against the specific confounder(s) the review team has pre-specified as most important for that outcome (one star), plus a second star if the study controlled for any additional relevant factor. Because the reviewer decides in advance which factor counts as “the most important,” Comparability should be defined in the review protocol before appraisal begins, not chosen after seeing which studies adjusted for what.
Outcome (cohort) / Exposure (case-control) (up to 3 stars)
For cohort studies this domain scores whether outcome assessment was independent or blind (e.g., record linkage) rather than self-report, whether follow-up was long enough for the outcome to plausibly occur, and whether follow-up of the cohort was adequate (a common threshold cited in NOS guidance and secondary literature is loss to follow-up below roughly 20%, or a described and defensible reason for higher attrition). For case-control studies it scores whether exposure was ascertained through a secure record or blinded interview rather than an unblinded interview or written self-report alone, whether cases and controls were assessed for exposure using the identical method, and whether non-response rates were reported and comparable between the two groups.
Worked example (illustrative, not a real published study)
The walkthrough below is a synthesized example built to show how the star system is applied in practice. It is not drawn from, or attributed to, any specific real study.
Suppose a reviewer is appraising a hypothetical retrospective cohort study comparing a nutritional exposure against a cardiovascular outcome:
- Selection: the exposed cohort is drawn from a population-representative registry (★), the non-exposed cohort is drawn from the same registry (★), exposure is ascertained from validated dietary-intake records rather than a single self-report (★), but the study does not clearly demonstrate the outcome was absent at baseline (no star). Selection subtotal: 3/4.
- Comparability: the analysis adjusts for age and sex, the review team’s pre-specified most-important confounder (★), and separately adjusts for smoking status as an additional factor (★). Comparability subtotal: 2/2.
- Outcome: outcome is ascertained via linked hospital-admission records rather than self-report (★), follow-up duration (8 years) is long enough for the outcome to occur (★), but loss to follow-up is 28% with no stated reason (no star). Outcome subtotal: 2/3.
Total: 7/9 stars — Selection 3, Comparability 2, Outcome 2.
Interpreting the total: what counts as a “good” NOS score
The NOS itself does not publish an official good/fair/poor cutoff on the OHRI page; the star total is a summary of a domain-by-domain appraisal, not an inherent grade. In practice, most reviewers who need a threshold use the conversion popularized by the US Agency for Healthcare Research and Quality (AHRQ) for its evidence reports, commonly reported as:
| Category | Typical threshold (widely used AHRQ-derived convention) |
|---|---|
| Good quality | 3–4 stars in Selection and 1–2 stars in Comparability and 2–3 stars in Outcome/Exposure |
| Fair quality | 2 stars in Selection and 1–2 stars in Comparability and 2–3 stars in Outcome/Exposure |
| Poor quality | 0–1 stars in Selection, or 0 stars in Comparability, or 0–1 stars in Outcome/Exposure |
Treat this table as a widely used convention rather than a fixed part of the official NOS documentation — different reviews, and different review teams, sometimes apply their own thresholds or report the domain scores without collapsing them into a category at all, so state whichever convention you use explicitly in your methods section rather than assuming readers will recognize an unstated cutoff.
A caution the Cochrane Handbook raises about star-based tools
The Cochrane Handbook for Systematic Reviews of Interventions is explicit that risk-of-bias assessment should generally be reported and interpreted at the level of individual domains rather than collapsed into a single composite score, because a summary number can mask which specific source of bias is driving a low rating and can be misused as a hard inclusion/exclusion cutoff. This is a live methodological tension for the NOS specifically: it is scored as a star total, which is convenient for reporting and for weighting studies in sensitivity analyses, but the Handbook’s general preference for structured, domain-based risk-of-bias judgment is one reason Cochrane reviews of non-randomized studies of interventions now typically use ROBINS-I rather than the NOS. The NOS nonetheless remains extremely widely used outside formal Cochrane reviews — particularly in general biomedical and public-health meta-analyses of cohort and case-control evidence — in large part because of its speed, its long track record, and the ease of training multiple reviewers to apply it consistently.
Newcastle-Ottawa Scale vs. ROBINS-I vs. GRADE: what each one actually does
| Tool | What it appraises | Output | Where it fits |
|---|---|---|---|
| Newcastle-Ottawa Scale (NOS) | Individual cohort or case-control studies | Star count per domain (max 9) | Study-level quality appraisal, most common outside formal Cochrane reviews |
| ROBINS-I | Individual non-randomized studies of interventions | Domain-level risk-of-bias judgment (Low / Moderate / Serious / Critical / No information) | Study-level appraisal, Cochrane’s preferred tool for NRSI |
| GRADE | The body of evidence for a specific outcome across all included studies | Certainty rating (High / Moderate / Low / Very low) | Evidence-level synthesis, applied after individual studies are already appraised |
These tools are not interchangeable substitutes: NOS and ROBINS-I both appraise individual studies (and a review generally uses one or the other, not both, for the same study set), while GRADE operates one level up, synthesizing the appraised evidence base for a given outcome into a certainty rating. See Heterogeneity in Meta-Analysis for how study-level appraisal connects to pooling decisions, and Systematic Review vs. Meta-Analysis for how quality appraisal fits into the broader synthesis workflow.
Frequently asked questions
Is there a Newcastle-Ottawa Scale for cross-sectional studies?
The official NOS from OHRI covers only cohort and case-control designs. A number of adapted versions for cross-sectional studies circulate in the published literature, but none is the official OHRI-endorsed instrument — if you use one, cite the specific adaptation and its source explicitly, since these unofficial versions differ from each other in item wording and star allocation.
Do reviewers need to double-score with the NOS?
Independent double appraisal (two reviewers scoring the same study separately, then resolving disagreements, often with a documented inter-rater agreement statistic) is standard systematic-review practice generally and is reported in the methods sections of most reviews that use the NOS, though it is a general good-practice convention rather than a rule written into the NOS instrument itself.
Can NOS star totals be used as a hard inclusion cutoff?
It is common in practice but is exactly the use the Cochrane Handbook’s general guidance on risk-of-bias tools cautions against: a single composite threshold can exclude a study for a reason unrelated to the outcome in question, or include a study that is well-appraised on two domains but seriously flawed on the third. Reporting domain-level stars alongside any threshold used, and running a sensitivity analysis with and without lower-scoring studies, is the more defensible approach.
Where can I get the official NOS coding forms?
The Ottawa Hospital Research Institute publishes the current coding manuals and forms directly on its site (ohri.ca), with separate downloadable forms for cohort studies and case-control studies. Always score from the current official form rather than a version reproduced in a secondary source, since item wording is what reviewers are actually trained against.
Related CASRAI pages
- Cohort Study: Design, Types, and How It Works
- Case-Control Study: Design, Odds Ratios, and Common Pitfalls
- Case-Control Study vs. Cohort Study
- Systematic Review vs. Meta-Analysis
- PRISMA and Systematic Review Methodology
- Heterogeneity in Meta-Analysis: I², τ², and Prediction Intervals
- Publication Bias
- Cochrane Handbook for Systematic Reviews of Interventions








