Skip to main content
v2026.11,610 entries · CC-BY 4.0
CASRAIRegulatory RadarCompliance intelligence, specialized for research administrationA daily digest of new regulatory and funding items from four official sources, a subscriber dashboard, and 150 questions a day to Ask CASRAI — grounded in cited sources. $49/month.See Regulatory Radar CASRAI · Own product

The Newcastle-Ottawa Scale (NOS): How to Score Study Quality

How the Newcastle-Ottawa Scale (NOS) star system works for cohort and case-control studies: the three domains, a worked scoring example, good/fair/poor interpretation, and how NOS compares to ROBINS-I and GRADE.

Ask about The Newcastle-Ottawa Scale (NOS): How to Score Study Quality

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

Last verified: August 16, 2026, against the Ottawa Hospital Research Institute (OHRI) Newcastle-Ottawa Scale page (ohri.ca) and the Cochrane Handbook for Systematic Reviews of Interventions. This is a methodology-orientation guide, not a substitute for the official NOS manuals and coding forms, which should be downloaded directly from OHRI before use in a real review.

The Newcastle-Ottawa Scale at a glance

Domain Cohort studies Case-control studies Max. stars
Selection Representativeness of the exposed cohort; selection of the non-exposed cohort; ascertainment of exposure; demonstration that the outcome was not present at the start of the study Adequacy of the case definition; representativeness of the cases; selection of controls; definition of controls 4
Comparability Comparability of cohorts on the basis of the design or analysis (control for the most important confounder, plus any additional confounder) Comparability of cases and controls on the basis of the design or analysis (same two-item structure) 2
Outcome / Exposure Assessment of outcome; was follow-up long enough for outcomes to occur; adequacy of follow-up of cohorts Ascertainment of exposure; same method of ascertainment for cases and controls; non-response rate 3
Total possible 9 stars 9

Each item can receive at most one star (the two Comparability items can together contribute up to two stars), awarded only when a study meets the specific criterion described on the relevant NOS coding form. A study earns a star for meeting a criterion — there is no partial credit within an item.

What the Newcastle-Ottawa Scale is

The Newcastle-Ottawa Scale (NOS) is a quality-assessment tool used to appraise non-randomized studies — specifically cohort studies and case-control studies — for inclusion in systematic reviews and meta-analyses. It was developed by G.A. Wells, B. Shea, D. O’Connell, J. Peterson, V. Welch, M. Losos and P. Tugwell, through a collaboration between the University of Newcastle (Australia) and the University of Ottawa (Canada) — the source of the tool’s name. It is now maintained and published by the Ottawa Hospital Research Institute (OHRI).

Randomized controlled trials have well-established risk-of-bias tools (notably Cochrane’s RoB 2). Observational studies needed an equivalent that could be applied consistently across reviewers, and the NOS was built to fill that gap using a star-rating system rather than a simple checklist, based on three broad perspectives common to case-control and cohort designs: how the study groups were selected, how comparable those groups were, and how exposure or outcome was ascertained.

The scale exists in two parallel versions with the same three-domain structure but different items within Selection and Outcome/Exposure, reflecting the different logic of the two designs — see the Case-Control Study vs. Cohort Study comparison for how the underlying designs themselves differ before scoring either one.

How scoring works, domain by domain

Selection (up to 4 stars)

For a cohort study, Selection asks whether the exposed cohort is truly representative of the average person with that exposure in the community (rather than, say, a selected group of volunteers), whether the non-exposed cohort was drawn from the same community as the exposed cohort, whether exposure was ascertained from a secure record or structured interview rather than self-report alone, and whether the study demonstrated that the outcome of interest was not already present when the cohort was assembled.

For a case-control study, Selection asks whether cases were defined with independent validation (not just self-report or database code alone), whether cases are consecutive or clearly representative rather than a selected series, whether controls were drawn from the community rather than, for example, hospital controls with a different exposure profile, and whether controls had no history of the outcome under study.

Comparability (up to 2 stars)

This is the domain most reviewers get wrong in practice: it is not a general judgment of “did the study control for confounders,” it is scored against the specific confounder(s) the review team has pre-specified as most important for that outcome (one star), plus a second star if the study controlled for any additional relevant factor. Because the reviewer decides in advance which factor counts as “the most important,” Comparability should be defined in the review protocol before appraisal begins, not chosen after seeing which studies adjusted for what.

Outcome (cohort) / Exposure (case-control) (up to 3 stars)

For cohort studies this domain scores whether outcome assessment was independent or blind (e.g., record linkage) rather than self-report, whether follow-up was long enough for the outcome to plausibly occur, and whether follow-up of the cohort was adequate (a common threshold cited in NOS guidance and secondary literature is loss to follow-up below roughly 20%, or a described and defensible reason for higher attrition). For case-control studies it scores whether exposure was ascertained through a secure record or blinded interview rather than an unblinded interview or written self-report alone, whether cases and controls were assessed for exposure using the identical method, and whether non-response rates were reported and comparable between the two groups.

Worked example (illustrative, not a real published study)

The walkthrough below is a synthesized example built to show how the star system is applied in practice. It is not drawn from, or attributed to, any specific real study.

Suppose a reviewer is appraising a hypothetical retrospective cohort study comparing a nutritional exposure against a cardiovascular outcome:

  • Selection: the exposed cohort is drawn from a population-representative registry (★), the non-exposed cohort is drawn from the same registry (★), exposure is ascertained from validated dietary-intake records rather than a single self-report (★), but the study does not clearly demonstrate the outcome was absent at baseline (no star). Selection subtotal: 3/4.
  • Comparability: the analysis adjusts for age and sex, the review team’s pre-specified most-important confounder (★), and separately adjusts for smoking status as an additional factor (★). Comparability subtotal: 2/2.
  • Outcome: outcome is ascertained via linked hospital-admission records rather than self-report (★), follow-up duration (8 years) is long enough for the outcome to occur (★), but loss to follow-up is 28% with no stated reason (no star). Outcome subtotal: 2/3.

Total: 7/9 stars — Selection 3, Comparability 2, Outcome 2.

Interpreting the total: what counts as a “good” NOS score

The NOS itself does not publish an official good/fair/poor cutoff on the OHRI page; the star total is a summary of a domain-by-domain appraisal, not an inherent grade. In practice, most reviewers who need a threshold use the conversion popularized by the US Agency for Healthcare Research and Quality (AHRQ) for its evidence reports, commonly reported as:

Category Typical threshold (widely used AHRQ-derived convention)
Good quality 3–4 stars in Selection and 1–2 stars in Comparability and 2–3 stars in Outcome/Exposure
Fair quality 2 stars in Selection and 1–2 stars in Comparability and 2–3 stars in Outcome/Exposure
Poor quality 0–1 stars in Selection, or 0 stars in Comparability, or 0–1 stars in Outcome/Exposure

Treat this table as a widely used convention rather than a fixed part of the official NOS documentation — different reviews, and different review teams, sometimes apply their own thresholds or report the domain scores without collapsing them into a category at all, so state whichever convention you use explicitly in your methods section rather than assuming readers will recognize an unstated cutoff.

A caution the Cochrane Handbook raises about star-based tools

The Cochrane Handbook for Systematic Reviews of Interventions is explicit that risk-of-bias assessment should generally be reported and interpreted at the level of individual domains rather than collapsed into a single composite score, because a summary number can mask which specific source of bias is driving a low rating and can be misused as a hard inclusion/exclusion cutoff. This is a live methodological tension for the NOS specifically: it is scored as a star total, which is convenient for reporting and for weighting studies in sensitivity analyses, but the Handbook’s general preference for structured, domain-based risk-of-bias judgment is one reason Cochrane reviews of non-randomized studies of interventions now typically use ROBINS-I rather than the NOS. The NOS nonetheless remains extremely widely used outside formal Cochrane reviews — particularly in general biomedical and public-health meta-analyses of cohort and case-control evidence — in large part because of its speed, its long track record, and the ease of training multiple reviewers to apply it consistently.

Newcastle-Ottawa Scale vs. ROBINS-I vs. GRADE: what each one actually does

Tool What it appraises Output Where it fits
Newcastle-Ottawa Scale (NOS) Individual cohort or case-control studies Star count per domain (max 9) Study-level quality appraisal, most common outside formal Cochrane reviews
ROBINS-I Individual non-randomized studies of interventions Domain-level risk-of-bias judgment (Low / Moderate / Serious / Critical / No information) Study-level appraisal, Cochrane’s preferred tool for NRSI
GRADE The body of evidence for a specific outcome across all included studies Certainty rating (High / Moderate / Low / Very low) Evidence-level synthesis, applied after individual studies are already appraised

These tools are not interchangeable substitutes: NOS and ROBINS-I both appraise individual studies (and a review generally uses one or the other, not both, for the same study set), while GRADE operates one level up, synthesizing the appraised evidence base for a given outcome into a certainty rating. See Heterogeneity in Meta-Analysis for how study-level appraisal connects to pooling decisions, and Systematic Review vs. Meta-Analysis for how quality appraisal fits into the broader synthesis workflow.

Frequently asked questions

Is there a Newcastle-Ottawa Scale for cross-sectional studies?

The official NOS from OHRI covers only cohort and case-control designs. A number of adapted versions for cross-sectional studies circulate in the published literature, but none is the official OHRI-endorsed instrument — if you use one, cite the specific adaptation and its source explicitly, since these unofficial versions differ from each other in item wording and star allocation.

Do reviewers need to double-score with the NOS?

Independent double appraisal (two reviewers scoring the same study separately, then resolving disagreements, often with a documented inter-rater agreement statistic) is standard systematic-review practice generally and is reported in the methods sections of most reviews that use the NOS, though it is a general good-practice convention rather than a rule written into the NOS instrument itself.

Can NOS star totals be used as a hard inclusion cutoff?

It is common in practice but is exactly the use the Cochrane Handbook’s general guidance on risk-of-bias tools cautions against: a single composite threshold can exclude a study for a reason unrelated to the outcome in question, or include a study that is well-appraised on two domains but seriously flawed on the third. Reporting domain-level stars alongside any threshold used, and running a sensitivity analysis with and without lower-scoring studies, is the more defensible approach.

Where can I get the official NOS coding forms?

The Ottawa Hospital Research Institute publishes the current coding manuals and forms directly on its site (ohri.ca), with separate downloadable forms for cohort studies and case-control studies. Always score from the current official form rather than a version reproduced in a secondary source, since item wording is what reviewers are actually trained against.

Related CASRAI pages

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →