Skip to main content
v2026.11,610 entries · CC-BY 4.0

ROBINS-I for Non-Randomised Studies: The Seven Domains, Applied to One Study

ROBINS-I’s seven bias domains, walked through practically for a single non-randomized study of an intervention: the target-trial framing the tool assumes, what evidence each domain actually requires, and a worked E-value example for quantifying confounding sensitivity.

Ask about ROBINS-I for Non-Randomised Studies: The Seven Domains, Applied to One Study

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

ROBINS-I (“Risk Of Bias In Non-randomised Studies of Interventions”) is the standard instrument for judging internal validity in a single non-randomized study comparing an intervention to a comparator — a cohort study, a case-control study nested in a defined cohort, a controlled before-after study, or an interrupted time series. This guide walks through applying it to one study: either designing a non-randomized study so it holds up against each domain, or critiquing a study you or someone else has already run. If you are instead assessing risk of bias across many studies for a systematic review or meta-analysis, ROBINS-I’s role there — alongside RoB 2, the traffic-light plot, and how the judgement feeds GRADE — is covered separately in Risk of Bias Assessment: RoB 2, ROBINS-I and the Traffic-Light Plot; this page does not repeat that synthesis-level material.

ROBINS-I was developed by Sterne et al. and published as “ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions” (BMJ 2016;355:i4919). It organizes bias into seven domains and produces a judgement on a five-level scale: Low, Moderate, Serious, or Critical risk of bias, or No information.

Start with the target trial, not the study you have

ROBINS-I’s first, easy-to-skip step is specifying the target trial: the hypothetical, pragmatic randomized trial that would answer your research question if it were ethical and feasible to run. Every domain judgement that follows is made relative to that hypothetical trial, not in the abstract — you are asking “how far does this non-randomized study depart from how the target trial would have handled eligibility, intervention assignment, follow-up, and analysis?” rather than judging the study against a generic standard.

Concretely, before touching the seven domains, write down:

  • Eligibility criteria the target trial would use, and the point in time (“time zero”) at which a participant becomes eligible and intervention status is assigned.
  • The intervention and comparator as the target trial would define them — specific enough that a person’s exposure at time zero could in principle be randomized.
  • The outcome and follow-up period the target trial would measure.
  • The confounders and co-interventions you expect to differ between intervention groups for reasons other than the intervention itself, and that also affect the outcome — listed before you look at Domain 1’s evidence, not after, so the confounding assessment isn’t reverse-engineered from what the data happen to show.

This step is the single biggest source of inconsistency when teams apply ROBINS-I without training: skipping straight to the domains without a written target trial and confounder list means two assessors reviewing the same study can reasonably disagree, because they’re implicitly comparing it against different hypothetical trials. For background on the broader target-trial framework this borrows from, see Target Trial Emulation: Applying RCT Design Principles to Real-World Data.

The seven domains, and what evidence each one actually requires

Domain 1: Bias due to confounding

This is the domain with no RoB 2 equivalent — randomization is designed to prevent exactly this, so non-randomized studies need it and RCTs don’t. It asks whether baseline differences between intervention groups, caused by something other than the intervention itself, could plausibly explain the observed association with the outcome.

To make this judgement you need to see, not infer: a baseline characteristics table comparing intervention and comparator groups on every pre-specified confounder; the analytic strategy used to handle them (restriction, matching, stratification, multivariable regression adjustment, propensity-score methods, or an instrumental-variable approach); and, if a confounder can change over the course of follow-up in response to the intervention (time-varying confounding), evidence the analysis accounted for that rather than adjusting only at baseline. A study that lists confounders but adjusts for them with a crude bivariate comparison, or that never mentions an important confounder specific to the clinical/policy area at all, cannot be rated better than Serious on this domain regardless of how well the other six domains score. See Confounding Variable, Propensity Score Matching: How It Works and What It Cannot Fix, and Inverse Probability Weighting: When It Beats Propensity Score Matching for the analytic methods this domain is judging.

Domain 2: Bias due to selection of participants into the study

This asks whether the way participants ended up in the analysis — including how “time zero” was defined — depended on both the intervention and the outcome. A common failure pattern: participants are only included if they survived long enough after starting the intervention to be observed taking it, which selects out early events in a way that flatters the intervention group. Evidence to look for: the exact eligibility and exclusion criteria, whether time zero for intervention and comparator groups is defined identically and aligned with when intervention status was actually determined, and whether any post-baseline event (e.g., loss to follow-up, a change in eligibility status) was used to decide inclusion.

Domain 3: Bias in classification of interventions

Was intervention status classified using information collected at or before the start of the intervention, and was the same classification procedure applied to every group? A study that determines exposure retrospectively from records that could be influenced by knowledge of the outcome (e.g., chart review conducted after an adverse event, where documentation is more thorough for cases) risks differential misclassification. Evidence needed: the data source and exact timing of intervention-status determination, and confirmation the same procedure applied symmetrically across comparison groups.

Domain 4: Bias due to deviations from intended interventions

The non-randomized analogue of RoB 2’s deviations domain: did participants receive an intervention that deviated from what was assigned or started, was this imbalanced between groups, and did the analysis handle it appropriately for the question being asked (an intention-to-treat-type effect vs. a per-protocol effect are different estimands and need different analytic strategies)? Evidence: description of switching, discontinuation, or co-intervention use after baseline, and whether the analytic approach matches the stated target-trial estimand rather than silently mixing the two.

Domain 5: Bias due to missing data

Were outcome, exposure, and confounder data reasonably complete, and is it plausible the analysis’s handling of missingness biased the result? Evidence: the proportion missing on each key variable, by group; whether missingness is plausibly related to the true (unobserved) value of that variable; and whether the study did anything beyond an unexamined complete-case analysis when missingness was non-trivial (multiple imputation, inverse-probability-of-missingness weighting, or at minimum a sensitivity comparison against complete cases).

Domain 6: Bias in measurement of outcomes

Was the outcome measured the same way, from the same source, at the same point in follow-up, across comparison groups — and could knowledge of which intervention a participant received have influenced how the outcome was measured or recorded? This matters most for subjective or judgement-dependent outcomes; an objective outcome from a source blind to intervention status (e.g., linked mortality data) is lower risk on this domain by construction.

Domain 7: Bias in selection of the reported result

Was the reported result selected, from among multiple outcome definitions, multiple analytic models, or multiple subgroups, on the basis of which one produced the most favorable or most publishable result? Evidence: a pre-registered protocol or statistical analysis plan predating the data, if one exists, checked against what was actually reported; and whether the paper reports every pre-specified outcome and subgroup or only a favorable subset.

How the overall judgement works

Each domain is rated Low, Moderate, Serious, or Critical risk of bias, or No information. The overall judgement for the study (or, in a review context, for a specific result within it) is driven by the single worst domain, not an average across domains — a study that scores well on six domains but is rated Critical on confounding is rated Critical overall, because a Critical rating on any one domain (bias severe enough to invalidate the comparison) is treated as invalidating the result, not merely discounting it. Resist the temptation to informally tally “5 green, 2 yellow” and call the study moderate; that is not how the algorithm works.

A worked example: quantifying how much unmeasured confounding could explain

Domain 1’s judgement is often qualitative — a narrative comparison of the confounders adjusted for against the confounders you expect matter. One way to make it more concrete when critiquing (or reporting) a single study is an E-value sensitivity analysis (VanderWeele TJ, Ding P, “Sensitivity Analysis in Observational Research: Introducing the E-Value,” Annals of Internal Medicine, 2017). The E-value is the minimum strength of association, on the risk-ratio scale, that an unmeasured confounder would need to have with both the intervention and the outcome, above and beyond the measured confounders already adjusted for, to fully explain away the observed association.

For a risk ratio RR ≥ 1, the E-value is RR + √(RR × (RR − 1)); for RR < 1, invert it first (1/RR) and apply the same formula. Take a hypothetical adjusted risk ratio of 1.80 for an intervention-outcome association, with a 95% confidence interval lower bound of 1.30 (both figures illustrative, not drawn from a real study):

  • E-value for the point estimate (RR = 1.80): 1.80 + √(1.80 × 0.80) = 3.00
  • E-value for the confidence-interval bound closer to the null (RR = 1.30): 1.30 + √(1.30 × 0.30) = 1.92

(Both values computed directly from the formula above, not estimated by eye.) Read this as: an unmeasured confounder would need to be associated with both the intervention and the outcome by a risk ratio of at least 1.92 — above and beyond every confounder already adjusted for — to move the confidence interval to include the null. Whether that’s plausible is a judgement call specific to the clinical or policy context, but it turns Domain 1’s assessment from “did they adjust for confounders?” into “how strong would an unmeasured one have to be, and is that realistic given what’s known about this exposure-outcome relationship?” — a genuinely sharper question than the qualitative version, and one worth running whenever a study reports enough information (point estimate and confidence interval) to compute it.

Designing a study against ROBINS-I vs. critiquing one after the fact

Applying the seven domains at the design stage, before data collection, is more useful than applying them retrospectively, because most of ROBINS-I’s Critical-risk failure patterns are only fixable in the design: a confounder that wasn’t measured can’t be adjusted for after the fact; a time-zero misalignment that created immortal time can’t be un-created without re-defining the cohort; an outcome-ascertainment method that wasn’t blinded to intervention status can’t be reblinded retrospectively. Using ROBINS-I as a design checklist means, for each domain, asking “what would I need to have collected, or how would I need to have defined this, for an assessor applying ROBINS-I to my finished study to rate this domain Low?” and building the study around that answer, rather than discovering the gap in a reviewer’s risk-of-bias table after the study is already run.

Critiquing an already-completed study — your own, a collaborator’s, or one you’re reading as a reader rather than a formal reviewer — still uses the same seven domains and the same target-trial framing, but the output is a judgement rather than a design decision: how much weight the result deserves given what the study could and couldn’t control for, which is a more calibrated way to read a single non-randomized study’s headline effect estimate than either accepting it uncritically or dismissing all non-randomized evidence outright.

Common mistakes

  • Skipping the target-trial and confounder pre-specification step. Jumping straight to Domain 1 without first writing down which confounders matter, independent of what the data show, is the most common source of inconsistent ROBINS-I judgements between assessors.
  • Treating the domain ratings as an average. The overall judgement is set by the worst domain, not a mean across the seven.
  • Assessing the study instead of the specific comparison. A study with multiple exposure comparisons or multiple outcomes can warrant a separate ROBINS-I judgement per comparison, since confounding and measurement concerns can differ by outcome even within the same dataset.
  • Confusing ROBINS-I with ROBINS-E. ROBINS-I assesses studies of an intervention — something assigned, chosen, or initiated as a treatment or policy decision. Studies of an exposure that isn’t an intervention decision (environmental or occupational exposure, for instance) use the related but distinct ROBINS-E tool instead.

Frequently asked questions

Can ROBINS-I be used on a study I’m designing, before any data exist?

Yes — that’s arguably where it adds the most value, since several of its Critical-risk failure patterns (an unmeasured confounder, a misaligned time zero, an unblinded outcome assessment) are only fixable in the design stage. Working through the seven domains as a pre-data checklist, against your specified target trial, is a legitimate and increasingly recommended use of the tool.

Does ROBINS-I apply to case-control studies?

Yes, when the case-control study is nested within a defined cohort with a clear time zero and eligibility criteria. A poorly defined case-control study where cases and controls aren’t drawn from a common, definable source population is harder to map onto ROBINS-I’s domains cleanly — see Case-Control Study vs. Cohort Study and Cohort Study: Design, Types, and How It Works for how the two designs differ on this point.

Is a Critical rating on one domain always fatal to the study’s conclusions?

It means the result is judged too much at risk of bias to be interpreted as reliable evidence for the specific comparison assessed — not that the study has no value at all. A Critical-rated comparison can still be useful for hypothesis generation or for informing which confounders a better-designed follow-up study should prioritize measuring.

How is this different from RoB 2?

RoB 2 assesses randomized trials, where randomization is assumed to have handled baseline confounding, so RoB 2 has five domains and no confounding domain. ROBINS-I assesses non-randomized studies, where confounding is a live threat that has to be assessed directly, hence the two additional domains (confounding, and selection into the study) ahead of the five that parallel RoB 2’s structure. See RCT vs. Observational Study for the underlying design distinction.

Related guides

Last verified 2026-08-29 against the ROBINS-I primary publication (Sterne et al., BMJ 2016;355:i4919) and the E-value primary publication (VanderWeele & Ding, Annals of Internal Medicine, 2017); the numeric E-value example above was computed directly from the published formula, not estimated. Consult the Cochrane Methods risk-of-bias resource pages directly before a formal assessment, since implementation guidance for ROBINS-I is updated periodically.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.