Written and maintained by CASRAI Editorial Board
Last updated
Grading of Recommendations Assessment, Development and Evaluation (GRADE) rates the certainty of a body of evidence across five domains: risk of bias, inconsistency, indirectness, imprecision, and publication bias. Of the five, imprecision is the one most often misread from a forest plot alone — a reviewer sees a narrow confidence interval that excludes the null and assumes the evidence is precise. GRADE’s own guidance says that assumption is not safe on its own. The imprecision domain asks a second question the confidence interval cannot answer by itself: does the meta-analysis actually contain enough participants to have been an adequately powered trial in the first place? That second question is answered by comparing the evidence base against its optimal information size (OIS).
What Optimal Information Size Represents
Optimal information size is the number of participants a single, adequately powered randomised trial would need to reliably detect the effect of interest. It is calculated the same way a conventional trial sample-size calculation is: from an assumed control-group event rate or baseline value, the minimally important or expected treatment effect, and conventional error rates (typically α = 0.05, power ≥ 80–90%). Whichever formula a reviewer uses — a proportions-based calculation for dichotomous outcomes or a means-based calculation for continuous outcomes — the logic is identical to any other power analysis and sample-size calculation: you specify the effect you need to be able to detect and the calculation returns the number of participants required to detect it with adequate certainty.
The difference is where GRADE applies that number. Instead of sizing a single planned trial before it starts, OIS is calculated retrospectively and applied to an already-pooled body of evidence, then compared against how many participants the meta-analysis actually accumulated across all included studies. It is, in effect, a power calculation run against the review itself rather than against any one trial inside it.
How GRADE Compares Accumulated Sample Size Against OIS
GRADE’s guidance on rating imprecision (the Guyatt et al. GRADE series paper, “GRADE guidelines 6: rating the quality of evidence—imprecision,” Journal of Clinical Epidemiology, 2011) sets out two separate criteria a reviewer checks, either of which can trigger a downgrade on its own:
- The optimal information size criterion. If the total number of participants across all studies contributing to the pooled estimate is smaller than the number that a conventionally calculated single-trial sample size would require to detect the effect of interest, consider rating down for imprecision — independent of what the confidence interval looks like.
- The confidence interval criterion. If the interval around the pooled effect is wide enough to span both “no important effect” and an effect large enough to change the recommendation — i.e., it crosses a clinically important threshold rather than just the null — that also supports downgrading, even when the point estimate itself is precise-looking.
These two checks are complementary, not redundant. A meta-analysis can pass one and fail the other. Reviewers building a GRADE Summary of Findings table apply the imprecision domain alongside the other four, and the OIS check is specifically the one that catches sparse-data pooled estimates a confidence interval alone would let through.
Why a Narrow, Statistically Significant CI Is Not Automatically Reassuring
This is the counter-intuitive part of the domain, and the reason GRADE built a second criterion instead of relying on interval width alone. A pooled estimate from two or three small trials can produce a narrow interval that excludes the null purely by chance — random variation in a small sample can understate the true uncertainty of the effect, especially with few events. Reading that narrow interval as “precise” would be a mistake: the evidence base still has not accumulated the number of participants a single adequately powered trial would have needed, so the estimate has not been tested against the sample size that would normally be required to trust it.
Understanding this requires reading the interval correctly in the first place — see how to calculate and interpret a confidence interval and how to read a forest plot for the mechanics GRADE assumes a reviewer already has. The practical upshot: a statistically significant result with a tight-looking interval can still be downgraded one level for imprecision if the accumulated sample size falls short of OIS. GRADE panels see this most often in early or thinly evidenced comparisons, where only a handful of small trials have been pooled and the apparent precision of the summary estimate outruns the actual amount of data behind it.
Estimating OIS for a Pooled Estimate in Practice
In practice, reviewers estimate OIS the same way they would size a hypothetical single trial testing the comparison of interest, then sum the number of participants (or, for dichotomous outcomes, sometimes the number of events) actually contributed by every study in the pooled analysis and compare the two figures. Some conventions in the literature use a fixed minimum threshold (several hundred events for a typical binary outcome) as a rough proxy when a full calculation is impractical, but the guidance itself frames OIS as case-specific: it depends on the assumed control-group risk and the effect size judged clinically important for that particular question, not a single number that applies across all outcomes. This is also why the same body of evidence can be judged imprecise for one outcome (a rare adverse event with few accumulated cases) while being adequately powered for another (a common outcome with a large accumulated sample) within the same review.
The choice between a fixed-effect and random-effects pooling model and the weighting method used to pool the estimate both affect the width of the resulting interval, which is part of why GRADE keeps the OIS check separate from the interval-width check — a model choice that narrows the interval does not change how many participants were actually studied.
OIS and Related Sequential-Monitoring Methods
The same underlying logic — that a pooled result should be measured against the sample size an adequately powered single trial would need — also underlies Trial Sequential Analysis (TSA), which applies OIS-style required-information-size calculations with formal sequential monitoring boundaries to control the risk of a false-positive conclusion from repeatedly updating a meta-analysis as new trials accumulate. TSA is a distinct, more formal statistical method with its own alpha-spending boundaries; GRADE’s imprecision domain is a qualitative certainty-rating judgment that borrows the same conceptual comparison rather than adopting TSA’s full monitoring-boundary apparatus. Reviewers should not treat the two as interchangeable, but understanding one clarifies the other.
Practical Implications for Reviewers and Guideline Panels
For anyone building or appraising a Summary of Findings table or working through a GRADE Evidence-to-Decision framework, three implications follow directly from how OIS works:
- Don’t stop at the confidence interval. A significant, narrow-looking pooled estimate from a small number of trials still needs its accumulated sample size checked against OIS before imprecision is rated “no serious concern.”
- Rare outcomes are especially exposed. Adverse-event and other low-event-rate outcomes are the most common place an OIS shortfall shows up, because the number of participants needed to detect a given relative effect grows quickly as baseline event rates fall — this is the same reason publication bias and imprecision are often assessed together for rare-outcome comparisons.
- Imprecision downgrades compound with the other four domains. A pooled estimate that also carries risk-of-bias concerns or comes from a meta-analysis of studies with indirect populations can end up rated down two or more levels once imprecision is added, even though no single domain looks severe in isolation.
GRADE certainty ratings sit alongside, but are not identical to, older evidence hierarchies — see where the traditional levels-of-evidence pyramid and GRADE diverge for how the two frameworks handle this differently. For the underlying methodology this domain sits inside, the Cochrane Handbook chapter guide and PRISMA methodology overview cover where imprecision fits in a full systematic review workflow, and sample-size justification for a grant application or protocol covers the equivalent calculation from the trial-design side rather than the review side.
Frequently Asked Questions
Is optimal information size the same thing as statistical power?
They’re the same underlying concept applied at different levels. Statistical power is the property of a specific analysis given its sample size, effect size, and alpha level. Optimal information size is the sample-size number that calculation produces when applied to the comparison a meta-analysis is assessing — effectively, “how many participants would an adequately powered single trial of this comparison need.” GRADE then checks whether the pooled evidence base reaches that number.
What happens if a meta-analysis fails the OIS criterion but the confidence interval still looks precise?
GRADE guidance treats the OIS shortfall as sufficient grounds on its own to consider downgrading for imprecision, independent of interval width. The two criteria are checked separately; failing either one supports a downgrade, and a reviewer does not need the confidence interval to also look wide before acting on an OIS shortfall.
Does every GRADE imprecision downgrade come from an OIS shortfall?
No. A wide confidence interval that spans both “no important effect” and a clinically important effect can trigger a downgrade even when the accumulated sample size is adequate. OIS and interval width are two separate, independently sufficient triggers within the same domain, not one combined test.
Can imprecision downgrade certainty by more than one level?
Yes. GRADE guidance allows a two-level downgrade for imprecision when the evidence is very sparse relative to OIS or the interval is extremely wide — for example, very few events, very few participants, or an interval spanning both a large benefit and a large harm. A one-level downgrade is more typical when the shortfall is moderate.
Who calculates OIS — the review authors or the guideline panel?
Typically the systematic review or guideline development team performing the GRADE assessment, usually the same methodologists building the Summary of Findings table, since the calculation needs the same baseline-risk and minimally-important-difference inputs the panel is already using to judge clinical importance elsewhere in the assessment.








