Written and maintained by CASRAI Editorial Board
Last updated
Dose-response meta-analysis pools not just whether an exposure has an effect, but how the effect changes as the dose or exposure level changes — turning a set of studies that each report several exposure-category risk estimates into one summary trend or curve. It is the standard approach when the included studies report results for three or more exposure categories (light/moderate/heavy, or specific dose bands) rather than a single exposed-vs-unexposed contrast, and the review question is “how much does risk change per unit of exposure,” not just “is there an association at all.”
What problem this solves that a standard meta-analysis doesn’t
A conventional pairwise meta-analysis pools one effect estimate per study — typically a single relative risk (RR), odds ratio (OR), or hazard ratio (HR) comparing an exposed group to an unexposed or reference group. That works when exposure is genuinely binary. It throws away information when the underlying studies instead report a graded set of estimates: for example, a cohort study on alcohol and a health outcome might report separate RRs for 1-2, 3-4, and 5+ drinks per day, each relative to non-drinkers. Collapsing that into a single “any exposure vs. none” contrast discards the shape of the relationship — and the shape is often the clinically or scientifically important part of the answer, particularly for questions like whether risk rises steadily, plateaus, or only appears above a threshold.
Dose-response meta-analysis instead treats each study’s set of category-specific estimates as data on a curve and models that curve across studies. The output is a pooled trend (a single slope, if the relationship is modelled as linear) or a pooled curve (if modelled flexibly), typically presented as a plot of relative risk against exposure level with a 95% confidence band.
The correlation problem, and why it needs its own method
The technical obstacle is that a study’s own exposure-category estimates are not independent of each other. If a cohort reports RRs of 1.3, 1.6, and 2.1 for three increasing dose categories, all three are calculated against the same reference group — so they share sampling error from that reference group and are statistically correlated. Naively treating each category as its own independent data point and running an ordinary regression on the log-RRs overstates precision and can bias the pooled trend.
The standard solution is the method described by Sander Greenland and Matthew Longnecker in “Methods for Trend Estimation from Summarized Dose-Response Data, with Applications to Meta-Analysis” (American Journal of Epidemiology, 1992;135(11):1301-1309) — almost universally referred to in the applied literature as the Greenland-Longnecker method. It reconstructs the approximate covariance matrix of a study’s log-RRs (or log-ORs/log-HRs) from summary data alone — the reported point estimates, their confidence intervals, and the case/person counts per category — without needing the original individual-level data. That reconstructed covariance structure is what lets a study’s several category estimates be combined correctly into one dose-response trend per study, and then pooled correctly across studies in the second stage.
What you need to extract from each included study
Because the method works from summarized category data, a systematic data extraction pass for a dose-response meta-analysis needs a specific, consistent set of fields per exposure category, not just one overall effect estimate per study:
- An assigned dose value per category. Studies rarely report a single dose value for a category (“5-9 cigarettes/day”) — the convention is to use the category midpoint when a closed range is given, and for the reference category, the value is fixed at the study’s own reference point (usually zero or the lowest category’s midpoint). Open-ended top categories (“20+”) are a recurring extraction problem; common practice is to assign a value based on the width of the preceding category or a distributional assumption, and to flag this as a source of uncertainty rather than treat it as exact.
- The effect estimate and its precision for each category relative to the study’s reference group — RR, OR, or HR, plus the confidence interval or standard error needed to reconstruct variance.
- The number of cases and the number of participants (or person-time) in each category, including the reference category — required by the Greenland-Longnecker covariance reconstruction, and also what most software implementations expect as direct input alongside the effect estimates.
- Which category is the reference, and whether it is the lowest exposure category or a separate “unexposed” group — this determines how the correlation structure is built and needs to be recorded explicitly, not assumed to be uniform across studies.
Piloting this extraction on a handful of included studies before running it at scale is worth doing specifically because open-ended categories and inconsistent reference groups are the most common source of downstream errors — see the companion guide on building and piloting a data extraction form for the general process this fits into.
Modelling the shape: linear trend vs. non-linear curve
Once category-level data are extracted and correlation-adjusted, there are two broad modelling choices:
Linear trend. The simplest model estimates a single slope — the change in log-relative-risk per unit increase in exposure — and assumes that relationship holds across the whole observed dose range. It answers “does risk rise (or fall) on average as exposure increases,” and produces one pooled number, easy to report and interpret, but it enforces a straight-line shape even where the real relationship plateaus, reverses, or is J-shaped.
Non-linear modelling. Where a straight line is not a defensible assumption, restricted cubic splines are the standard flexible alternative — piecewise cubic polynomials joined smoothly at a small number of “knots” along the exposure range, constrained to be linear beyond the outermost knots so the curve does not behave erratically at the extremes of sparse data. Vincenzo Bagnardi and colleagues’ methodological work on flexible meta-regression functions for aggregate dose-response data (published in the mid-2000s) is the standard reference for applying restricted cubic splines and fractional polynomials in this setting. Three to five knots, often placed at fixed percentiles of the pooled exposure distribution, is typical; more knots allow more flexibility but need more data to estimate reliably. Fractional polynomials are a less commonly used alternative that fits a small number of power-transformed terms rather than a piecewise spline.
The practical choice between linear and non-linear modelling should be driven by the scientific question and by a formal test for non-linearity (comparing the linear and spline models), not chosen after inspecting which one looks better — deciding the shape post hoc on the basis of the results is a form of the same selective-reporting problem covered in the guide on pre-specification in subgroup and sensitivity analysis.
One-stage vs. two-stage estimation
There are two structurally different ways to get from per-study category data to a pooled dose-response curve:
Two-stage. First, within each study, the Greenland-Longnecker method (or an equivalent) is used to fit a study-specific dose-response curve (linear or spline) from that study’s own category data. Second, the resulting study-specific curves — represented as coefficients with their estimated covariance — are pooled across studies using multivariate random-effects meta-analysis, conceptually similar to how meta-regression pools study-level estimates, but pooling a vector of spline or slope coefficients per study rather than a single effect size. This is the more established, more widely implemented approach.
One-stage. All studies’ category-level data points are modelled together in a single mixed-effects model, with study-level random effects and covariance structure imposed directly, rather than summarizing each study first and pooling second. A one-stage model can be more efficient and more stable when individual studies contribute only a small number of exposure categories — the two-stage approach can struggle to even estimate a study-specific curve from very few points — but it is more complex to specify and less widely implemented in off-the-shelf software.
In practice, most published dose-response meta-analyses use the two-stage approach, implemented via the dosresmeta package in R (built specifically around this method, and documented with worked examples in its CRAN vignette), sometimes alongside Stata’s user-written drmeta/glst commands, which implement closely related Greenland-Longnecker-based routines.
Where this connects to GRADE and causal inference
A demonstrated dose-response gradient is not just a modelling nicety — it is one of the explicit criteria the GRADE approach uses to upgrade the certainty of evidence from observational studies, alongside a large effect size and the absence of plausible confounding that would explain the association away. See the GRADE evidence-certainty dictionary entry and the fuller GRADE Evidence-to-Decision framework guide for how this fits into an overall certainty rating. It also echoes one of Bradford Hill’s classic considerations for judging whether an observed association is likely causal: a biological gradient, where risk increases consistently with exposure, is treated as evidence in favor of causation, though — as with all of Hill’s considerations — it is a judgment aid, not a mechanical test, and a dose-response pattern can still arise from confounding that itself varies by exposure level.
Because dose-response meta-analyses most often pool observational (cohort or case-control) studies rather than randomized trials, per-study risk-of-bias assessment matters as much here as in any other observational evidence synthesis — see the guide on ROBINS-I for non-randomised studies for the companion appraisal step that should sit alongside the quantitative modelling, and the guide on heterogeneity in meta-analysis for interpreting how much the study-specific curves actually agree with each other before trusting a single pooled shape.
Common pitfalls
- Assigning an arbitrary value to open-ended top or bottom categories without sensitivity-testing how much the pooled trend changes under a different assumption.
- Choosing linear vs. non-linear after seeing the results, rather than pre-specifying the modelling approach and testing for non-linearity formally.
- Treating each category estimate as an independent data point and skipping the Greenland-Longnecker (or equivalent) correlation adjustment entirely — this understates uncertainty and is the single most common methodological shortcut to check for when appraising a published dose-response meta-analysis.
- Extrapolating the pooled curve beyond the range of exposure actually observed in the included studies, where data are sparse and the model’s behavior is least reliable.
Frequently asked questions
Do I need dose-response meta-analysis if my studies only report exposed vs. unexposed?
No — if every included study reports a single binary contrast, there is no per-study dose-response shape to model, and a standard pairwise meta-analysis is the right approach. Dose-response methods apply once studies report three or more ordered exposure categories.
How many exposure categories does a study need to contribute?
At least three (including the reference) to estimate any trend at all, and more are needed to support non-linear modelling with restricted cubic splines. Studies contributing only two categories can still be included in a two-stage pooled analysis as long as at least some included studies contribute three or more, but they cannot support a study-specific spline on their own.
Is restricted cubic spline modelling always better than a linear trend?
Not automatically — it is more flexible, but that flexibility costs precision, especially with a modest number of studies. A pre-specified test for non-linearity (or a fit-statistic comparison) should guide the choice, and reporting a linear trend as a secondary, simpler summary alongside a spline curve is common practice.
What software actually implements the Greenland-Longnecker method?
The dosresmeta package in R is the most widely used, purpose-built implementation, covering both one-stage and two-stage, linear and spline models. Stata users commonly rely on the user-written glst and drmeta commands, which implement closely related routines.








