Written and maintained by CASRAI Editorial Board
Last updated
A tornado diagram is the standard way to report one-way deterministic sensitivity analysis (DSA) in a health-economic decision model: a horizontal bar chart, ranked from the largest bar at the top to the smallest at the bottom, that shows how much the model’s output changes when a single input parameter is moved across its plausible range while every other parameter is held at its base-case value. The name comes from the shape — the widest bars cluster at the top and the chart narrows toward the bottom, resembling the funnel of a tornado.
Tornado diagrams are a reporting output of cost-effectiveness analysis and cost-utility analysis models submitted to health technology assessment (HTA) bodies, and their inclusion is expected practice under the CHEERS 2022 reporting checklist‘s requirement to characterize uncertainty. This guide covers how a tornado diagram is built, how to read one, and where it breaks down as evidence — specifically its inability to capture parameter interactions or joint uncertainty, which is why regulators and HTA agencies also require a full probabilistic sensitivity analysis alongside it.
What one-way sensitivity analysis is testing
A decision-analytic model (a decision tree, a Markov cohort model, or a discrete-event simulation) produces a single output for its base case — typically an incremental cost, an incremental effect measured in quality-adjusted life years (QALYs), or a derived summary statistic like the incremental cost-effectiveness ratio (ICER). That base case rests on dozens of input parameters: drug acquisition costs, event probabilities, utility values, discount rates, resource-use assumptions, and so on. Each of those parameters is itself an estimate with genuine uncertainty around it — drawn from a clinical trial’s confidence interval, a costing study’s standard error, or an expert-elicited plausible range.
One-way DSA asks a narrow, mechanical question of each parameter in turn: if this single input were actually at the low end of its plausible range, or the high end, and everything else stayed exactly at its base-case value, how far would the model’s output move? The procedure, repeated across the full parameter list, is:
- Run the model once at base-case values for every parameter to establish the reference output (e.g., the base-case ICER or net monetary benefit).
- For one parameter, substitute its low bound (commonly the lower limit of a 95% confidence interval, or another pre-specified plausible minimum) while every other parameter stays at base case. Re-run the model and record the output.
- Repeat step 2 with that same parameter’s high bound. This gives a low-output and a high-output for that one parameter, and the difference between them is its “swing” or range of influence.
- Reset the parameter to base case, move to the next parameter, and repeat the whole low/high pair for it.
- Once every parameter has been swung individually, rank them by the width of their swing — largest range at the top, smallest at the bottom — and plot each as a horizontal bar anchored at the base-case value, extending left and right to the low- and high-output results.
Because only one input moves at a time, this is deterministic in the strict sense: each bar comes from a single, reproducible model run pair, not from a distribution or a random draw. That is both the method’s main strength (it is transparent and easy to audit — a reviewer can re-run any one bar by hand) and its central limitation, covered below.
Choosing the parameter ranges
The credibility of a tornado diagram depends entirely on how the low/high range for each parameter was chosen, and this should always be stated explicitly in the model’s methods — the CHEERS 2022 reporting standard requires it. Common, defensible sources for a parameter’s range are the 95% confidence interval from the trial or meta-analysis it was drawn from, a costing source’s reported standard error converted to a interval, or, where no empirical range exists, a plausible minimum/maximum set by clinical or expert judgment and labelled as such. A tornado diagram built on arbitrarily chosen ranges (a blanket ±20% applied to every parameter regardless of how well- or poorly-characterized it actually is) is a weaker piece of evidence than one built on parameter-specific empirical ranges, and reviewers at HTA bodies routinely query this.
How to read a tornado diagram
Reading a tornado diagram correctly comes down to three things:
- Bar order is the ranking of influence. The parameter at the top has the widest swing in the model’s output across its plausible range — it is the input the result is most sensitive to. A parameter near the bottom barely moves the output even at its extremes, so refining that estimate further would do little to change the model’s conclusion.
- The vertical reference line is the base-case result. Every bar is anchored to it. A bar that crosses the reference line only slightly in one direction and far in the other is asymmetric — the parameter’s effect on the output is not linear across its range, which is itself informative (it may signal a threshold effect, such as a probability approaching a boundary).
- Whether any bar crosses a decision-relevant threshold matters more than its raw width. If the outcome is an ICER being compared against a willingness-to-pay threshold, the question isn’t just which bar is widest — it’s whether any bar’s low or high end crosses the threshold and flips the cost-effectiveness conclusion. A parameter with a moderate swing that crosses the threshold is more decision-relevant than a wider swing that never approaches it.
Because bars are ordered top-to-bottom by influence, a tornado diagram functions as a triage tool: the handful of parameters at the top are the ones worth prioritizing for better data (a more precise cost estimate, a longer follow-up for an event-probability estimate, an additional utility elicitation study) because they are the ones actually driving model uncertainty. Parameters low on the chart are, for this model’s purposes, effectively settled — more precision on them would not change the conclusion.
Illustrative worked example
The figures below are an illustrative composite constructed to demonstrate the method, not a result from any real published evaluation or CASRAI’s own analysis of any actual intervention. Any resemblance to a specific drug, device, or published cost-effectiveness study is coincidental.
Consider a simplified cost-utility model comparing a new maintenance therapy against standard care over a lifetime horizon, with a base-case ICER of $42,000 per QALY gained. Five inputs are varied one at a time across their 95% confidence intervals:
| Parameter | Base case | Range tested | Resulting ICER range |
|---|---|---|---|
| Annual drug acquisition cost | $18,000 | $14,500 – $22,500 | $31,000 – $54,000 |
| Utility gain per QALY on treatment | 0.12 | 0.08 – 0.16 | $33,500 – $57,000 |
| Annual event-avoidance probability | 0.22 | 0.15 – 0.30 | $38,000 – $47,500 |
| Discount rate (costs and effects) | 3.5% | 1.5% – 5.0% | $40,500 – $44,000 |
| Time on treatment before discontinuation | 4 years | 2 – 6 years | $41,000 – $43,500 |
Ranked by the width of the swing (highest to lowest), the tornado diagram would place utility gain per QALY at the top, drug acquisition cost second, event-avoidance probability third, and the two remaining parameters — discount rate and treatment duration — as narrow bars near the bottom. If the willingness-to-pay threshold under consideration is $50,000 per QALY, the diagram immediately shows that the conclusion is genuinely uncertain: both the top two parameters’ high-output ends cross that threshold, meaning plausible-but-unfavorable values for either input alone would flip the intervention from cost-effective to not. The bottom two parameters, by contrast, never come close to the threshold at either extreme — refining the discount-rate assumption further would not change the recommendation, however methodologically tidy it might feel to do so.
What a tornado diagram cannot tell you
The one-way method’s core limitation follows directly from how it is built: because every bar is generated by moving exactly one parameter while holding all others fixed, a tornado diagram cannot show what happens when multiple uncertain parameters are unfavorable at the same time, and it treats every parameter as if its uncertainty were independent of every other parameter’s. Neither assumption holds in most real models. Specifically:
- No parameter interactions. If a low event-avoidance probability and a high drug cost are both plausible simultaneously, their combined effect on the ICER is not visible anywhere on a tornado diagram — each was tested in isolation, with the other held at base case.
- No correlation structure. Real parameters are often correlated (a sicker trial population might simultaneously have lower baseline utility and higher event rates); one-way DSA has no mechanism to represent that relationship.
- No joint probability distribution over the output. A tornado diagram shows a range for each input individually, not a probability that the overall result falls above or below a given threshold once every parameter’s uncertainty is considered together.
- The choice of range is a judgment call, not a distributional statement. A 95% confidence interval used as a DSA range is a statement about one parameter estimated in isolation; it says nothing about how likely the *combination* tested is.
This is exactly the gap that probabilistic sensitivity analysis (PSA) is designed to close. Instead of moving one parameter at a time, PSA assigns each uncertain parameter a probability distribution reflecting its estimated uncertainty (a beta distribution for a probability, a gamma or lognormal distribution for a cost, and so on), then runs the model thousands of times — typically via Monte Carlo simulation — drawing every parameter simultaneously from its distribution on each iteration. The output is a full distribution of possible ICERs (or net benefits) that reflects joint parameter uncertainty and any specified correlation, usually summarized as a cost-effectiveness acceptability curve showing the probability the intervention is cost-effective across a range of willingness-to-pay thresholds. The NICE technology appraisal process and most HTA bodies expect PSA as the primary characterization of decision uncertainty, with a one-way tornado diagram retained alongside it because it does something PSA’s aggregate output does not: it names, individually, which specific inputs are worth better data.
The two techniques are complementary, not substitutes for one another, and the ISPOR-SMDM Modeling Good Research Practices Task Force’s report on parameter estimation and uncertainty (Briggs AH, Weinstein MC, Fenwick EA, Karnon J, Sculpher MJ, Paltiel AD. “Model Parameter Estimation and Uncertainty: A Report of the ISPOR-SMDM Modeling Good Research Practices Task Force–6.” Value in Health, 2012;15(6):835–42) is the standard methodological reference recommending exactly this combination: deterministic one-way analysis (tornado diagrams) to identify influential parameters and communicate them transparently, alongside probabilistic analysis to characterize overall decision uncertainty for the reader who needs to know how confident the model’s conclusion actually is. Neither is a “better” method in isolation — they answer different questions, and a submission that reports only one is incomplete by current good-practice standards.
Related sensitivity-analysis techniques exist outside decision modeling and are worth distinguishing from this one-way DSA/tornado approach: sensitivity analysis in meta-analysis tests how robust a pooled effect estimate is to excluding individual studies or changing inclusion criteria, and prior sensitivity analysis in Bayesian statistics tests how a posterior estimate changes under alternative prior specifications. Both share the “vary one thing, see what moves” logic with tornado-diagram DSA, but none of them are interchangeable — each answers a question specific to its own modeling context.
FAQ
Is a tornado diagram the same thing as a sensitivity analysis?
No. A tornado diagram is a way of reporting the results of one specific type of sensitivity analysis — one-way deterministic sensitivity analysis. Sensitivity analysis more broadly is any systematic test of how a model’s conclusions change under different assumptions, and includes probabilistic sensitivity analysis and scenario analysis as well, neither of which is typically presented as a tornado diagram.
Why are the bars different widths?
Bar width represents how much the model’s output (e.g., the ICER) changes when that specific parameter is moved across its plausible low-to-high range, holding everything else fixed. A wide bar means the output is highly sensitive to that parameter’s uncertainty; a narrow bar means the output barely moves even at the parameter’s extremes.
Can a tornado diagram show two parameters being uncertain at once?
Not directly. Each bar isolates exactly one parameter. Some modelers add a supplementary multi-way or scenario analysis to test specific combinations of concern, but that is a separate analysis reported alongside the tornado diagram, not something the diagram itself shows.
Does CHEERS 2022 require a tornado diagram specifically?
The CHEERS 2022 checklist requires that a health-economic evaluation characterize uncertainty in its parameters and report the methods and results of that characterization; it does not mandate a specific chart format. A tornado diagram is the conventional way modelers satisfy that requirement for deterministic one-way results, but the substantive requirement is disclosure of the ranges tested and their effect on the result, however it is displayed.
What determines which parameters get tested in a one-way sensitivity analysis?
In principle, every parameter with a genuine, quantifiable range of uncertainty is a candidate. In practice, modelers typically test all parameters with an empirical range available (from a trial, meta-analysis, or costing source) and may omit structural assumptions that are instead tested through separate scenario analyses. The full parameter list and the source of each range should be reported in the model’s methods, not just the resulting chart.








