Written and maintained by CASRAI Editorial Board
Last updated
The Naranjo algorithm — formally the Adverse Drug Reaction Probability Scale — is a ten-question instrument that converts a case narrative into a numeric probability band: definite, probable, possible or doubtful. It was published by Naranjo and colleagues at the University of Toronto in Clinical Pharmacology & Therapeutics in August 1981, and it is still the most widely applied causality instrument in the published case-report literature.
This page is about the method: what each of the ten questions is actually asking, how the arithmetic behaves, what the published reliability evaluations measured, and where the scale is documented to fail. It is written for a professional performing a causality assessment — a pharmacovigilance scientist, a hospital pharmacist, a safety physician, a quality reviewer signing off an individual case safety report. It is not clinical advice, and no drug–reaction pair on this page is asserted as causal.
The single most important thing to know before you use it: Naranjo was designed for controlled trials and registration studies, not for routine clinical practice, and three of its ten items — rechallenge, placebo administration and toxic drug concentration — are unavailable in most real post-market cases. That is not a marginal caveat. It is arithmetically why the scale almost never returns definite, a pattern confirmed in every large inter-rater study below.
What the Naranjo algorithm is, and what it was built for
Causality assessment answers one question: given this patient, this drug and this event, how much of the observed association is attributable to the drug? Before 1981 that judgment was almost entirely unstructured, which made it unauditable and unreproducible across reviewers. Naranjo et al. proposed a fixed set of questions with fixed weights so that two assessors working from the same record would, in principle, arrive at the same category.
The design intent matters and is routinely lost. NIDDK’s LiverTox monograph on the scale states plainly that it “was also designed for use in controlled trials and registration studies of new medications, rather than in routine clinical practice.” A controlled-trial setting is one where rechallenge is sometimes protocolised, placebo genuinely is administered, and plasma concentrations are genuinely measured. Move the same instrument to a spontaneous post-market report and three of its ten inputs vanish.
Naranjo sits alongside, not above, the other structures a safety function runs: MedDRA coding standardises what the event was, expedited reporting rules govern when it must be transmitted, and causality assessment governs how confidently the drug is implicated. The three are independent — a case can be MedDRA-coded, reportable, and causality-doubtful all at once.
The ten questions, with the exact scoring
Each item is answered Yes, No, or Do not know. “Do not know” is always zero. Total scores run from −4 to +13.
| # | Question | Yes | No | Do not know |
|---|---|---|---|---|
| 1 | Are there previous conclusive reports on this reaction? | +1 | 0 | 0 |
| 2 | Did the adverse event appear after the suspected drug was administered? | +2 | −1 | 0 |
| 3 | Did the adverse event improve when the drug was discontinued or a specific antagonist was administered? | +1 | 0 | 0 |
| 4 | Did the adverse event reappear when the drug was readministered? | +2 | −1 | 0 |
| 5 | Are there alternative causes that could on their own have caused the reaction? | −1 | +2 | 0 |
| 6 | Did the reaction reappear when a placebo was given? | −1 | +1 | 0 |
| 7 | Was the drug detected in blood or other fluids in concentrations known to be toxic? | +1 | 0 | 0 |
| 8 | Was the reaction more severe when the dose was increased, or less severe when the dose was decreased? | +1 | 0 | 0 |
| 9 | Did the patient have a similar reaction to the same or similar drugs in any previous exposure? | +1 | 0 | 0 |
| 10 | Was the adverse event confirmed by any objective evidence? | +1 | 0 | 0 |
What each question is actually asking
- Q1 — prior conclusive reports (+1 / 0 / 0)
- Is the reaction already established for this substance or class — labelled, or supported by published cases? This is a prior-probability item, not a within-case item. It is also the only question that rewards a well-characterised drug, which means a genuinely novel reaction for a new molecule is structurally penalised one point relative to a known one. Note the asymmetry: absence of prior reports scores 0, not −1, so the scale declines to treat novelty as evidence against.
- Q2 — temporality (+2 / −1 / 0)
- Did exposure precede the event? This is the heaviest positive item and the closest thing the scale has to a necessary condition. It tests sequence only, not plausibility of latency — a reaction beginning three years after first exposure scores identically to one beginning three hours after. That insensitivity to how long is the root of the long-latency problem discussed below.
- Q3 — dechallenge (+1 / 0 / 0)
- Did the event resolve or improve on withdrawal, or on giving a specific antagonist? A positive dechallenge is real evidence. But the item is under-specified in two directions: it sets no window for what counts as improvement “on” withdrawal, and it gives no credit for the informative negative case where a reaction is irreversible by mechanism and therefore cannot dechallenge. Both score 0 or +1 with no way to distinguish them.
- Q4 — rechallenge (+2 / −1 / 0)
- Did the event recur on re-exposure? Together with Q2 this is the scale’s strongest evidence — and for any serious reaction, deliberately re-exposing the patient is unethical. In practice this item is answered “Do not know” and scores 0 in exactly the cases where the causality question matters most. Inadvertent re-exposure occasionally supplies a genuine positive; when it does, record how it arose.
- Q5 — alternative causes (−1 / +2 / 0)
- Could the underlying disease, a comorbidity, or another agent have produced this on its own? Note the reversed polarity and the weighting: ruling alternatives out is worth +2, the joint-heaviest positive in the instrument, while finding one costs only −1. This is the single most judgment-laden item on the scale, and it is where most inter-rater disagreement originates — “could on their own have caused” is a threshold each assessor sets privately. It is also the item that misfires on drug–drug interactions, below.
- Q6 — placebo (−1 / +1 / 0)
- Did the reaction recur when placebo was substituted? A blinded-trial item with no post-market analogue. It also creates a live scoring ambiguity that materially moves totals: if the patient never received placebo, is the correct answer “No” (+1, because the reaction demonstrably did not recur on placebo) or “Do not know” (0, because the condition was never tested)? The scale does not adjudicate this, and the two readings differ by a full point on every case an assessor scores. Whichever convention your organisation adopts, write it down and apply it uniformly — otherwise your own assessors are not scoring the same instrument.
- Q7 — toxic concentration (+1 / 0 / 0)
- Was the drug measured in blood or another fluid at a concentration known to be toxic? This presumes an available, validated assay with an established toxic threshold — true for a small minority of drugs and essentially never for idiosyncratic, non-dose-related reactions. LiverTox notes specifically that reliance on drug levels “is rarely helpful in idiosyncratic drug induced liver disease.”
- Q8 — dose–response (+1 / 0 / 0)
- Did severity track the dose in either direction? Genuine evidence of a pharmacological mechanism when present. Absent for immune-mediated and idiosyncratic reactions by definition, and unobservable in any case where the drug was given at one dose and stopped.
- Q9 — prior similar reaction (+1 / 0 / 0)
- Has this patient reacted this way before, to this drug or a related one? A within-patient replication and strong evidence when documented. It depends entirely on the completeness of the medication history, so a thin record scores 0 for absence of data rather than absence of the phenomenon.
- Q10 — objective confirmation (+1 / 0 / 0)
- Is there an objective correlate — laboratory value, imaging, biopsy, measured physiological change — as opposed to symptom report alone? The most reliably answerable item on the scale for a well-documented case, and the one most improved by good source documentation.
The score bands
| Total score | Category | What the original paper means by it |
|---|---|---|
| ≥ 9 | Definite | Reasonable temporal sequence or an established toxic concentration; a recognised response to the drug; confirmed by improvement on withdrawal and recurrence on re-exposure. |
| 5 – 8 | Probable | Reasonable temporal sequence; a recognised response; confirmed by withdrawal but not by re-exposure; not reasonably explained by the patient’s clinical state. |
| 1 – 4 | Possible | A temporal sequence; possibly a recognised pattern; could also be explained by the patient’s disease. |
| ≤ 0 | Doubtful | The event is more plausibly attributed to something other than the drug. |
Read the definite definition again: it requires confirmation by re-exposure. The band definition itself concedes that the top category is reserved for cases where rechallenge occurred.
The arithmetic nobody publishes: why definite is nearly unreachable
Most published summaries of the Naranjo algorithm stop at the two tables above. The consequence of those tables is more useful than the tables themselves. Take a serious, non-rechallengeable, non-assayable reaction — the class that generates most real causality work — and score it as favourably as the evidence honestly allows:
| Item | Answer | Score |
|---|---|---|
| Q1 prior reports | Yes — reaction is labelled | +1 |
| Q2 temporality | Yes | +2 |
| Q3 dechallenge | Yes — resolved on withdrawal | +1 |
| Q4 rechallenge | Do not know — re-exposure would be unethical | 0 |
| Q5 alternative causes | No — none identified | +2 |
| Q6 placebo | Do not know — never administered | 0 |
| Q7 toxic concentration | Do not know — no assay | 0 |
| Q8 dose–response | Do not know — single dose level | 0 |
| Q9 prior similar reaction | No prior exposure documented | 0 |
| Q10 objective evidence | Yes — laboratory confirmation | +1 |
| Total | +7 — “Probable” | |
That is a textbook-strong case: labelled reaction, clean temporality, positive dechallenge, no competing cause, objective confirmation. It lands mid-band in probable and cannot reach definite. Push every remaining item to its maximum — grant a documented prior reaction in the same patient (Q9 +1) and an observed dose–response (Q8 +1), both uncommon — and the total reaches exactly 9, the bare floor of definite. With Q4, Q6 and Q7 unanswerable, 9 is the ceiling, and only a case that scores maximally on all seven remaining items attains it.
Two things follow, and they are the practical core of using this instrument honestly:
- A probable Naranjo score on a serious post-market case is not a weak finding. It is very often the strongest result the instrument is capable of producing for that case class. Reading it as “the tool could not confirm causality” misreads a structural ceiling as an evidentiary shortfall.
- Score compression is expected, not anomalous. If your case series returns almost nothing outside possible and probable, the instrument is behaving exactly as its arithmetic dictates. Do not tune your assessors toward the tails to fix it.
What the published evaluations actually found
The reliability literature is where honest use of Naranjo has to start, and the findings are consistent enough to state plainly.
Inter-rater reliability is moderate at best
Gallagher and colleagues (PLoS ONE 2011;6(12):e28096), developing the Liverpool tool, ran seven assessors across 80 observational cases and 37 published case reports — 819 causality assessments. Naranjo returned a kappa of 0.45 on the first 40 cases and remained “moderate” on the second 40. The Liverpool instrument scored 0.48 and then 0.60 (“good”) on the same material.
Pandit and colleagues (Journal of Patient Safety 2024;20(4):236–239) had five clinical-pharmacology raters independently assess 100 individual case safety reports on both Naranjo and WHO-UMC. Fleiss kappa across raters was 0.297 (fair) for Naranjo and 0.200 (slight) for WHO-UMC. Their conclusion was that the two scales carry a similar degree of subjectivity and that “more robust and less subjective scales are required.”
A counterweight worth recording: Pradhan and colleagues (Journal of Evaluation in Clinical Practice 2025;31(3):e70110) obtained a weighted kappa of 0.92 between two reviewers on 12 serious adverse events in a Canadian hospital. That is excellent agreement — but it is two trained reviewers, twelve cases, and a pilot design, and the same study reported sensitivity of 1.00 against specificity of 0.31. High reliability is achievable with few, well-calibrated, well-briefed assessors on richly documented cases; it does not generalise to a distributed reporting network.
The top category is almost never assigned
This is the empirical confirmation of the arithmetic above, and it is the finding most often omitted from summaries of the scale.
- Gallagher 2011, first 40 cases (280 assessments): Naranjo assigned its highest category 8 times; the Liverpool tool assigned its highest 125 times. Second 40 cases: Naranjo 4, Liverpool 133. Naranjo concentrated 172 and then 185 of 280 assessments into probable.
- Pandit 2024: across 100 reports and five raters — 500 assessments — no report was placed in the highest Naranjo category by any rater.
An instrument whose top band is effectively unpopulated is, in practice, a three-category instrument. Plan your case-handling rules around that.
Naranjo and WHO-UMC disagree substantially
More and colleagues (Journal of Pharmacological and Toxicological Methods 2024;127:107514) applied both systems to 399 anonymised individual case safety reports with blinded raters. WHO-UMC classed 53.3% as “Certain”; Naranjo classed 96.74% as “Probable”. Cohen’s kappa between the two systems was 0.22 — “minimal”.
The operational implication is direct: a causality category is only interpretable if the instrument is named alongside it. “Probable” carries different evidence on Naranjo than “Probable/Likely” does on WHO-UMC, and switching instruments mid-programme silently breaks any trend you are computing across your case series. Pandit 2024 found substantial-to-fair agreement between the scales depending on which rater was doing the assessing — the disagreement is not a constant offset you can correct for.
Where Naranjo is documented to fail
Long-latency reactions
Q2 tests sequence, not interval. A reaction emerging after months or years of continuous exposure scores the same +2 as one emerging within hours, and no item anywhere in the scale asks whether the observed latency is consistent with a plausible mechanism. LiverTox states the scale “is not weighted for the most critical elements in judging the likelihood of drug induced liver injury, such as specific time to onset, criteria for time of recovery, and list of critical diagnoses to exclude.” Long latency also degrades Q3 (dechallenge is slow, ambiguous, or confounded by other interventions) and Q5 (the longer the exposure, the more competing causes have accumulated).
Drug–drug interactions
This is the sharpest structural failure. Where the reaction arises from an interaction, the interacting partner is by construction an alternative cause that could on its own have caused the reaction — or at minimum the assessor must decide whether it is. Q5 therefore penalises the case by up to three points relative to a monotherapy case (−1 instead of +2), precisely because the mechanism is an interaction. The scale carries no item that represents a joint or potentiated effect. Assessing each suspect drug separately, the routine workaround, scores each one down for the presence of the other and can return two possible ratings where the real answer is one interaction.
Rechallenge is unavailable exactly where it is needed
Deliberate re-exposure after a serious reaction is not ethically defensible, so Q4’s +2 is unreachable for the case class that matters most. Combined with Q6 (no placebo outside a blinded trial) and Q7 (no assay for most idiosyncratic reactions), three items — worth up to four points and, on the Q6 ambiguity, a fifth — default to zero in ordinary post-market work.
The placebo item can subtract points
LiverTox flags a further oddity: points are subtracted if the reaction reappears on placebo, which “does not apply to the usual case” outside a controlled trial. An item that can only ever be neutral or negative in your setting is dead weight in the numerator of a score you are treating as a probability.
It was never built for your setting
All of the above is one root cause: an instrument designed for controlled trials and registration studies is being applied to spontaneous reports, retrospective chart review and post-market surveillance. LiverTox’s own comparison finds Naranjo “easier to apply, but has less sensitivity and specificity” than RUCAM for drug-induced liver injury. Easier to apply is a real virtue — it is most of why Naranjo dominates the case-report literature — but it is not the same virtue as accuracy.
A citation caution
The LiverTox monograph’s opening line dates the scale to 1991. The primary source is 1981: Naranjo CA, Busto U, Sellers EM, et al. “A method for estimating the probability of adverse drug reactions.” Clin Pharmacol Ther. 1981 Aug;30(2):239–45 (doi:10.1038/clpt.1981.154; PMID 7249508). Cite the 1981 paper.
The alternatives, and what each does differently
WHO-UMC causality categories
The Uppsala Monitoring Centre system is criteria-based rather than arithmetic. Instead of summing weighted answers, the assessor matches the case against definitions for six categories: Certain, Probable/Likely, Possible, Unlikely, Conditional/Unclassified, and Unassessable/Unclassifiable.
Two of those six have no Naranjo equivalent and are the system’s main advantage. Conditional/Unclassified holds a case pending more data; Unassessable/Unclassifiable records that the report is too incomplete or contradictory to judge. Naranjo has no way to say “this report cannot be assessed” — an unassessable case simply accumulates “Do not know” zeros and emerges as possible or doubtful, which reads as a substantive finding rather than as missing data. If your case series contains many thin reports, that distinction is not cosmetic: it is the difference between a data-quality problem you can see and one you cannot. WHO-UMC is the system recommended under the Pharmacovigilance Programme of India, among other national programmes.
Source note. UMC’s own definition document for these categories now sits behind a login on who-umc.org and could not be retrieved for verbatim quotation while writing this page. The six category names above are as reproduced consistently across the peer-reviewed pharmacovigilance literature, but for the authoritative criteria under each category, go to UMC directly rather than to any secondary rendering of the table — this one included.
Liverpool ADR Causality Assessment Tool
The Liverpool tool (Gallagher et al., PLoS ONE 2011;6(12):e28096) is a decision-tree flowchart, not a score. The assessor follows branching questions to a terminal category, which removes the “sum of weights” step entirely and with it the score-compression problem. It was developed by an expert focus group specifically to address Naranjo’s failure to use the full range of its own categories — and in head-to-head testing it did exactly that, assigning its top category 125 and 133 times per 280 assessments where Naranjo assigned 8 and 4. Its inter-rater reliability reached “good” (kappa 0.60) on the second case set against Naranjo’s persistent “moderate”. The authors’ own conclusion is appropriately restrained: further assessment by different investigators in different settings is needed to fully assess its utility. It originated in a paediatric context.
RUCAM
The Roussel Uclaf Causality Assessment Method is organ-specific — built for hepatotoxicity, and scoring the elements Naranjo omits: specific time-to-onset windows, the course of enzyme values after withdrawal, defined exclusion of competing diagnoses, and risk factors. It is harder to apply than Naranjo and more sensitive and specific for liver injury. The general lesson generalises past the liver: where a validated organ-specific instrument exists for the reaction you are assessing, it will usually beat a general-purpose one, because the elements that discriminate causality are organ-specific and a general instrument cannot weight them.
When a structured algorithm helps, and when judgement should override
The case for a structured instrument is not that it is more accurate than an expert. It is that it is legible: it forces the same ten considerations onto every case, it leaves an auditable trace, and it stops an assessment from resting on an unstated intuition. Those are real and they are why the algorithm belongs in a quality system.
Use the algorithm as the primary instrument when:
- You are assessing volume — a case series, a periodic safety report, a literature screen — and need consistency across assessors more than depth on any one case.
- The assessment will be read by someone who cannot re-examine the underlying record, such as a regulator, an ethics committee, or a journal reviewing a case report.
- The reaction is pharmacologically predictable, dose-related, short-latency and monotherapy — the profile the ten items were designed around.
- You need an unbiased first pass before expert opinion is brought in, or a check on it afterwards.
Let documented clinical judgement override the score when:
- The score is depressed by structural unanswerability rather than by weak evidence. Three “Do not know” zeros on Q4/Q6/Q7 are a property of the setting, not a finding about the drug. Say so in the narrative rather than reporting the number alone.
- A drug–drug interaction is the plausible mechanism. Q5 is actively misleading here. Assess the interaction as the exposure and reason in prose.
- The latency is long, or the reaction is immune-mediated or idiosyncratic. Q2, Q7 and Q8 all lose their discriminating power at once.
- A validated organ-specific instrument exists for the reaction in question. Use it, and if you also report Naranjo, report both and name both.
- The report is too incomplete to assess. Do not let “Do not know” zeros manufacture a possible or doubtful rating out of missing data. Record the case as unassessable — which is a reason to consider WHO-UMC, which has a category for exactly this.
- The score contradicts a strong specific finding — a positive inadvertent rechallenge, a diagnostic biopsy, a validated biomarker. A single decisive piece of evidence outweighs a general-purpose sum, and the sum has no item to carry it.
Judgement overriding a score is legitimate and expected. Judgement silently overriding a score is not: the reason is what makes it auditable, and an unexplained divergence between a documented score and a reported category is exactly what an inspector will ask about.
Documenting an assessment so it survives review
Recording only the total and the band throws away everything that makes the assessment reviewable. A defensible record carries:
- The instrument and version, named. “Naranjo (ADR Probability Scale), Naranjo et al. 1981” — not “causality assessment performed”. Given the 0.22 kappa between Naranjo and WHO-UMC, an unnamed category is uninterpretable.
- All ten item answers, not just the total. The pattern of zeros is the finding. A 7 built from three unanswerable items is a different case from a 7 built from equivocal evidence, and only the itemised record distinguishes them.
- A one-line justification for Q5. Which alternatives were considered, and why each was or was not judged sufficient on its own. This is the item that drives inter-rater disagreement; it is the one worth the sentence.
- Your organisation’s Q6 convention, applied consistently — “No” or “Do not know” where placebo was never administered. Write it into the SOP.
- Assessor identity and date, and for anything consequential, an independent second assessment. Published kappas in the 0.3–0.5 range are the argument for dual assessment on serious cases; they are not an argument against using the tool.
- Any override, with its reason, recorded next to the score rather than in place of it.
This is ordinary GxP documentation practice applied to a judgement rather than to a measurement: contemporaneous, attributable, and complete enough that a second reviewer can reach the same conclusion or identify precisely where they diverge.
Frequently asked questions
What is a good Naranjo score?
There is no single threshold. ≥9 is definite, 5–8 probable, 1–4 possible, ≤0 doubtful. But for serious post-market cases the achievable ceiling is typically 9 and realistically 7, because rechallenge, placebo and toxic-concentration items default to zero. Judge a score against what that case class can attain, not against the −4 to +13 range.
Why does the Naranjo scale rarely give a “definite” result?
Because Q4 (+2, rechallenge), Q6 (up to +1, placebo) and Q7 (+1, toxic concentration) are unavailable in most non-trial settings, capping the realistic maximum at 9 — the exact floor of the definite band. The empirical record matches: Gallagher 2011 recorded 8 and then 4 top-category assignments per 280, and Pandit 2024 recorded none across 500 assessments.
Naranjo or WHO-UMC — which should we use?
They measure differently and agree only minimally (kappa 0.22 in More 2024 across 399 reports). Naranjo is faster, more mechanical and dominant in the case-report literature; WHO-UMC is criteria-based and uniquely able to record a case as Conditional/Unclassified or Unassessable/Unclassifiable. Choose one per programme, name it in every record, and do not switch mid-series without re-assessing the back catalogue.
Can Naranjo assess a drug–drug interaction?
Not well. It has no item for a joint effect, and Q5 treats the interacting partner as an alternative cause — costing up to three points relative to a monotherapy case precisely because an interaction is the mechanism. Assess the interaction narratively and record why the score understates it.
How reliable is the Naranjo algorithm between assessors?
Moderate. Kappa 0.45 across seven assessors in Gallagher 2011; Fleiss kappa 0.297 across five raters in Pandit 2024. A small Canadian pilot (Pradhan 2025, 12 cases, two reviewers) reached 0.92 weighted kappa with sensitivity 1.00 and specificity 0.31 — good agreement is achievable with few, well-calibrated assessors, but does not generalise to a distributed network.
Does a low Naranjo score mean the drug did not cause the reaction?
No. It means the ten items, as answered, did not accumulate points — which can equally reflect a thin report, an unavailable rechallenge, or a missing assay. Distinguish “the evidence points away from the drug” from “the instrument could not interrogate this case”, and record which one you mean.
Is Naranjo required by regulators?
No regulator mandates this specific instrument. Causality assessment obligations attach through frameworks such as ICH E2A and 21 CFR 314.80, and through Good Pharmacovigilance Practices, but the choice of tool is generally yours to make, justify in your procedures, and apply consistently.
Related reading
- Pharmacovigilance in clinical research — how AE, SAE and SUSAR reporting obligations are triggered, which is a separate question from causality.
- Side effect vs adverse event — the terminology distinction that determines whether a causality assessment is even in scope.
- FDA Adverse Event Reporting System (FAERS) and the FDA Sentinel Initiative — the spontaneous and active-surveillance systems whose individual case reports are what causality assessment is applied to.
- VAERS reporting — the vaccine-specific spontaneous system, and what its data can and cannot support.
- Adverse event of special interest (AESI) and the IND safety report — where a causality judgement becomes a reporting trigger.
- Adverse event reporting to the IRB — the parallel human-subjects reporting track.
- Pharmacovigilance certification — where formal training in these methods sits.
- CARE case report guidelines — the reporting standard for the published case reports in which Naranjo scores most often appear.
- Safety and pharmacovigilance — the wider cluster this method sits within.
Primary sources
- Naranjo CA, Busto U, Sellers EM, Sandor P, Ruiz I, Roberts EA, Janecek E, Domecq C, Greenblatt DJ. A method for estimating the probability of adverse drug reactions. Clin Pharmacol Ther. 1981 Aug;30(2):239–45. doi:10.1038/clpt.1981.154. PMID 7249508
- LiverTox: Clinical and Research Information on Drug-Induced Liver Injury. Adverse Drug Reaction Probability Scale (Naranjo) in Drug Induced Liver Injury. Bethesda (MD): NIDDK; updated 4 May 2019. NBK548069 — source of the item table, point values and score bands reproduced above.
- Gallagher RM, Kirkham JJ, Mason JR, Bird KA, Williamson PR, Nunn AJ, Turner MA, Smyth RL, Pirmohamed M. Development and inter-rater reliability of the Liverpool adverse drug reaction causality assessment tool. PLoS ONE. 2011;6(12):e28096. PMID 22194808
- Pandit S, Soni D, Krishnamurthy B, Belhekar MN. Comparison of WHO-UMC and Naranjo scales for causality assessment of reported adverse drug reactions. J Patient Saf. 2024 Jun 1;20(4):236–239. PMID 38345209
- More SA, Atal S, Mishra PS. Inter-rater agreement between WHO-Uppsala Monitoring Centre system and Naranjo algorithm for causality assessment of adverse drug reactions. J Pharmacol Toxicol Methods. 2024;127:107514. PMID 38768933
- Pradhan P, Corbin S, Todkar S, et al. Serious adverse events: a replicability and validation study of Naranjo causality assessment tool in a Canadian clinical setting. J Eval Clin Pract. 2025 Apr;31(3):e70110. PMID 40275458
- Uppsala Monitoring Centre — who-umc.org, for the authoritative WHO-UMC category criteria (see the source note above).








