From disproportionality signal to WHO-UMC verdict — structuring an ICSR, scoring it on the Naranjo scale, and classifying drug-event causality
Modern pharmacovigilance begins not with a clinician's suspicion but with statistics run overnight against hundreds of millions of individual case safety reports (ICSRs). Disproportionality analysis asks a narrow question: is this drug-event pair reported more often than would be expected if drug and event were independent? A positive statistical signal is the trigger — not the conclusion — for the causality assessment work that follows.
A 2×2 contingency table underlies every disproportionality statistic: reports mentioning the drug of interest and the event of interest (a), the drug with other events (b), other drugs with the event (c), and other drugs with other events (d), all counted across the full spontaneous-reporting database.
Proportional Reporting Ratio (PRR) = [a/(a+b)] / [c/(c+d)] • Evans criteria (2001) flag a signal when PRR ≥ 2, χ² ≥ 4, and n ≥ 3 cases — the working default at many national centers including the UK MHRA Yellow Card scheme.
Reporting Odds Ratio (ROR) = (a×d) / (b×c) • Preferred by EudraVigilance's EVDAS screening algorithm; a signal is flagged when the lower bound of the 95% confidence interval exceeds 1.0, since this is more conservative than the point estimate alone.
Empirical Bayes Geometric Mean (EBGM) — the Multi-item Gamma Poisson Shrinker (MGPS): • Used operationally inside FDA's Sentinel Initiative and by Oracle/Lincoln Technologies' signal-management platforms. • Shrinks noisy small-count ratios toward 1 using a Bayesian prior, reducing false positives from rare drug-event pairs with only 3–4 reports. • A signal is typically actioned when EB05 (5th-percentile bound of the EBGM posterior) ≥ 2.
All three approaches share the same blind spot: they measure reporting disproportionality, not biological causation. A disproportionate signal can arise from channeling bias, media attention (the "Weber effect" of reporting spikes after launch), notoriety bias, or genuine pharmacological causation — which is exactly why every statistical signal must pass through structured case-level causality assessment before any regulatory action is taken.
VigiRank, WHO-UMC's machine-learning-assisted triage tool, combines disproportionality with case-quality features (completeness score, number of well-documented reports, literature co-mentions) to rank ~150 new candidate signals per quarter for human assessor review — cutting manual screening load by roughly 40% without lowering sensitivity for genuine safety signals.
Spontaneous report databases like FAERS, EudraVigilance, and VigiBase are passive surveillance systems: reporting is voluntary (in most jurisdictions), reporting rates vary by drug age, country, and publicity, and there is no reliable denominator of total drug exposure. This makes any single disproportionality score an estimate of reporting association, confounded by:
• Notoriety bias: a drug under media scrutiny (e.g., following a black-box warning) receives a surge of reports for unrelated events simply because clinicians are watching more closely. • Channeling bias: a new drug is preferentially prescribed to sicker patients who were failed by first-line therapy, inflating apparent event rates that reflect underlying disease severity, not drug toxicity. • Under-reporting: only an estimated 1–10% of adverse drug reactions are ever reported to spontaneous systems (the "iceberg phenomenon" described by Rawlins, 1988). • Duplicate reports: the same case reported by patient, physician, and manufacturer can inflate case counts before deduplication algorithms (probabilistic record linkage on patient age, sex, event date, and drug) merge them.
Because of these limitations, every jurisdiction's pharmacovigilance guidance — ICH E2E, EMA GVP Module IX (Signal Management), and FDA's Guidance for Industry on Postmarketing Safety Reporting — treats a disproportionality signal strictly as the entry point into a structured case-review and causality-assessment workflow, never as a stand-alone basis for regulatory action.
A raw individual case safety report arrives as a mixture of free text, lab values, and dosing records. Before any causality algorithm can run, the case must be transcoded into a standardized, machine-comparable structure: ICH E2B(R3) XML fields for exposure and outcome, MedDRA hierarchical terms for the reported reaction, and an explicit exposure-to-onset timeline extracted from often-inconsistent narrative dates.
The Medical Dictionary for Regulatory Activities (MedDRA), maintained by the ICH and MSSO, is the mandatory terminology for adverse event coding in FAERS, EudraVigilance, and VigiBase alike. A trained pharmacovigilance coder (or increasingly, an NLP auto-coding engine validated against human-coded gold standards) maps the verbatim reported term to the most specific applicable MedDRA Preferred Term (PT), which rolls up through Lowest Level Term (LLT) → Preferred Term (PT) → High Level Term (HLT) → High Level Group Term (HLGT) → System Organ Class (SOC).
Coding precision matters enormously for signal work: "swelling of the leg" might code to Peripheral swelling, Oedema peripheral, or — if the narrative supports it — the more clinically specific Deep vein thrombosis. Under-specific coding dilutes a true signal across multiple loosely related PTs; over-specific coding can fragment a real signal into terms too rare to reach statistical significance individually. This is why WHO-UMC and EMA both recommend Standardised MedDRA Queries (SMQs) — curated groupings of PTs representing a single medical concept (e.g., the SMQ "Anaphylactic reaction") — for signal detection, rather than single-PT counts alone.
ICH E2B(R3), harmonized under the ICH E2 series alongside E2A (definitions), E2D (postmarketing reporting), and E2E (pharmacovigilance planning), defines the XML message structure that every regulatory ICSR transmission must conform to — whether submitted by a manufacturer to FDA via the FDA Electronic Submissions Gateway, or to EMA via the EudraVigilance Gateway.
Key structured fields the causality assessor depends on: • Drug section: substance, brand name, indication, dose, route, start/stop dates, action taken with drug (withdrawn, dose reduced, unchanged) • Reaction section: MedDRA PT, onset date, outcome (recovered, recovering, not recovered, fatal, unknown), seriousness criteria (hospitalization, life-threatening, disability, death, congenital anomaly, other medically important) • Rechallenge/dechallenge fields: explicit coded outcome, not free text — critical because these two fields alone can contribute up to 5 of the 13 possible Naranjo points • Narrative case summary: free text field the assessor reads to catch nuance the structured fields miss (e.g., a confounding new drug started the same week)
Case completeness scoring (WHO-UMC's completeness score, and the analogous vigiGrade algorithm) weights the presence of these key fields — a case missing onset date or dechallenge outcome receives a lower documentation-quality weight in signal strength calculations, independent of the disproportionality statistics.
A 2019 Uppsala Monitoring Centre audit found that only about 35% of ICSRs in VigiBase contained a codeable rechallenge outcome — meaning two-thirds of reports are structurally incapable of ever reaching the WHO-UMC "Certain" category, regardless of how compelling the rest of the narrative is.
Published by Naranjo et al. in Clinical Pharmacology & Therapeutics (1981), the Naranjo scale remains the most widely used individual-case causality instrument worldwide because of its simplicity: ten yes/no/unknown questions, each with a fixed point weight, summed to a single interpretable score. It was designed for bedside and case-review use without requiring a trained pharmacovigilance physician, and it remains embedded in FDA MedWatch review workflows and most hospital ADR committees.
Each question is answered Yes / No / Unknown (Do not know), with points awarded only for definitive answers:
1. Are there previous conclusive reports of this reaction? Yes +1 / No 0 / Unknown 0 2. Did the adverse event appear after the suspected drug was administered? Yes +2 / No −1 / Unknown 0 3. Did the reaction improve when the drug was discontinued or a specific antagonist administered (dechallenge)? Yes +1 / No 0 / Unknown 0 4. Did the reaction reappear when the drug was readministered (rechallenge)? Yes +2 / No −1 / Unknown 0 5. Are there alternative causes that could on their own have caused the reaction? Yes −1 / No +2 / Unknown 0 6. Did the reaction reappear when a placebo was given? Yes −1 / No +1 / Unknown 0 7. Was the drug detected in blood/body fluids at a concentration known to be toxic? Yes +1 / No 0 / Unknown 0 8. Was the reaction more severe when the dose was increased, or less severe when decreased? Yes +1 / No 0 / Unknown 0 9. Did the patient have a similar reaction to the same or similar drug in any previous exposure? Yes +1 / No 0 / Unknown 0 10. Was the adverse event confirmed by any objective evidence (lab test, imaging, biopsy)? Yes +1 / No 0 / Unknown 0
Score interpretation: • ≥9 — Definite adverse drug reaction • 5–8 — Probable adverse drug reaction • 1–4 — Possible adverse drug reaction • ≤0 — Doubtful adverse drug reaction
In routine practice, questions 6 (placebo rechallenge) and 7 (toxic drug level) are almost always answered "Unknown" for spontaneous reports, since neither is ethically or practically performed outside controlled trials — meaning real-world Naranjo scores are structurally capped well below the theoretical maximum of 13 for the vast majority of ICSRs.
The Naranjo scale's enduring appeal is its transparency and reproducibility: a moderate-to-substantial interrater kappa of 0.6–0.8 has been reported across validation studies, far higher than unstructured global clinical judgment (κ often below 0.4). Because it is a fixed-weight checklist, two independent reviewers scoring the same case will usually land within one or two points of each other — a property regulators value for auditability.
Its limitations are well documented: • Weights were expert-derived, not empirically fitted to a gold-standard causality dataset, so the numeric point values are somewhat arbitrary • It performs poorly for reactions with long or variable latency (e.g., drug-induced malignancy, tardive dyskinesia) where "improved on dechallenge" is not a meaningful concept • It does not account for the strength of mechanistic/pharmacological plausibility beyond question 7 • It was validated primarily on adult general-medicine populations and translates imperfectly to pediatric, oncology, or vaccine safety contexts
Competing and complementary instruments in clinical use: • WHO-UMC categorical system — qualitative, used for regulatory case classification (see Stage 4) • French method (Bégaud, 1985) — separates chronological and semiological criteria into an intrinsic score, then layers a bibliographic extrinsic score • Liverpool ADR Causality Assessment Tool (LCAT, 2015) — a modernized decision-tree successor addressing several Naranjo weaknesses • RUCAM (Roussel Uclaf Causality Assessment Method) — the specialized scale for drug-induced liver injury, incorporating rechallenge and exclusion-of-alternative-cause criteria similar to Naranjo but weighted for hepatic enzyme kinetics
Most pharmacovigilance centers run Naranjo alongside WHO-UMC in parallel rather than choosing one — the numeric score supports triage and trending across a case series, while the WHO-UMC category is what actually appears in regulatory correspondence and periodic safety update reports (PSURs/PBRERs).
Developed by the Uppsala Monitoring Centre for the WHO Programme for International Drug Monitoring, the WHO-UMC system is the causality taxonomy behind VigiBase and the shared reference language across national pharmacovigilance centers in 170+ member countries. Unlike Naranjo's numeric score, WHO-UMC assigns each case to one of six categories based on temporal relationship, dechallenge/rechallenge, and plausibility of alternative explanations — assessed holistically rather than by additive scoring.
WHO-UMC causality assessment is a structured qualitative judgment, not a point sum, applied by an assessor weighing four dimensions: time relationship, absence of alternative explanation, dechallenge response, and rechallenge response (where available), plus pharmacological plausibility.
• Certain — event with plausible time relationship to drug intake, cannot be explained by disease or other drugs, response to withdrawal (dechallenge) is plausible, event is pharmacologically/phenomenologically definitive, and rechallenge is positive if necessary. The strictest category — genuinely rare in spontaneous-report data because it functionally requires a positive rechallenge. • Probable/Likely — reasonable time relationship, unlikely to be attributed to disease or other drugs, dechallenge response clinically reasonable, rechallenge not required for this category. • Possible — reasonable time relationship, but could also be explained by disease or other drugs; dechallenge information lacking or unclear. • Unlikely — temporal relationship makes a causal association improbable, and other drugs, chemicals, or underlying disease provide plausible explanations. • Conditional/Unclassified — event reported as an adverse reaction, but more data is essential for proper assessment, or additional data is under examination. • Unassessable/Unclassifiable — report suggesting a reaction cannot be judged because information is insufficient or contradictory, and cannot be supplemented or verified.
Because "Certain" requires a positive rechallenge that is rarely performed for ethical reasons once a serious reaction has occurred, the overwhelming majority of even well-documented, mechanistically obvious ADR reports in VigiBase settle at "Probable/Likely" — this is expected behavior of the system, not a failure of the case.
Thalidomide-induced phocomelia — the case that catalyzed the 1968 founding of the WHO Programme for International Drug Monitoring — would today be classified WHO-UMC "Probable/Likely" rather than "Certain" for any individual case, since ethical rechallenge was never performed; the "Certain" designation across the drug class only emerged from the aggregate epidemiological and mechanistic evidence base, illustrating why signal-level and case-level causality conclusions can diverge.
| Product | Indication | Trial Design | Key Result |
|---|---|---|---|
| Certain | Rechallenge positive, no alternative cause | Plausible time course, dechallenge conclusive | Strongest regulatory weight; rare (<2% of cases) |
| Probable / Likely | Reasonable time relationship | Dechallenge clinically reasonable; rechallenge not required | Modal category for well-documented ICSRs |
| Possible | Plausible but confounded by disease/co-medication | Dechallenge unclear or absent | Common; drives further data requests |
| Unlikely | Improbable temporal relationship | Alternative explanation more plausible | Typically closed without further signal action |
| Conditional / Unclassified | Reaction reported, data pending | Case flagged for supplementary follow-up | Placeholder pending additional records |
| Unassessable | Insufficient or contradictory information | Cannot be supplemented or verified | Excluded from case-series signal strength |
A single case causality assessment rarely changes a product label on its own. Signal validation aggregates causality-assessed cases into a case series, weighs it against background incidence and mechanistic plausibility using the Bradford Hill viewpoints, and routes the conclusion through formal regulatory committees — EMA's PRAC, FDA's Office of Surveillance and Epidemiology, or national authorities feeding back into WHO-UMC — toward a defined set of possible actions.
Individual WHO-UMC or Naranjo assessments answer "how likely is it that drug X caused event Y in this patient?" Signal validation asks the population-level question: "does the totality of evidence support a causal association between drug X and event Y at all?" Regulators lean on Sir Austin Bradford Hill's 1965 viewpoints, adapted for pharmacovigilance:
• Strength of association — magnitude of the ROR/PRR/EBGM signal • Consistency — replicated across independent databases (FAERS, EudraVigilance, VigiBase) and geographies • Temporality — exposure consistently precedes event across the case series • Biological plausibility — consistent with the drug's pharmacology (e.g., an anticoagulant and bleeding events) • Dose-response — higher exposure or dose correlating with event frequency or severity • Analogy — class-effect precedent from pharmacologically similar drugs • Coherence — consistent with preclinical toxicology and clinical trial safety data
No single viewpoint is necessary or sufficient; PRAC assessors and FDA safety reviewers weigh them jointly, documented in a formal signal assessment report referencing the individual case causality categories established in Stage 4 as primary supporting evidence.
Once a signal is validated, EMA GVP Module IX and equivalent FDA/Sentinel processes define a graded response ladder:
1. No action / continued monitoring — signal noted in the next Periodic Safety Update Report (PSUR/PBRER), reassessed at the following cycle 2. Additional data request — manufacturer required to submit a cumulative case review, targeted pharmacoepidemiological study, or post-authorization safety study (PASS) 3. Label update — addition to Section 4.8 (undesirable effects) of the EU Summary of Product Characteristics, or the Warnings and Precautions / Adverse Reactions sections of the US Prescribing Information 4. Direct Healthcare Professional Communication (DHPC) — an urgent, EMA-coordinated letter distributed to prescribers and pharmacists within days of a PRAC recommendation, used when the safety information must reach clinical practice before the next label revision cycle 5. Risk Minimization Measures (RMMs) — restricted indications, mandatory monitoring programs, or in the US, a formal REMS (Risk Evaluation and Mitigation Strategy) 6. Suspension or withdrawal — reserved for signals where the risk-benefit balance is judged unfavorable for any authorized use; tracked publicly via EMA's Article 107i urgent EU procedures
All EU signal validation activity is tracked through EPITT (EMA's signal-tracking database), which logs roughly 350 candidate signals entering formal validation annually across all centrally and nationally authorized products — of which historically only 15–20% progress to a label change, underscoring how much case-level causality screening exists precisely to filter statistical noise before it reaches prescribers.
The 2012 EMA-wide review of new oral anticoagulants (dabigatran, rivaroxaban) for major bleeding signals is a textbook validation pathway: EudraVigilance disproportionality flagged the signal, national centers WHO-UMC-classified several hundred cases as Probable/Likely, Bradford Hill coherence was confirmed against known pharmacology, and PRAC issued a DHPC plus label strengthening on renal-function-based dosing within one review cycle — while explicitly concluding the benefit-risk balance for the drug class remained positive.