Quality-adjusted life years, incremental cost-effectiveness ratios, and HTA decision thresholds in health economics
Every health economic evaluation begins with a single deceptively simple idea: not all years of life are equal in quality, and not all health improvements are equal in value. The Quality-Adjusted Life Year (QALY) combines length of life and quality of life into one number by weighting each year lived by a "utility" score anchored at 0 (dead) and 1 (perfect health), measured with validated preference-based instruments like the EQ-5D-5L.
A Quality-Adjusted Life Year (QALY) is the standard unit of health outcome in cost-effectiveness analysis. It is calculated as:
QALY = Σ (utility weight in period t) × (time spent in period t)
One year lived in perfect health (utility = 1.0) contributes 1.0 QALY. One year lived at a utility of 0.5 (e.g., moderate chronic pain and reduced mobility) contributes 0.5 QALY. Death contributes 0 utility, and some severe states (e.g., end-stage disease with untreated pain) can be valued below zero — "worse than dead" — under standard utility elicitation methods.
Graphically, a patient's QALYs over time are the area under a curve plotting utility (y-axis, 0–1) against time (x-axis, years). A longer life at lower quality and a shorter life at higher quality can yield the same total QALYs — this equivalence is precisely what makes QALYs useful for comparing fundamentally different interventions (a hip replacement vs a cancer drug vs a vaccination program) on one common scale.
The QALY framework was formalized in the 1970s (Torrance, Zeckhauser, Weinstein) and has since become the dominant outcome metric for health technology assessment (HTA) in over 30 countries, including the UK, Canada, Australia, and much of the EU.
The EQ-5D, developed by the EuroQol Group, is the most widely used preference-based health-related quality of life instrument in economic evaluation. The 5L (five-level) version asks patients to self-report their health across five dimensions:
• Mobility — walking about • Self-care — washing and dressing • Usual activities — work, study, housework, family, leisure • Pain / discomfort • Anxiety / depression
Each dimension is rated on 5 levels (no problems, slight, moderate, severe, extreme/unable), generating a 5-digit health-state descriptor (e.g., "21321") out of 3,125 possible combinations. This descriptor is then converted into a single utility index using a country-specific value set — derived from large general-population surveys using time trade-off (TTO) or discrete choice experiment (DCE) methods, where members of the public are asked to trade off years of life against health states to reveal their preferences.
Other instruments include the SF-6D (derived from the SF-36 quality-of-life survey), the HUI3 (Health Utilities Index), and disease-specific mapped utilities. Regulators such as NICE express a strong preference for the EQ-5D in reference-case analyses because it is generic, validated, and permits comparison across completely different disease areas.
Because a QALY gained today is generally valued more than a QALY gained in 20 years (positive time preference, plus the option value of resource flexibility), both future costs and future QALYs are discounted to present value:
PV = Σ [ value_t / (1+r)^t ]
Most HTA bodies use r = 3.5% (UK NICE) or r = 3% (US Second Panel on Cost-Effectiveness, WHO-CHOICE) per year for both costs and health outcomes. Discounting matters enormously for interventions with benefits far in the future — e.g., childhood vaccination programs or one-time curative gene therapies — because 30 years of undiscounted QALYs shrink substantially once discounted (a QALY 30 years out at 3.5% discount is worth roughly 35% of an undiscounted QALY).
Some jurisdictions apply differential discounting for very-long-horizon or curative pediatric interventions (the UK allows a lower discount rate for QALYs when effects extend beyond 30 years), reflecting an active methodological and ethical debate about how much weight society should give to future health relative to costs incurred today.
No treatment is evaluated in isolation — cost-effectiveness analysis is always comparative. A new drug's value is defined relative to the next-best alternative (usually current standard of care), by comparing two simulated (or trial-derived) health trajectories: how long each strategy keeps patients alive, and at what quality of life, integrated into total QALYs for each arm.
A cost-effectiveness model (commonly a partitioned-survival model or a Markov cohort model) simulates two arms over a defined time horizon — ideally the patient's remaining lifetime:
• Standard of care (SoC) arm: uses trial or real-world data on overall survival, progression-free survival, and health-state-specific utilities for the existing standard treatment. • New treatment arm: uses trial data (commonly from the pivotal randomized controlled trial) on the same endpoints for the new intervention, often extrapolated beyond the trial's observed follow-up using parametric survival curves (Weibull, log-normal, log-logistic, Gompertz) fitted to Kaplan-Meier data.
QALYs for each arm are calculated by multiplying time spent in each health state by that state's utility weight, summing over the horizon, and discounting. The incremental QALY gain (ΔQALY) is simply:
ΔQALY = QALY(new treatment) − QALY(standard of care)
Graphically this is the area between the two utility-over-time curves — a longer, higher curve for the new treatment vs a shorter and/or lower curve for standard of care.
Extrapolating survival curves beyond the trial's observed follow-up is one of the single largest sources of uncertainty and controversy in HTA submissions — small differences in which parametric curve is chosen can change the estimated QALY gain, and therefore the ICER, by 2× or more.
Two dominant modeling architectures are used in oncology and chronic disease HTA:
• Partitioned-survival models (PSM): directly use overall survival (OS) and progression-free survival (PFS) curves to partition the cohort at each time point into three mutually exclusive states — progression-free, progressed, and dead — without needing an explicit transition matrix. Simple, transparent, and widely used in oncology submissions, but can be criticized for not explicitly modeling transition probabilities.
• Markov (state-transition) models: define a finite set of health states (e.g., "mild," "moderate," "severe," "dead") and transition probabilities between them at each cycle. Well-suited to chronic, relapsing-remitting conditions (e.g., rheumatoid arthritis, multiple sclerosis) where patients can move back and forth between states rather than following a strictly progressive path.
Both approaches converge on the same output: total discounted QALYs and total discounted costs per arm, which feed directly into the ICER calculation. Increasingly, more granular patient-level microsimulation and discrete-event simulation models are used when individual patient heterogeneity or complex treatment-sequencing decisions matter.
Consider a new oncology drug vs standard chemotherapy:
Standard of care: mean survival 2.1 years at average utility 0.62 → 2.1 × 0.62 ≈ 1.30 QALYs (discounted lifetime total, simplified undiscounted illustration)
New treatment: mean survival 3.4 years at average utility 0.68 (better tolerability, less time in the "progressed" health state) → 3.4 × 0.68 ≈ 2.31 QALYs
Incremental QALY gain: ΔQALY = 2.31 − 1.30 ≈ 1.01 QALYs
This single number — roughly one additional year of perfect-health-equivalent life — becomes the denominator of the ICER in the next stage. Note that the QALY gain blends two distinct effects: extra survival time (1.3 extra years) and improved quality during that time (utility 0.68 vs 0.62) — separating these two contributions is often reported explicitly in HTA submissions as "life-years gained" and "QALYs gained" side by side.
The ICER is the single number that health technology assessment bodies scrutinize most: the additional cost required to gain one additional QALY, compared against a willingness-to-pay (WTP) threshold representing the value society (or a payer) places on health. Plotting the incremental cost and incremental QALY on the cost-effectiveness plane makes the accept/reject decision visually explicit.
The Incremental Cost-Effectiveness Ratio is defined as:
ICER = ΔCost / ΔQALY = [ Cost(new) − Cost(comparator) ] / [ QALY(new) − QALY(comparator) ]
It answers: "how many additional dollars must be spent to gain one additional QALY by switching from the comparator to the new intervention?" The cost-effectiveness plane plots ΔQALY on the x-axis and ΔCost on the y-axis, dividing the space into four quadrants around the origin (the comparator's own position):
• NE quadrant (more costly, more effective): the typical location for most new drugs — costs more, but also works better. Whether this is accepted depends on where the point falls relative to the WTP threshold line. • SE quadrant (less costly, more effective): "dominant" — strictly better and cheaper. Automatic accept. • NW quadrant (more costly, less effective): "dominated" — strictly worse. Automatic reject. • SW quadrant (less costly, less effective): a trade-off requiring judgment — cheaper but worse outcomes; rarely how new branded drugs are priced, more relevant to de-adoption or generic substitution decisions.
The WTP threshold is drawn as a straight diagonal line through the origin with slope equal to the threshold value (e.g., $50,000/QALY). Points falling below/right of this line (lower cost per QALY than the threshold) are judged cost-effective; points above/left are not.
A point can have a favorable ICER purely by chance — high uncertainty near the WTP line is why probabilistic sensitivity analysis and the CEAC (next stage) are required alongside the single ICER point estimate.
There is no single universal WTP threshold — different health systems have adopted very different explicit or implicit values:
• UK NICE: explicit threshold range of £20,000–£30,000 per QALY for most technologies; up to ~£50,000/QALY for end-of-life criteria and specialized "highly specialized technologies" (ultra-rare diseases can be evaluated with much higher implied thresholds, sometimes £100,000+/QALY, reflecting severity-of-disease modifiers introduced in 2022). • United States: no single official government threshold (Medicare is statutorily barred from using cost-effectiveness as the sole criterion for coverage), but the independent, non-governmental Institute for Clinical and Economic Review (ICER US, unrelated in name only to the ICER metric) commonly references $100,000–$150,000 per QALY as a benchmark range in its value assessments, influencing payer negotiations even without regulatory force. • Canada (CADTH): typically references $50,000 CAD per QALY as an informal benchmark, with flexibility case by case. • Germany (G-BA/IQWiG): does not use a fixed QALY threshold at all — instead uses "efficiency frontier" methodology comparing added benefit and added cost against existing therapies in the same indication. • WHO historical guideline (largely superseded): 1–3× national GDP per capita per QALY/DALY averted — heavily criticized as arbitrary and abandoned in WHO's own 2014 guidance in favor of opportunity-cost-based, budget-impact-informed thresholds.
• Zolgensma (onasemnogene abeparvovec), a one-time gene therapy for spinal muscular atrophy, launched at ~$2.1 million per dose. ICER (US) analyses estimated the therapy could be cost-effective at a $150,000/QALY threshold given the severity of untreated SMA and lifetime QALY gains for infants, illustrating how a huge sticker price can still fall within a WTP threshold when the health gain (potentially decades of quality life vs early death) is large enough — while critics argued the analysis relied on highly uncertain long-term durability assumptions.
• Sovaldi (sofosbuvir) for hepatitis C launched around $84,000 for a 12-week curative course (~$1,000/pill), triggering major public and payer controversy over near-term budget impact — despite ICER analyses generally finding it cost-effective per QALY (curing HCV avoids cirrhosis, liver transplant, and hepatocellular carcinoma costs), the sheer number of eligible patients created an affordability crisis distinct from cost-effectiveness.
• Many oncology drugs, especially in later-line, heavily pretreated indications with modest survival gains (weeks to a few months), routinely produce ICERs exceeding $150,000/QALY and sometimes $300,000+/QALY, making them frequent subjects of NICE rejections or restricted "Cancer Drugs Fund" conditional approvals in the UK pending further real-world evidence.
A single ICER point estimate hides enormous underlying uncertainty — in survival extrapolation, utility values, drug acquisition cost, and resource use. Probabilistic sensitivity analysis (PSA) propagates this uncertainty through Monte Carlo simulation, and the resulting cloud of simulated ICERs is summarized in a Cost-Effectiveness Acceptability Curve (CEAC): the probability the intervention is cost-effective as a function of the WTP threshold.
Cost-effectiveness models rest on dozens of uncertain inputs: hazard ratios from a single clinical trial (itself a sample estimate with confidence intervals), utility values elicited from small patient samples, drug acquisition costs that vary by contract, and long-term survival extrapolations beyond observed data. Two complementary approaches quantify how this uncertainty affects the conclusion:
• Deterministic (one-way / tornado) sensitivity analysis: vary one parameter at a time across its plausible range (e.g., 95% CI) while holding all others at their base-case value, and observe how the ICER changes. Results are typically displayed as a tornado diagram ranking parameters by their influence on the ICER.
• Probabilistic sensitivity analysis (PSA): assign a probability distribution to every uncertain parameter simultaneously (not just one at a time), then run Monte Carlo simulation — typically 1,000 to 10,000 iterations — drawing a random value for every parameter from its distribution in each iteration, recalculating costs and QALYs for both arms, and recording the resulting simulated ICER. This produces a scatter cloud of thousands of possible ICER outcomes reflecting joint parameter uncertainty, which is what most modern HTA reference cases (NICE, ICER US) require as standard practice.
NICE explicitly requires PSA in reference-case submissions specifically because it captures correlated uncertainty across many parameters simultaneously — something a one-way tornado analysis, however useful for identifying key drivers, cannot represent.
Distribution choice matters — each PSA parameter uses a distribution matching its statistical properties:
• Utility values (bounded 0–1): Beta distribution, parameterized from the mean and standard error observed in the utility elicitation study. • Costs (positive, right-skewed): Gamma or log-normal distribution, since costs cannot be negative and often have a long right tail (a few patients incur very high resource use). • Relative treatment effects (hazard ratios, odds ratios, always positive): log-normal distribution on the log scale, derived from the trial's reported confidence interval. • Transition probabilities in multi-state Markov models: Dirichlet distribution, which correctly constrains a full row of transition probabilities to sum to 1 across all destination states simultaneously.
Each of the (say) 10,000 Monte Carlo iterations draws one random value from every distribution, runs the full health-economic model with that specific parameter set, and stores the resulting (ΔQALY, ΔCost) pair. The result is a scatter of 10,000 points on the cost-effectiveness plane — visually, a cloud centered near the base-case ICER but with a spread reflecting how sensitive the model is to input uncertainty.
The CEAC converts the PSA scatter cloud into a decision-relevant summary. For any candidate WTP threshold λ, each simulated iteration is classified as cost-effective if:
λ × ΔQALY(i) − ΔCost(i) > 0 (equivalently, Net Monetary Benefit NMB(i) = λ·ΔQALY(i) − ΔCost(i) > 0)
The proportion of the 10,000 iterations satisfying this condition at a given λ is the probability the intervention is cost-effective at that threshold. Repeating this calculation across a full range of λ values (e.g., $0 to $200,000/QALY) traces the CEAC — typically an S-shaped curve rising from 0% (at very low WTP) toward 100% (at very high WTP), crossing 50% near the base-case ICER.
A steep CEAC indicates low decision uncertainty (the conclusion is robust across a wide range of plausible thresholds); a shallow, gradually rising CEAC indicates the accept/reject decision is highly sensitive to exactly which threshold is used — precisely the situation where committees like NICE's appraisal committee spend the most deliberation time, and where Value of Information (VOI) analysis may be commissioned to decide whether further research (e.g., a confirmatory trial) is worth funding before a final reimbursement decision.
The ICER and CEAC are inputs to a real institutional decision: will a health system pay for this treatment, and at what price? Health technology assessment (HTA) bodies around the world formalize this judgment, weighing cost-effectiveness evidence against disease severity, unmet need, budget impact, and — increasingly — negotiating price directly with manufacturers to bring the ICER within an acceptable range.
• NICE (National Institute for Health and Care Excellence, England/Wales): the most influential and methodologically rigorous HTA body globally. Uses an explicit £20,000–£30,000/QALY reference range, a severity modifier (extra QALY weighting for more severe conditions, introduced 2022), and special pathways for end-of-life care and highly specialized technologies (ultra-rare diseases). A "yes/no" appraisal committee decision directly determines NHS funding and, in effect, sets a de facto UK price ceiling.
• ICER (Institute for Clinical and Economic Review, United States): an independent non-profit, not a government body — the US has no binding national cost-effectiveness threshold. ICER publishes value assessments referencing a $100,000–$150,000/QALY range; its reports influence payer formulary decisions and price negotiations even though adoption is voluntary, and are increasingly cited by the Centers for Medicare & Medicaid Services (CMS) in the new Medicare Drug Price Negotiation Program created by the Inflation Reduction Act.
• CADTH (Canada's Drug Agency, formerly Canadian Agency for Drugs and Technologies in Health): reviews drugs for public formulary listing across Canadian provinces, generally referencing ~$50,000 CAD/QALY informally, with pan-Canadian Pharmaceutical Alliance (pCPA) price negotiations following a positive recommendation.
• G-BA / IQWiG (Germany): unique "early benefit assessment" (AMNOG) process — does not use a QALY threshold at all, instead classifies "added benefit" (major, considerable, minor, none, or negative) versus an appropriate comparator, which then anchors price negotiation between the manufacturer and the statutory health insurance umbrella association (GKV-Spitzenverband).
When an ICER exceeds a jurisdiction's threshold, the outcome is rarely an outright rejection alone — manufacturers frequently respond with confidential discounts, patient access schemes, or outcomes-based managed-entry agreements that lower the effective net price until the ICER falls within an acceptable range, without changing the publicly listed price.
A favorable ICER answers "is this a good use of health resources per patient treated?" — but it does not answer "can the health system afford to pay for all eligible patients this year?" Budget impact analysis (BIA) answers the second, complementary question by projecting total incremental spending over a defined horizon (typically 1–5 years) as:
Budget Impact = (Eligible population × uptake rate × net price) in the "new treatment" world minus the equivalent spending in the "current standard of care" world
Sovaldi/hepatitis C is the canonical case: individually cost-effective (avoiding cirrhosis and transplant costs over a lifetime) yet a major near-term budget shock because millions of patients became simultaneously eligible for a curative but expensive therapy, forcing payers to ration access via prior-authorization criteria despite a favorable long-run ICER. Many HTA frameworks now formally incorporate an affordability/budget-impact threshold alongside the cost-effectiveness threshold, and some (e.g., value-based pricing frameworks in several EU states) explicitly discount the acceptable price further as the eligible population grows.
Value-based pricing (VBP) sets a drug's price such that its ICER, at that price, meets the payer's WTP threshold — working the ICER formula backward: solving ΔCost such that ΔCost/ΔQALY = λ_threshold gives the maximum acceptable price. In practice, negotiations blend this economic logic with strategic factors (competitor pricing, international reference pricing, political visibility of the disease area, and manufacturer R&D cost recovery needs).
Managed-entry agreements (MEAs) are increasingly used for high-uncertainty, high-cost therapies (especially one-time gene and cell therapies like Zolgensma, Luxturna, and CAR-T products):
• Outcomes-based agreements: payment is contingent on the patient achieving a pre-specified clinical outcome (e.g., staying alive and free of a specific complication at 2 years); if the outcome is not met, the manufacturer refunds part or all of the payment. • Annuity/installment payments: a $2 million one-time gene therapy is paid over 5 years instead of upfront, reducing immediate budget shock and allowing partial clawback if durability of effect is not sustained. • Price-volume agreements: price per unit decreases as cumulative volume/spend crosses pre-agreed thresholds, capping total budget exposure.
These mechanisms let payers accept therapies with a favorable point-estimate ICER but genuinely high uncertainty about long-term durability — directly addressing the decision uncertainty the CEAC makes visible in the previous stage — without simply rejecting the technology outright.
| Product | Indication | Trial Design | Key Result |
|---|---|---|---|
| NICE (England/Wales) | £20,000–£30,000/QALY (up to ~£50k+ end-of-life, severity-weighted) | Single/multiple technology appraisal, severity modifier since 2022 | Binding NHS funding decision |
| ICER (United States) | $100,000–$150,000/QALY (informal, non-binding) | Independent value assessment reports, no statutory authority | Influences payer & PBM negotiation |
| CADTH (Canada) | ~$50,000 CAD/QALY (informal reference) | Reimbursement recommendation, then pCPA price negotiation | Pan-provincial price alignment |
| G-BA / IQWiG (Germany) | No fixed QALY threshold — efficiency frontier method | AMNOG early benefit assessment vs appropriate comparator | Direct price negotiation post-launch |