Benchmarking industry-wide clinical phase transition probabilities by therapeutic area to forecast a program's likelihood of approval
Every probability-of-success (PoS) benchmark ultimately rests on one thing: a large, carefully curated database of what actually happened to thousands of real drug-development programs. Before any transition rate can be calculated, programs from hundreds of sponsors — large pharma, biotech, academic spin-outs — must be identified, classified, and tracked from first-in-human dosing through to regulatory outcome (or abandonment).
Well-known industry PoS benchmarking exercises (in the spirit of BIO/Amplion/QLS-style clinical development success-rate studies) pull program-level records primarily from clinical trial registries (ClinicalTrials.gov and equivalents), company disclosures, regulatory databases, and commercial pipeline-tracking services. Each program is tagged with its lead indication, therapeutic area, modality (small molecule, biologic, cell/gene therapy), and every phase transition it passed through or failed at.
Building the database is mostly definitional plumbing rather than statistics: a "program" has to be defined consistently (one investigational compound in one indication, not one NCT number), phase entry has to be dated to when a program first entered that phase rather than when it was announced, and outcomes have to be classified into a small number of mutually exclusive buckets — advanced, discontinued, or still active/censored at the time of the analysis cut.
The single most common methodological error in home-grown PoS estimates is survivorship bias — counting only programs that eventually became visible successes (because they were licensed, publicized, or IPO'd on) while silently dropping the much larger number of programs that were quietly discontinued. A rigorous database has to capture failures at least as diligently as successes.
Real-world development is messy in ways that complicate clean aggregation. The same molecule is often tested in multiple indications simultaneously, each with its own independent phase-transition outcome — a kinase inhibitor might fail in one solid tumor type while succeeding in a hematologic malignancy. A rigorous database therefore tracks compound-indication pairs, not compounds alone, and de-duplicates trials so that a single program is not double-counted across its Phase 2a, Phase 2b, and pivotal Phase 2/3 studies.
Censoring is the other central issue: at the moment a snapshot is taken, some fraction of programs are still actively enrolling or awaiting a regulatory decision — their ultimate fate is unknown. These programs cannot simply be excluded (that reintroduces survivorship bias) or counted as failures (that understates PoS); they must be handled with time-to-event methods that treat "still in progress" as right-censored observation, similar to how clinical survival analysis handles patients who are alive at last follow-up.
A 20-year historical window maximizes sample size and lets stratified therapeutic-area cuts remain statistically meaningful, but it also blends together very different regulatory and scientific eras — pre- and post-biomarker-driven oncology trial design, the rise of accelerated approval pathways, and shifting FDA/EMA evidentiary standards are all folded into one pooled number.
A 5-year recent window better reflects the current regulatory and scientific environment (companion diagnostics, adaptive trial designs, more precise patient selection) but trades away sample size, making the resulting phase-specific probabilities noisier, especially for less common indications. Neither vintage choice is objectively "correct" — it is a bias-variance tradeoff that any credible benchmarking exercise has to disclose explicitly rather than silently pick one and present it as ground truth.
With a clean database in hand, the next step converts raw program outcomes into the four numbers that matter most: the empirical probability of advancing from Phase 1 to Phase 2, Phase 2 to Phase 3, Phase 3 to Filing, and Filing to Approval. Each is a simple ratio in principle — advances divided by attempts — but getting that ratio right requires care about what counts as an "attempt" and what counts as an "advance."
For each gate, the empirical transition probability is:
P(phase X → phase X+1) = (programs that entered phase X+1) ÷ (programs that entered phase X)
Applied sequentially across all four gates, this produces a full phase-by-phase "success curve" for the pooled population. The overall Likelihood of Approval (LOA) from Phase 1 is simply the product of all four conditional probabilities — because a program has to clear every gate in sequence, the joint probability of clearing all four is the multiplication of each conditional probability, not their average.
This is the single most important arithmetic fact in the whole discipline, and it explains why LOA from Phase 1 is always dramatically lower than any individual gate probability: even four gates each around 60% multiply down to roughly 13%, and real gate probabilities are usually more uneven than that, pulling pooled LOA further down.
Because probabilities compound multiplicatively, small changes in the weakest gate move overall LOA far more than equivalent changes in a strong gate. A program's single biggest lever for improving its probability-weighted valuation is almost always the phase with the lowest historical transition rate — for most therapeutic areas, that is Phase 2 to Phase 3.
The four gates are not equally risky, and the pattern is remarkably consistent across therapeutic areas: transition probability generally rises as programs progress. Phase 1 to Phase 2 mainly screens out programs with unacceptable safety/tolerability or clearly wrong pharmacokinetics — a relatively coarse filter. Phase 2 to Phase 3 is where efficacy is tested for the first time in a reasonably powered way, and it is where the majority of "the drug simply doesn't work as hypothesized" failures occur — this is why it is consistently the lowest transition rate of the four.
By the time a program reaches Phase 3 to Filing, it has already survived two successive efficacy and safety filters, so the population attempting Phase 3 is heavily enriched for genuinely promising candidates — a form of selection that mechanically raises the transition rate even before considering that pivotal trials are typically better designed based on earlier-phase learnings. Filing to Approval is the least risky gate of all: by definition, a sponsor only files when internal data already strongly supports approval, so this gate mostly screens out administrative and manufacturing-quality issues rather than efficacy failures.
It is essential to distinguish conditional transition probabilities (the four gate-specific rates described above, each conditioned on having reached that phase) from unconditional probabilities (the chance, from Phase 1, of ever reaching a later phase). A Phase 3 to Filing rate of 62% is a conditional probability — it says nothing about how many Phase 1 programs ever make it that far.
Mixing the two up is a frequent source of error in internal forecasting decks: a business-development analyst might see "Phase 3 success rate: 62%" and apply it directly to a Phase 1 asset's valuation, effectively assuming the asset has already survived Phase 1 and Phase 2 for free. The correct approach for a Phase 1 asset is always to chain-multiply from wherever the asset currently sits through every remaining gate to Approval — never to apply a single downstream conditional rate in isolation.
A single pooled LOA number — computed by averaging across every therapeutic area — is almost never the right number to apply to any specific program. Structural differences in disease biology, endpoint validation, regulatory pathway maturity, and patient population homogeneity drive transition rates that vary by a factor of two or more between the best- and worst-performing therapeutic areas.
Oncology consistently shows among the lowest overall LOA of any major therapeutic area in illustrative industry benchmarking exercises, for reasons rooted in the biology and trial design of cancer drug development rather than any deficiency in the science itself. Tumor biology is enormously heterogeneous — even within a single histological tumor type, genomic subtypes respond very differently to the same mechanism, so a drug that shows a strong signal in an early, less-selected population frequently dilutes in a larger, more heterogeneous pivotal trial.
Oncology pivotal trials also frequently rely on surrogate endpoints (progression-free survival, objective response rate) that regulators and payers increasingly want confirmed by mature overall-survival data, adding a layer of downstream risk even after a positive pivotal readout. Combination-regimen requirements — where a new agent must show benefit on top of an evolving standard-of-care backbone that itself keeps improving — raise the bar continuously, and the sheer number of competing oncology programs racing for the same biomarker-defined patient populations intensifies both scientific and enrollment risk.
Infectious disease and hematology programs illustratively show meaningfully higher transition rates at nearly every gate. Infectious disease benefits from well-validated, often objective endpoints (microbiological clearance, viral load reduction) that leave less room for the endpoint-definition disputes that plague other areas, plus a regulatory environment shaped by strong public-health incentives (QIDP designation, priority review vouchers) that reward well-run programs with faster, more predictable pathways.
Hematology and many rare/orphan diseases benefit from a related dynamic: smaller, more genetically homogeneous patient populations, mechanistically well-understood single-gene or single-pathway biology, and — critically — orphan drug regulatory frameworks that permit smaller pivotal trials, single-arm designs against historical controls, and surrogate-endpoint-based accelerated approval more readily than most other therapeutic areas. All three of these factors mechanically raise the recorded historical transition rate.
None of this means oncology drugs are scientifically worse bets than infectious-disease drugs — it means the average oncology *trial* carries structurally higher biological and statistical risk than the average infectious-disease trial. Applying an unstratified pooled LOA to an oncology asset systematically overstates its probability of success; applying it to a rare-disease asset systematically understates it.
Because portfolio composition varies so much across sponsors and eras, a pooled "industry average" LOA is highly sensitive to which therapeutic areas happen to dominate the sample in a given study window. A benchmarking dataset weighted heavily toward oncology (which represents a large share of all clinical-stage programs industry-wide) will report a lower pooled LOA than one balanced evenly across therapeutic areas — not because drug development generally got harder, but purely because of composition.
This is precisely why credible PoS benchmarking always reports therapeutic-area-stratified cuts alongside (never instead of) the pooled headline number, and why any team applying an external benchmark to an internal asset should match on therapeutic area, and where data allows, on modality and target class as well, rather than reaching for the single easiest-to-remember pooled figure.
| Product | Indication | Trial Design | Key Result |
|---|---|---|---|
| Oncology | |||
| Infectious Disease | |||
| Cardiovascular | |||
| Neurology | |||
| Rare Disease |
External, therapeutic-area-stratified benchmarks are most valuable not as a replacement for internal judgment, but as a discipline-enforcing reference point: a specific program's internally assumed PoS — the number a project team plugs into its own forecasts — can and should be compared directly against the relevant industry benchmark to surface whether the internal view is optimistic, conservative, or well-calibrated.
Project teams are, almost by construction, systematically optimistic about their own asset's probability of success. The scientists and clinicians closest to a program have typically already seen encouraging early signals that motivated continued investment, they are professionally and often financially invested in the program's success, and they are naturally more attuned to reasons their specific mechanism might beat the historical average than to the base-rate reasons it might not.
This is not a criticism of individual judgment so much as a well-documented structural bias in forecasting under uncertainty — the same phenomenon shows up in software project timelines, sales pipeline forecasts, and venture capital return projections. The corrective is not to distrust internal teams, but to systematically anchor every internal PoS assumption to an external, empirically grounded reference class before it enters a valuation model or a portfolio prioritization decision.
A useful governance rule of thumb: any internal PoS assumption that sits meaningfully above the relevant therapeutic-area benchmark should require an explicit, written justification tied to a specific, verifiable differentiator (validated biomarker-selected population, best-in-class potency/selectivity data, a novel mechanism with strong genetic human-validation) — not general optimism about the team or the science.
Company-specific benchmarking is mechanically simple once the reference class is chosen correctly: identify the program's current phase and therapeutic area (and modality, where the benchmark allows that level of granularity), pull the corresponding external gate-by-gate transition probabilities, and plot the program's internally assumed probabilities against that curve, gate by gate rather than only at the aggregate LOA level.
A program that assumes, say, a 55% Phase 2 to Phase 3 transition rate in an oncology indication where the therapeutic-area benchmark sits closer to 28% is not automatically wrong — but the burden of proof shifts decisively onto the team to explain why this specific program should be expected to outperform the historical base rate by roughly two-fold, rather than the benchmark team having to prove the program is unremarkable.
Company-specific benchmarking against external PoS data shows up at several recurring decision points across a development organization: portfolio Go/No-Go committees use it to sanity-check project-team forecasts before committing further capital to a phase transition; business development and licensing teams use it as a first-pass discipline check on a target company's own PoS claims during in-licensing or M&A diligence, precisely the kind of benchmarked assumption that later feeds a risk-adjusted net present value (rNPV) model; and portfolio-level capital allocation exercises use therapeutic-area-stratified benchmarks to compare risk-adjusted return across a diversified pipeline on a consistent, externally anchored basis rather than each project team's self-reported optimism.
Once transition probabilities have been benchmarked, stratified, and sanity-checked against a specific program's own assumptions, the final step is mechanical but consequential: chain-multiplying the four gate probabilities together into a single forward-looking Likelihood of Approval, and feeding that number directly into risk-adjusted valuation models elsewhere in the portfolio.
For a hypothetical new program entering Phase 1 in a given therapeutic area, the forward-looking Likelihood of Approval is:
LOA = P(1→2) × P(2→3) × P(3→Filing) × P(Filing→Approval)
Each factor is pulled from the appropriate therapeutic-area-stratified benchmark (adjusted, where justified, for company-specific differentiators established in the benchmarking step). Because the terms multiply rather than average, LOA is highly sensitive to the weakest link in the chain — a program with three strong gates and one weak gate will have an LOA much closer to what the weak gate alone would suggest than a simple average of the four gates would imply.
As a program advances and clears each gate in reality, LOA is recalculated using only the remaining downstream gates — a Phase 2 program's LOA no longer includes the Phase 1 to Phase 2 term at all, since that risk has already resolved. This is why LOA should be understood as a moving, risk-resolving quantity rather than a single fixed number assigned once at inception.
This same chain-multiplied LOA is the direct input to risk-adjusted net present value (rNPV) modeling used elsewhere in portfolio and transaction valuation work — every dollar of projected future cash flow from an unapproved asset is scaled by this probability before being discounted back to present value, which is why getting the benchmarking right upstream has an outsized effect on downstream valuation.
A benchmarked LOA describes the average historical outcome for programs that resemble this one along the dimensions the benchmark can observe — phase, therapeutic area, sometimes modality. It cannot, by construction, capture everything that makes a specific asset unusually strong or unusually weak: a validated genetic-target rationale in humans, an unusually clean early safety profile, a companion diagnostic that meaningfully de-risks patient selection, or conversely a crowded competitive landscape or a marginal early efficacy signal.
The correct mental model is that a therapeutic-area-stratified benchmark provides a well-anchored prior, not a final answer — company-specific evidence should be allowed to move the estimate away from the benchmark, but only with the kind of explicit, falsifiable justification described in Stage 4, and ideally documented so the assumption can be revisited if new data arrives.
Forward PoS benchmarking is rarely an end in itself — it is almost always an input to a larger valuation or capital-allocation exercise. In rNPV models used for internal portfolio prioritization, licensing deal structuring, or M&A target valuation, the chain-multiplied LOA at each future phase transition scales the probability-weighted cash flows that ultimately determine an asset's risk-adjusted value; the same benchmarking discipline described across this page underlies target-valuation exercises like risk-adjusted M&A valuation models, where a mismatch between a target's self-reported PoS assumptions and industry-benchmarked reality is one of the most common sources of overpayment in biotech dealmaking.
Because of this direct linkage, disciplined PoS benchmarking is not merely a scientific exercise — it is a capital-allocation control, and getting the therapeutic-area stratification and company-specific adjustment steps right has direct, quantifiable consequences for how much an organization should be willing to pay for, or invest in, a given clinical-stage asset.