A perpetual master protocol that adds and retires investigational arms over time against one shared, continuously evolving control
A perpetual platform trial inverts the classical model of clinical research. Instead of writing a new protocol, assembling a new site network, and negotiating a new statistical analysis plan for every drug, sponsors and academic consortia build one durable master protocol — one IRB/ethics approval, one master informed consent template, one data infrastructure, one Bayesian statistical engine — and then plug investigational arms into it indefinitely, the way apps plug into an operating system.
A platform trial's master protocol is structurally different from a conventional randomized controlled trial in five ways:
• Single regulatory shell: one Investigational New Drug (IND) umbrella or one national CTA can host multiple sponsor-owned investigational products, each governed by a product-specific appendix rather than a freestanding protocol • Shared infrastructure: one central IRB/EC (or reliance agreement network), one electronic data capture build on a CDISC SDTM-compliant backbone, one core eligibility screening funnel, one biostatistics and data-management core • Shared, continuously refreshed control arm: patients not currently eligible for or randomized to an experimental arm receive standard-of-care control, pooled longitudinally across the life of the platform rather than re-recruited for each new comparison • Domain-based factorial structure: platforms such as REMAP-CAP randomize patients simultaneously across multiple treatment domains (antivirals, corticosteroids, immune modulators), each domain evaluated with its own adaptive rule • No pre-specified closing date: the protocol is amended, not replaced, when arms are added or retired — the master protocol document itself can accumulate 20–40 amendments over a decade of operation
Governance sits above any single drug: a platform steering committee (sponsor representatives, methodologists, patient advocates, regulators as observers) approves each new arm's entry criteria, sample size, and stopping rules before activation — decoupling scientific and operational review of the chassis from the review of any individual molecule.
STAMPEDE (Systemic Therapy in Advancing or Metastatic Prostate cancer: Evaluation of Drug Efficacy), launched by the UK Medical Research Council in 2005, has enrolled more than 11,000 men across 10 research arms under one continuously running multi-arm multi-stage (MAMS) protocol, directly changing standard of care four separate times (docetaxel, abiraterone, zoledronic acid, abiraterone+enzalutamide) without ever closing and reopening as a new trial.
Regulators have converged on formal frameworks for perpetual platforms over the past decade:
• FDA "Master Protocols: Efficient Clinical Trial Design Strategies to Expedite Development of Oncology Drugs and Biologics" (final guidance, 2022) codifies umbrella, basket, and platform trial subtypes and expectations for arm-level type-I error control • FDA "Adaptive Designs for Clinical Trials of Drugs and Biologics" (2019) sets expectations for pre-specification, simulation reports, and firewalled access to unblinded interim data • FDA Complex Innovative Trial Design (CID) Meeting Program, funded under PDUFA VI/VII, gives sponsors a formal venue to pressure-test a platform's Bayesian design and simulated operating characteristics with the review division before the first patient is enrolled • EMA "Reflection paper on the use of extrapolation and complex/innovative clinical trial designs" (EMA/CHMP, 2020) and subsequent qualification opinions extend similar principles across the EU network • ICH E20 (adaptive designs for clinical trials), under development since 2022, aims to harmonize simulation-report expectations and interim-analysis firewall standards across ICH regions • ICH E6(R3) Good Clinical Practice (2025 revision) explicitly accommodates protocol appendix structures and risk-based quality management suited to continuously evolving master protocols
The throughline across all of these documents is the same: regulators will accept a perpetual, self-modifying protocol only if the statistical design that governs arm entry, randomization, and exit is fully pre-specified and simulated in advance — improvisation mid-trial is exactly what these frameworks are built to prevent.
The defining operational advantage of a platform trial is speed of arm activation. Because the legal, ethical, and statistical scaffolding already exists, a new investigational arm can begin randomizing patients in weeks rather than the 12–18 months typical of standing up a de novo Phase 2/3 trial — the sponsor negotiates a substudy appendix, not a new protocol from a blank page.
Activating a new arm follows a standardized onboarding checklist rather than a bespoke protocol-writing exercise:
1. Sponsor proposes candidate regimen and biomarker-defined eligibility to the platform steering committee, along with a target effect size and prior distribution for the Bayesian model 2. Independent statistical simulation of operating characteristics: false-positive rate contribution to overall platform type-I error, expected sample size under the null and alternative, expected time to graduation or futility 3. Substudy appendix drafted against the master protocol template — inheriting eligibility screening, safety reporting (MedDRA-coded adverse event capture, expedited reporting per ICH E2B(R3)), and data standards (CDISC SDTM/ADaM) verbatim 4. Single IRB/EC reviews only the appendix's novel content — investigational product, dosing, arm-specific stopping rules — not the entire master protocol machinery again 5. Drug supply, pharmacy manual, and site activation proceed against a pre-qualified site network already trained on the master protocol's core procedures
This modular onboarding is why Lung-MAP (Lung Cancer Master Protocol), a SWOG/Friends of Cancer Research/FDA collaboration for squamous and non-squamous non-small-cell lung cancer, has cycled through more than ten biomarker-matched sub-protocols since 2014 using a single shared genomic screening infrastructure (a common next-generation sequencing panel run once per patient, routing each patient to whichever open sub-study matches their mutation).
A new arm entering an established platform faces a subtle statistical hazard: should it be compared only against patients randomized to control during the same calendar window (concurrent control), or against the full accumulated control pool including patients enrolled years earlier (non-concurrent, or "borrowed," control)?
Standard-of-care drifts over the life of a decade-long platform — supportive care improves, diagnostic staging becomes more sensitive, competing approved therapies enter practice. Pooling non-concurrent controls naively risks confounding calendar time with treatment effect. The accepted mitigation strategies are:
• Concurrency requirement: primary efficacy analysis restricted to control patients randomized during the arm's active enrollment window; non-concurrent controls used only in sensitivity or exploratory Bayesian borrowing models • Time-trend adjustment models: regression or hierarchical Bayesian models that explicitly estimate a secular drift term and adjust the treatment-control contrast for it • Meta-analytic-predictive (MAP) priors: non-concurrent control data are down-weighted via a discount factor calibrated to the estimated heterogeneity between calendar eras, formalized in FDA and EMA guidance on external/historical control borrowing
Nearly every regulatory rejection of a platform-trial submission traces back to inadequate handling of this concurrency question — it is the single most scrutinized statistical design choice in a platform's pre-specification package.
The FDA's 2022 master protocol guidance explicitly states that the primary analysis of a new arm should generally use only concurrently randomized controls, reserving non-concurrent, platform-wide control pooling for supportive or Bayesian-borrowing sensitivity analyses — a direct response to time-trend confounding concerns raised during review of early oncology platform submissions.
Once multiple arms are running concurrently, a platform trial can do something a fixed-ratio RCT cannot: let the randomization ratio itself evolve. At each scheduled interim look, the Bayesian model recomputes the posterior probability that each arm is superior to control, and the allocation ratio shifts — probabilistically, not deterministically — toward arms that are performing well, reducing the number of patients assigned to arms that are underperforming.
The canonical RAR rule, used in I-SPY2 and its descendants, follows a "play-the-winner" logic implemented through posterior sampling rather than simple win-rate tracking:
1. Prior specification: each arm-control contrast is assigned a prior distribution over the treatment effect (often a weakly informative Beta or normal prior on the log-odds ratio of the primary endpoint, e.g., pathologic complete response) 2. Posterior update: as outcomes accrue, the model computes the posterior probability that each experimental arm is superior to control, typically via Markov Chain Monte Carlo (MCMC) sampling of a hierarchical Bayesian logistic or time-to-event model 3. Allocation ratio update: the randomization probability for arm k is set proportional to a power function of its posterior probability of being the best arm, √(p_k), which dampens the swing so that a lucky early streak does not immediately capture the majority of remaining patients 4. Allocation floor: every active arm retains a minimum randomization probability (commonly ~10%) so that no arm is starved of accrual before its stopping rule can trigger, preserving estimability of the treatment effect 5. Burn-in period: the first 60–100 patients per arm are typically randomized at a fixed 1:1 (or 1:1:1…) ratio before adaptive allocation switches on, preventing early noise from distorting the ratio
The practical effect: patients enrolled later in a platform's life are systematically more likely to receive whichever regimen the accumulating evidence favors — an ethical argument frequently cited in favor of RAR designs, since fewer patients are randomized to arms already trending toward futility.
RAR is not statistically free. The methodological literature identifies three recurring costs that platform statisticians must budget for during pre-trial simulation:
• Variance inflation: because the allocation ratio departs from the variance-optimal fixed ratio, the standard error of the final treatment effect estimate can be 5–15% larger than under a matched fixed-ratio design of the same total sample size — RAR often needs modestly larger N to preserve power • Time-trend confounding: if unmeasured secular changes correlate with the changing allocation ratio, treatment-effect estimates can be biased; this compounds the concurrent-control problem from Stage 2 and is why time-trend covariates are built into the hierarchical model, not bolted on afterward • Operational bias: unblinded staff or investigators who can infer that a shifting allocation ratio signals which arm is winning may unconsciously alter enrollment or eligibility assessment behavior; strict firewalling of interim results to an independent statistical center (never shared with site investigators) is a mandatory design control
Despite these costs, simulation studies consistently show that under moderate-to-strong true treatment effects, RAR platforms randomize 15–30% fewer patients to the inferior arm compared to fixed 1:1 randomization while reaching a graduation decision in a similar or shorter overall timeline — the central efficiency argument for adopting Bayesian RAR in disease areas with high patient burden per trial arm, such as glioblastoma (GBM AGILE) and early triple-negative breast cancer (I-SPY2).
I-SPY2's Bayesian RAR engine uses biomarker signatures (HER2, hormone-receptor, MammaPrint risk) to run parallel adaptive randomizations within molecular subtypes simultaneously — a single patient's allocation probability across candidate arms is computed from her own signature profile, not a single platform-wide ratio, allowing more than ten regimens to be evaluated against a shared control within the same running infrastructure since 2010.
Every active arm in a perpetual platform is reviewed at scheduled interim looks by an independent Data Safety Monitoring Board (DSMB, sometimes iDMC) that sees unblinded Bayesian posterior and predictive-probability output the platform steering committee and sponsors do not. Each arm carries two live decision boundaries at every look: a graduation (efficacy) boundary and a futility boundary, both pre-specified and simulated before the arm was ever activated.
Rather than a single p-value computed once at a fixed sample size, a Bayesian platform arm carries a continuously updated predictive probability of success (PPOS): given the data observed so far, what is the probability that, if the arm continued enrolling to its maximum planned sample size, it would cross the pre-specified graduation threshold?
Graduation rule: if the posterior probability that the experimental arm is superior to concurrent control exceeds a high pre-specified bound — commonly 0.975 to 0.985, calibrated by simulation to control the platform-wide false-positive rate across all arms ever tested — the arm "graduates." Graduation typically means the arm exits the adaptive randomization phase and either (a) proceeds directly to a confirmatory registration-enabling analysis using the accumulated platform data, or (b) is recommended for an independent Phase 3 trial powered on the observed effect size.
Futility rule: if the predictive probability of eventually crossing the graduation boundary falls below a low pre-specified bound — commonly 0.05 to 0.10 — the arm is stopped for futility. Patients are no longer randomized to it, freeing allocation probability for remaining and future arms.
Both boundaries are computed at every scheduled DSMB look, not just at a single terminal analysis — this is what allows platform arms to exit (in either direction) far earlier than a fixed-sample trial while maintaining well-characterized frequentist operating characteristics (type-I error, power) established through tens of thousands of pre-trial Monte Carlo simulation replicates.
A platform that tests dozens of arms over a decade faces a multiplicity problem a single two-arm trial never encounters: if each arm is tested at a nominal one-sided alpha of 0.025, and twenty arms cycle through the platform over its lifetime, the family-wise false-positive rate across the whole platform could balloon well above the conventional 5% ceiling without correction.
Platform statisticians manage this with several complementary tools:
• Arm-wise alpha allocation: because arms are evaluated sequentially and rarely borrow information from each other's primary contrast, most platforms (following FDA 2022 guidance) treat each arm-control comparison as an independently controlled test, provided the shared-control multiplicity from repeated comparisons against the same growing control pool is itself modeled correctly • Graduation boundary calibration: setting the posterior threshold at 0.975–0.985 rather than the nominal frequentist 0.95-equivalent already builds in a stringency margin that offsets the effect of interim looks and multiple candidate arms • Independent confirmatory step: many platforms treat graduation as "graduation to Phase 3," not as registration-grade evidence in itself — STAMPEDE and Lung-MAP both funnel graduated arms into a confirmatory analysis or an independent randomized comparison before regulatory submission • Full pre-specification and simulation reporting: FDA CID meetings require sponsors to submit tens of thousands of simulated platform trajectories under a range of null and alternative scenario mixtures, demonstrating that the realized false-positive rate across the arm roster stays within the pre-agreed platform-wide budget
The DSMB's charter — a document as carefully negotiated as the master protocol itself — spells out exactly which of these mechanisms govern each specific interim decision, since improvising a multiplicity correction mid-trial would undermine the entire pre-specification argument regulators require.
REMAP-CAP's corticosteroid domain reached a graduation-equivalent decision for hydrocortisone in critically ill COVID-19 patients within months of the pandemic's onset in 2020 by reusing an already-running adaptive infrastructure and a pre-specified Bayesian response-adaptive rule — a speed of evidence generation that a conventional freestanding RCT, needing to be designed, approved, and activated from scratch, could not have matched.
The final and most conceptually radical feature of a platform trial is that it is not designed to end. Arms graduate, arms are dropped for futility, and new arms are activated on the same infrastructure indefinitely — the platform's identity persists across "eras" the way a hospital persists across generations of chief residents, accumulating institutional memory (a growing shared-control dataset, a maturing Bayesian model, a battle-tested operational playbook) that makes every subsequent arm faster and cheaper to test than the one before it.
As a platform accumulates control-arm data across successive investigational arms and calendar eras, the temptation and opportunity to borrow strength statistically grows. Two complementary techniques formalize this:
• Network meta-analysis (NMA): once several arms have graduated or been dropped, their relative effect estimates versus the shared control can be synthesized into an indirect comparison network, generating effect estimates between two experimental regimens that were never directly randomized against each other — a structure impossible to construct from isolated two-arm trials • Hierarchical Bayesian borrowing across eras: control-arm outcome models are updated as a living hierarchical structure, with era-specific random effects capturing secular drift (see Stage 2) while still allowing partial pooling of information across the full accumulated control history, increasing effective control-arm sample size for future arms
This compounding-data-asset property is the core long-run economic argument for platform trials over serial standalone RCTs: published health-economic analyses of oncology and critical-care platforms estimate per-arm cost reductions on the order of 30–50% relative to freestanding Phase 2/3 programs, driven primarily by reused infrastructure, reused control data, and shortened arm-activation timelines rather than by any reduction in per-patient data-collection rigor.
A trial without a planned closing date requires governance structures conventional RCTs never need:
• Living master protocol version control: every arm activation, graduation, futility stop, and eligibility amendment is logged as a dated, numbered amendment; STAMPEDE's master protocol has passed through dozens of amendments across two decades, each independently re-reviewed by the ethics committee and MHRA/regulatory authority of record • Rotating steering committee membership: methodologists, clinicians, and patient representatives cycle through multi-year terms, with institutional knowledge transferred through a maintained design-history file rather than depending on any single investigator's tenure • Sunset and re-charter review: even "perpetual" platforms undergo periodic (typically 5-year) external scientific review of whether the chassis itself — endpoints, eligibility architecture, statistical model — remains fit for purpose as disease biology and standard of care evolve; a platform can be formally re-chartered rather than simply persisting on inertia • Data-sharing and legacy planning: because control-arm and graduated-arm data outlive any individual sponsor relationship, platforms increasingly pre-negotiate data-sharing terms (aligned with CDISC standards for interoperability) so that de-identified cumulative datasets remain usable by the wider research community after any single sponsor exits
Regulatory and health-technology-assessment bodies (EMA, FDA, and payer-facing bodies such as ICER in the U.S. and NICE in the UK) increasingly treat a platform's accumulated track record — its calibration history, its historical false-positive and false-negative rates across dozens of tested arms — as evidence in itself when evaluating the credibility of the platform's next graduation decision, a form of institutional reputation that a one-off RCT can never establish.
GBM AGILE, launched for glioblastoma in 2019 by the Global Coalition for Adaptive Research, was explicitly designed from inception to run for decades across multiple biomarker-defined subtypes and successive investigational regimens on one shared Bayesian adaptive backbone — treating "closing the trial" not as a milestone to plan for, but as a contingency to be avoided for as long as the underlying disease continues to need new therapies.