HomeMaster Protocol & Adaptive Trial DesignUmbrella Trial Biomarker-Driven Arms

🧩 Umbrella Trial Biomarker-Driven Arms

This simulation demonstrates an umbrella trial design with multiple treatment arms based on patient biomarkers. It helps researchers and clinicians understand how to tailor cancer treatments to individual patients' genetic profiles, improving the efficacy of therapy.

Master Protocol & Adaptive Trial Design2DModerate60 FPS
umbrella-trial-design ↗ Open standalone

Master Protocol Architecture — Building the Screening Platform Once

An umbrella trial inverts the traditional one-drug-one-protocol model: instead of writing a new protocol, new IRB packet, and new site-activation process for every investigational agent, sponsors build a single master protocol for one tumor type — here, advanced non-small-cell lung cancer (NSCLC) — under which any number of biomarker-defined sub-studies can open, close, and be added over the life of the trial. The screening infrastructure, the statistical analysis plan template, the data management system, and the central IRB are shared assets, amortized across every arm that will ever run.

  • 2018 / 2022: FDA Master Protocols Guidance (draft then final guidance for industry)
  • ~700: Sites in Lung-MAP (S1400) (NCI National Clinical Trials Network)
  • ~9–12 mo: Time saved per new arm (vs. de novo standalone trial startup)
  • 1: Shared central IRB coverage (covers every current & future sub-study)

Umbrella vs. basket vs. platform — disambiguating master protocol designs

The FDA's 2018/2022 Master Protocols guidance formally distinguishes three related but distinct designs, and umbrella trials are frequently confused with the other two:

Umbrella trial: ONE tumor type (or histology), MULTIPLE biomarker-defined arms, each testing a different targeted or IO agent matched to a specific molecular alteration. Example: Lung-MAP (S1400) — one disease (squamous or non-squamous NSCLC), five to eight concurrent sub-studies, each keyed to a genomic signature.

Basket trial: ONE drug or biomarker, MULTIPLE tumor types. Example: larotrectinib's pivotal trials enrolled any TRK fusion-positive tumor regardless of primary site — pancreatic, thyroid, sarcoma, infantile fibrosarcoma — all pooled under one efficacy claim because the biomarker, not the organ, defines biology.

Platform trial: perpetual master protocol with prespecified rules for adding and dropping arms indefinitely, typically layered on top of either umbrella or basket logic. I-SPY2 (breast cancer, neoadjuvant) and REMAP-CAP (pneumonia) are platform trials that have run for a decade-plus, cycling dozens of investigational regimens through a shared adaptive statistical engine.

Lung-MAP is technically an umbrella-platform hybrid: umbrella because every arm shares the NSCLC screening funnel and biomarker panel, platform because arms are added and retired continuously rather than being fixed at trial launch.

Shared infrastructure — what actually gets reused across arms

The efficiency gain of an umbrella design comes from concretely reusable components, not just conceptual elegance:

• Central IRB: a single reliance agreement (per the NIH 2018 single-IRB-for-multisite policy and 45 CFR 46.114) covers every sub-study; new arms activate via an amendment rather than a fresh submission, cutting site-activation timelines from ~7 months to ~4–6 weeks per arm. • Shared screening log: every enrolled patient's NGS result populates one central registry (built on a CDISC SDTM-based oncology data model), so a patient who screens negative for arm A's biomarker is automatically cross-checked against every other open arm without re-consent. • Common control arm: the biomarker-negative/all-comers arm frequently serves as the reference against which several targeted arms are compared, avoiding the need for each sub-study to recruit its own separate control cohort. • Reusable statistical analysis plan template: each new arm inherits a validated Bayesian interim-monitoring framework (see Stage 4), reducing biostatistics start-up time from months to days. • Shared safety infrastructure: one Data and Safety Monitoring Board (DSMB) oversees all arms, with arm-specific charters nested inside a master DSMB charter, and one pharmacovigilance pipeline feeds FDA FAERS and, for multinational trials, EMA EudraVigilance under ICH E2B(R3) case-safety-report formatting.

Lung-MAP (SWOG S1400), launched in 2014 as the first FDA-collaborative master protocol in oncology, has screened over 4,000 patients and activated more than 15 sub-studies through a single infrastructure — several of which produced pivotal data supporting accelerated approvals (e.g., the S1400I substudy contributing to nivolumab/ipilimumab combination evidence) without each arm needing to stand up its own trial from scratch.

Central NGS Panel — Routing Every Patient to Their Matching Arm

The molecular triage step is the mechanical heart of an umbrella trial: a single centralized next-generation sequencing assay reads out every actionable alteration simultaneously, and a rules engine — not a treating physician's manual chart review — assigns the patient to whichever open arm matches. Because the panel is broad and the screening log is centralized, adding a sixth biomarker arm next year requires no new blood draw or biopsy; the historical NGS data can simply be re-queried.

  • 324: Panel size (genes) (FoundationOne CDx-class comprehensive panel)
  • 9 days: Median NGS turnaround (tissue; ~5–7 days for ctDNA liquid biopsy)
  • ~55–65%: Actionable alteration rate (of advanced NSCLC (EGFR/ALK/ROS1/KRAS/etc.))
  • FDA PMA: Companion diagnostics (CDx) required (per biomarker-drug pairing, 21 CFR 814)

From tissue to triage decision — the assay-to-assignment pipeline

The screening pipeline runs as a defined, auditable sequence:

1. Specimen acquisition: formalin-fixed paraffin-embedded (FFPE) tumor tissue block (preferred, ≥20% tumor content) or, when tissue is insufficient or unsafe to obtain, ctDNA from a peripheral blood draw (liquid biopsy) using a validated plasma-based NGS assay.

2. Comprehensive genomic profiling (CGP): hybrid-capture NGS across a panel of 300+ genes (FoundationOne CDx, Guardant360 CDx, or a trial-specific validated equivalent) reporting single-nucleotide variants, small indels, copy-number alterations, gene fusions, tumor mutational burden (TMB), and microsatellite instability (MSI) status in one report.

3. Immunohistochemistry overlay: PD-L1 tumor proportion score (TPS) via a validated IHC assay (22C3, SP263, or equivalent), run in parallel because PD-L1 expression is a protein-level readout NGS does not directly capture.

4. Rules-engine triage: a prespecified decision tree (hard-coded into the trial's electronic data capture system) maps the combined genomic + IHC result to exactly one open arm, using a strict hierarchy (e.g., EGFR/ALK/ROS1 > KRAS G12C > high PD-L1 > biomarker-negative) to prevent a single patient from qualifying for two overlapping arms simultaneously.

5. Central confirmation and randomization release: the coordinating center confirms eligibility against the current list of open arms (which can change week to week under a platform-style master protocol) before releasing the patient into that arm's specific randomization schedule.

The companion diagnostic problem — one screening test, many regulatory dossiers

Each biomarker-drug pairing tested inside the umbrella eventually needs its own FDA-cleared or -approved companion diagnostic (CDx) under 21 CFR 814 if the drug reaches approval — even though all the arms share one screening NGS panel operationally. This creates a structural tension the trial's regulatory strategy must resolve upfront:

• Panel-based CDx bridging studies: sponsors run concordance analyses between the broad research-use NGS panel used for triage and the narrower, single-analyte or panel-based CDx that will carry the eventual drug label, so that the pivotal efficacy data generated using the research panel can be bridged to the labeled diagnostic. • Local vs. central testing discordance: local hospital NGS platforms used for initial clinical decision-making do not always agree with the trial's central confirmatory assay — published concordance for common actionable driver alterations runs approximately 90–95%, meaning 5–10% of patients require re-testing or are found ineligible on central confirmation. • Retrospective re-querying: because genomic data is stored centrally, when a sixth arm opens 18 months into the trial for a newly druggable alteration (e.g., a rare MET exon 14 skipping mutation), patients already screened — but not yet assigned or already assigned to the biomarker-negative arm — can be retrospectively flagged and offered the new arm without a repeat biopsy.

In Lung-MAP, central-lab NGS turnaround and rules-engine triage reduced the historical multi-test, multi-visit biomarker workup (sequential single-gene PCR/FISH tests taking 3–6 weeks) to a single 9-day comprehensive panel result — the operational change that made running five to eight concurrent sub-studies off one blood draw or biopsy logistically possible in the first place.

Biomarker-Matched Arms — Independently Powered, Commonly Governed

Once triaged, each patient enters a sub-study that is statistically self-contained — powered, analyzed, and reported on its own efficacy endpoint — while still operating under the master protocol's shared governance, safety monitoring, and (frequently) a shared control arm. Within each arm, response-adaptive randomization increasingly favors the regimen showing the stronger signal, so patients enrolled later in the arm's life have a higher probability of receiving the better-performing treatment.

  • 50–120: Typical arm sample size (per biomarker sub-study (Simon two-stage or Bayesian))
  • Bayesian RAR: Randomization method (response-adaptive randomization, Thompson-sampling family)
  • ~60–70%: Shared control arm usage (of umbrella arms borrow the all-comers control)
  • 4–6: Median concurrent open arms (in mature platform-style umbrella trials)

Independent power, shared governance — the statistical contract between arms

Each arm is designed as if it were a standalone Phase 2 study — with its own null hypothesis, its own sample-size calculation, and its own primary endpoint (typically objective response rate per RECIST v1.1, sometimes progression-free survival) — but nested inside common master-protocol machinery:

• Arm-level design: a Simon two-stage design (screening for futility after ~15–20 patients before committing to full enrollment) or a Bayesian design with continuous monitoring is used per arm, since driver-mutation subgroups are often too rare to support a fixed, large, frequentist Phase 3 sample size on their own. • Type I error allocation: because the master protocol runs multiple simultaneous hypothesis tests, the overall trial-wide false-positive rate must be controlled using a graphical multiple-testing approach (Bretz–Maurer–Posch weighted Bonferroni graphs are the most common), allocating a fraction of the overall alpha to each arm rather than letting each arm test at an unadjusted 0.05. • Common control borrowing: when several arms share the biomarker-negative or standard-of-care control, a Bayesian hierarchical model (BHM) allows partial "borrowing of strength" across arms' control data — increasing effective control-arm sample size and statistical power for each individual comparison without literally re-randomizing more patients to control. • Common data standards: every arm reports using the same CDISC SDTM Oncology domain structure and RECIST-based tumor-assessment CRFs, so pooled cross-arm safety analyses (Stage 5) are mechanically straightforward even though efficacy is analyzed arm-by-arm.

Response-adaptive randomization within an arm

Inside arms that test more than one dose, schedule, or combination against a shared control (a common structure when a biomarker match has more than one plausible regimen), response-adaptive randomization (RAR) reallocates future patients toward whichever regimen has accrued the stronger interim response signal:

Allocation probability update (Thompson-sampling style): P(assign to regimen k) ∝ P(regimen k is best | current data)

This posterior probability is recomputed at each prespecified interim look using a Bayesian beta-binomial model on response (responder/non-responder) or a Bayesian normal model on a continuous endpoint, so allocation smoothly drifts — typically over 30–60% of the arm's total enrollment window — from a 1:1 start toward a skewed ratio (e.g., 75:25 or 80:20) favoring the better regimen, without ever completely abandoning randomization (a floor allocation, often 10–15%, is retained for every regimen including control to preserve unbiased comparison and safety surveillance).

This differs fundamentally from a simple non-randomized "pick the winner" design: RAR still generates a formally randomized, statistically valid comparison, it simply exposes fewer trial participants to the inferior regimen as evidence accumulates — a meaningful ethical advantage in oncology, where the inferior arm may mean a lower-response therapy for a life-threatening disease.

Representative arm structure inside an NSCLC umbrella master protocol

ProductIndicationTrial DesignKey Result
EGFR ex19del/L858R~14% of NSCLC adenocarcinoma3rd-generation EGFR TKI vs. 1st/2nd-gen comparatorORR 70–80%, median PFS 18–20 mo
ALK fusion~4–5% of NSCLC adenocarcinomaNext-gen ALK inhibitor vs. prior-gen ALK TKICNS penetration, ORR 60–70%
KRAS G12C~13% of NSCLC adenocarcinomaCovalent KRAS G12C inhibitor ± combinationFirst targeted option for a historically undruggable driver
PD-L1 high (TPS≥50%)~25–30% of NSCLCAnti-PD-1 monotherapy vs. chemo-IODurable responses, favorable toxicity vs. chemo
Biomarker-negativeRemaining ~20–25%Chemo-IO backbone; serves as shared controlCommon comparator anchoring cross-arm inference

Graduation, Futility Pruning, and Mid-Trial Arm Addition

The defining operational advantage of a master protocol is that it does not run to a single fixed endpoint — it runs to a schedule of prespecified interim looks at which each arm's data are evaluated independently against efficacy and futility boundaries. Arms that are clearly working graduate early to expansion cohorts and downstream regulatory packaging; arms that are clearly not working are stopped and their patients, sites, and drug supply reallocated; and arms for newly emerging biomarker-drug pairs can be added without disturbing the arms already running.

  • Every 15–25 pts: Interim look cadence (per arm, or fixed calendar intervals)
  • Post. prob. >0.95: Efficacy stopping boundary (that ORR exceeds historical control)
  • Post. prob. <0.05: Futility stopping boundary (that ORR will ultimately exceed threshold)
  • ~40%: Arms pruned historically (Lung-MAP) (of opened sub-studies closed for futility)

Bayesian predictive probability — the graduation/futility decision engine

Each arm's interim monitoring plan is built on a Bayesian predictive probability (PP) framework rather than a classical group-sequential O'Brien-Fleming boundary, because PP naturally accommodates the master protocol's asynchronous, rolling-enrollment structure where arms open and close on different calendars:

At each interim look with n patients evaluable: • A Beta(a,b) prior on the true response rate is updated with observed responses to a Beta(a+r, b+n-r) posterior, where r = observed responders. • Predictive probability of eventual success = P(final posterior response rate exceeds the prespecified target, e.g., historical control + 15–20 percentage points | current data), integrated over all possible remaining-patient outcomes to full arm enrollment. • Efficacy graduation boundary: PP > 0.95 → arm graduates, enrollment may expand into a larger confirmatory cohort or the arm is flagged for regulatory packaging. • Futility boundary: PP < 0.05 → arm is closed; remaining drug supply, site capacity, and the DSMB's attention reallocate to open or incoming arms. • Continue zone: 0.05 ≤ PP ≤ 0.95 → arm continues to next interim look unchanged.

This PP framework is why umbrella/platform trials report data continuously rather than at one terminal analysis — investigators, sponsors, and the DSMB always know each arm's current probability of ultimate success, and can act on it immediately rather than waiting for a fixed calendar-driven interim.

Lung-MAP's S1400 platform has closed roughly 40% of its opened sub-studies for futility at a first or second interim look, each closure typically occurring after only 15–25 evaluable patients rather than the 80–120 a standalone Phase 2 would have required to reach the same conclusion — the central efficiency claim of adaptive master protocols: negative answers arrive faster and expose fewer patients to an ineffective regimen.

Adding arms mid-trial without disturbing what is already running

A platform-style umbrella protocol is written so a new arm can be appended via protocol amendment rather than a new trial:

• Master statistical analysis plan template: pre-approved by the central IRB and, for registrational intent, discussed with FDA in advance (frequently under a Special Protocol Assessment or Type C/Type D meeting) so a new arm's design only needs sponsor-specific parameters filled in — sample size, boundaries, comparator — rather than a full de novo statistical review. • Screening log re-query: patients already profiled by the central NGS panel whose alteration matches a newly opened arm (but who were previously ineligible for any open arm and defaulted to the biomarker-negative bucket) can be identified retrospectively and offered re-consent into the new arm. • Independent DSMB charter per arm nested under the master DSMB: new arms get their own safety-monitoring charter reviewed and appended at the same board meeting, without requiring the board to re-review the entire platform's accumulated safety history. • Drug supply and site logistics: because the enrollment infrastructure, consent language, and data systems are already validated, a genuinely new arm can typically activate its first site within 4–8 weeks of protocol amendment approval — an interval that would be 6–12 months for a standalone trial.

Pooled Safety, Accelerated Approval, and Companion-Diagnostic Label Expansion

The master protocol's final payoff arrives when a graduated arm converts into a regulatory submission. Because all arms share pharmacovigilance infrastructure, pooled cross-arm safety data materially strengthens the adverse-event signal-detection power available to any single arm's package, even while each biomarker subgroup's efficacy claim stands on its own independently powered data.

  • Subpart H: FDA Accelerated Approval pathway (21 CFR 314.500 series; surrogate endpoint (ORR))
  • ~6–9 mo: Median time to filing after graduation (vs. 18–24 mo for a standalone Phase 2→3 program)
  • >1,000 pts: Pooled safety database (typical umbrella) (across all arms combined, strengthens rare-AE detection)
  • Mandatory: Confirmatory trial requirement (post-marketing per Accelerated Approval commitment)

From graduated arm to accelerated approval filing

A graduated arm's pathway to market leverages the surrogate-endpoint flexibility of FDA's Accelerated Approval program (21 CFR Part 314 Subpart H, and the parallel Subpart E for serious/life-threatening conditions), which was built for exactly this situation — a molecularly defined subgroup with a compelling but not yet survival-confirmed response signal:

• Surrogate endpoint: objective response rate (ORR) and duration of response (DoR) per RECIST v1.1, assessed by both investigator and blinded independent central review (BICR), substitute for the overall-survival endpoint a traditional approval would require. • Single-arm sufficiency in rare subgroups: for biomarker subgroups too small to power a randomized comparison within a reasonable timeframe (e.g., a driver mutation present in 2–4% of the tumor type), FDA has repeatedly accepted single-arm data from an umbrella sub-study with a historical-control comparison, provided the effect size is large and the natural history well characterized — the same regulatory logic used for basket-trial approvals like larotrectinib and entrectinib. • Companion diagnostic co-review: the arm's CDx (Stage 2) undergoes parallel PMA review so the drug label and the diagnostic label are approved concurrently, preventing a gap where the drug is approved but no cleared test exists to identify eligible patients. • Mandatory confirmatory trial: Accelerated Approval always carries a post-marketing requirement (PMR) to complete a confirmatory trial verifying clinical benefit (typically PFS or OS); the FDA's 2021–2023 accelerated-approval-withdrawal enforcement wave (following ODAC review of several indications with stalled confirmatory data) has made timely confirmatory-trial completion a hard requirement rather than a formality.

Pooled pharmacovigilance and international regulatory alignment

Even though efficacy claims stay arm-specific, safety surveillance benefits from full pooling across the master protocol:

• Shared pharmacovigilance pipeline: all serious adverse events across every arm feed one case-processing workflow, coded in MedDRA and formatted per ICH E2B(R3), flowing to FDA FAERS and, for global trials, EMA EudraVigilance under EMA GVP Module VI (adverse reaction reporting) and Module IX (signal management). • Increased rare-AE detection power: pooling ~1,000+ patients across arms (versus 50–120 in any single arm) materially raises the statistical power to detect an uncommon but serious class-effect toxicity — e.g., an immune-related adverse event common to every PD-1-containing arm — that no individual sub-study would be powered to characterize alone. • ICH E20 alignment: ICH's E20 guideline on adaptive clinical trial designs (finalized 2023–2024) codifies internationally harmonized expectations for exactly this master-protocol structure, reducing the historical divergence between how FDA and EMA reviewed adaptive designs and easing simultaneous multi-region filings. • Health-technology-assessment (HTA) and payer evidence: bodies such as ICER in the US build cost-effectiveness models arm-by-arm because each biomarker subgroup has a distinct comparator, price, and effect size — a single blended umbrella-trial cost-effectiveness number is not clinically meaningful, so payer value assessments mirror the trial's own biomarker-stratified structure. CMS coverage determinations for the paired companion diagnostic typically follow FDA's CDx approval within the same review cycle.

The Lung-MAP infrastructure directly enabled biomarker-matched sub-studies to reach FDA submission roughly 40–50% faster than comparable standalone trials for the same indications, because screening, IRB, DSMB, and pharmacovigilance systems did not need to be rebuilt for each new agent — several sub-studies' data have supported label expansions in genomically defined NSCLC subgroups without a separate ground-up trial for each molecule.
⚙ Under the hood

This simulation demonstrates an umbrella trial design with multiple treatment arms based on patient biomarkers. It helps researchers and clinicians understand how to tailor cancer treatments to individual patients' genetic profiles, improving the efficacy of therapy.

CanvasBiomedicine

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)