HomePatient-Derived Xenograft (PDX) Mouse ModelPDX Cohort Drug Efficacy Trial ("Mouse Clinical Trial")

🐭 PDX Cohort Drug Efficacy Trial ("Mouse Clinical Trial")

A clinical trial on mice using PDX models to test the efficacy of a treatment.

Patient-Derived Xenograft (PDX) Mouse Model2DModerate60 FPS
pdx-mouse-clinical-trial ↗ Open standalone

PDX Panel Assembly — Capturing Tumor Heterogeneity in a Living Biobank

A patient-derived xenograft (PDX) is created by implanting a fragment of a patient's tumor directly into an immunodeficient mouse, bypassing cell culture entirely. Because tissue architecture, stromal admixture, and clonal heterogeneity are largely preserved through early passages, a panel of dozens to hundreds of independently engrafted PDX models functions as a surrogate patient population — a "mouse clinical trial" that can be dosed, biopsied, and analyzed in ways no human cohort permits.

  • >500: Models in major PDX biobanks (NCI PDMR, PDXNet, Novartis PDXE)
  • 40–75%: Primary engraftment success (depends on tumor type, stage)
  • 8–16 wk: Time to first-passage tumor (subcutaneous or orthotopic)
  • 30–100+: Typical trial panel size (models spanning subtypes)

Why xenografts, not cell lines

Traditional cancer pharmacology relies on immortalized cell lines grown for decades in plastic flasks — a process that strips away tumor architecture, selects for the fastest-dividing clones, and erases the stromal, vascular, and immune microenvironment that shapes real drug response. PDX models sidestep this by transplanting minced patient tumor fragments (2–4 mm³) directly into the flank or orthotopic site of an immunodeficient mouse (NSG, NOG, or nude strains lacking functional T, B, and NK cells).

Across the first 2–4 passages, PDX tumors retain the histology, mutational spectrum, copy-number landscape, and even much of the intratumoral clonal architecture of the originating patient biopsy — properties documented to erode rapidly in 2D cell culture. Human stroma is gradually replaced by mouse stroma by passage 3–4, but the epithelial tumor compartment — and the driver mutations it carries — remains patient-derived.

This fidelity is precisely what a "mouse clinical trial" needs: not one artificial model of "breast cancer," but dozens of genuinely distinct breast cancers, each behaving according to its own molecular biology.

The Novartis PDX Encyclopedia (Gao et al., Nature Medicine 2015) profiled 1,075 PDX models across 6 tumor types against 62 anti-cancer agents — the founding demonstration that a PDX panel could statistically recapitulate the heterogeneous response rates seen in actual clinical trials.

Panel design — spanning genomic subtypes deliberately

A well-designed PDX efficacy panel is not a random convenience sample — it is deliberately constructed to span the molecular subtypes relevant to the drug's hypothesized mechanism. For a targeted therapy, the panel typically stratifies models by:

• Driver mutation status (e.g., KRAS G12C, EGFR exon 19del, BRAF V600E) — usually enriched relative to population frequency so enough events exist to test the hypothesis • Transcriptomic subtype (e.g., TNBC, HER2-enriched, luminal A/B in breast cancer; classical vs. basal in pancreatic cancer) • Copy-number and mutational burden strata, to probe whether genomic instability predicts sensitivity • Prior therapy exposure, mirroring the line-of-therapy the drug would enter clinically

Each model is typically represented in the panel by one or two mice per treatment arm — an extreme departure from classical pharmacology, where large groups (n=8–10) of a single model are used to control biological noise. In a PDX panel, the "replicate" is not the mouse — it is the model: heterogeneity across dozens of tumors substitutes for repeated measurement within one tumor, and statistical power comes from panel breadth rather than arm depth.

Quality control before dosing begins

Before any drug is administered, every model entering the panel must pass quality gates:

• Histopathological concordance: H&E sections compared to the original patient biopsy to confirm tumor grade and architecture were preserved through engraftment • Short tandem repeat (STR) profiling: confirms the xenograft still matches the originating patient and has not been cross-contaminated by another line in the vivarium • Growth kinetics screening: models must reach a baseline tumor volume of 150–250 mm³ within a defined window and show reproducible exponential growth before randomization • Passage number cap: most panels restrict use to passage ≤6, since murine stromal replacement and genetic drift accelerate beyond this point

Only models clearing all four gates are randomized into treatment and vehicle arms — ensuring that any efficacy signal reflects drug biology, not biobank artifact.

Randomized 1×1×1 Cohort Dosing — Building a Population Trial from Single Mice

The defining innovation of the PDX "mouse clinical trial" design, formalized by Gao et al. (2015), is the 1×1×1 architecture: one animal per model per treatment arm, replicated across a large panel of models, rather than the classical approach of many animals bearing one model. This inverts the unit of statistical power from "mice within a tumor" to "tumors within a population" — mirroring how a real Phase II trial enrolls many different patients rather than dosing one patient many times.

  • QDx21: Standard dosing schedule (once daily × 21 days, IP or PO)
  • n=1–3: Typical arm size (per model) (vehicle and treatment)
  • 150–250 mm³: Randomization trigger volume (baseline tumor size)
  • 25–150: Dose range tested (this panel) (mg/kg, 5-level titration)

The 1×1×1 design and its statistical logic

In a classical xenograft study, a single PDX model is grown in 8–10 mice per arm; the study asks "does this drug shrink this tumor?" and answers it with high within-model precision. A PDX panel trial asks a different question — "what fraction of a heterogeneous patient population responds?" — and 1×1×1 design answers it by trading within-model replicates for across-model breadth.

Each of the (say) 60 models in the panel contributes one tumor-bearing mouse to the vehicle arm and one to the treatment arm — for a 60-model panel, a 120-mouse total trial. Model identity is the unit of analysis: each animal pair (vehicle vs. treatment, same PDX line) functions like a matched patient enrolled in a two-arm trial, and the panel-wide response rate is computed by pooling matched pairs across all 60 "patients" simultaneously.

This design amplifies statistical power against the central question that matters clinically — response rate across a heterogeneous population — while modestly sacrificing power to detect small effects in any single model. It also dramatically reduces total animal usage relative to running the same panel with n=8–10 mice per model per arm, a meaningful 3Rs consideration.

A 60-model panel run 1×1×1 uses roughly 120–180 mice total. Running the same panel with classical n=8 arms would require 960–1,440 mice — an 8-fold difference that makes population-scale preclinical trials logistically and ethically feasible.

Dosing schedule, route, and vehicle control design

Dose and schedule are chosen to mimic the intended clinical regimen as closely as species allometry allows:

• QDx21 (once-daily dosing for 21 consecutive days) is the most common schedule for small-molecule and targeted agents, long enough to distinguish durable regression from transient growth delay • Route: intraperitoneal (IP) injection is common for early efficacy screening due to reliable bioavailability; oral gavage (PO) is used when testing the clinical formulation directly • Dose selection: a maximum-tolerated-dose (MTD) range-finding study in non-tumor-bearing mice precedes the efficacy panel; the efficacy dose is typically set at 60–80% of MTD to preserve a safety margin while maximizing exposure • Vehicle arm: receives the identical formulation carrier (e.g., 0.5% methylcellulose, DMSO/PEG cosolvent) on the identical schedule — controlling for injection stress, handling, and formulation vehicle effects on tumor growth, not just "no drug" as a null condition

Body weight is tracked alongside tumor volume throughout dosing; a loss >20% from baseline triggers a mandatory dose holiday or study withdrawal under standard IACUC-approved protocols, since PDX panels test both efficacy and a first-pass tolerability signal.

Randomization and blinding within a panel study

Even within a single PDX model, tumor-bearing mice are randomized to vehicle vs. treatment only after reaching the baseline volume window (150–250 mm³) and only in a way that balances starting tumor size between arms — an unbalanced baseline is one of the most common sources of spurious "efficacy" in xenograft literature.

Best practice includes:

• Block randomization within each model: mice from the same PDX line are ranked by baseline volume and alternately assigned to arms, preventing systematic size bias • Blinded volumetric assessment: caliper measurements performed by staff unaware of treatment assignment where feasible • Pre-specified analysis plan: which endpoint (best response, %TGI at day 21, time-to-progression) will define "response" is fixed before unblinding, preventing post-hoc threshold shopping across a panel where hundreds of comparisons are possible

Tumor Growth Curves & RECIST-like Response Criteria for Mice

Twice-weekly caliper measurements convert each animal into a longitudinal growth curve, and each model's trajectory is scored against response criteria deliberately modeled on RECIST 1.1 (the human solid-tumor imaging standard) but rescaled for the compressed, rapid kinetics of a 21-day mouse study. The result is a single categorical label — CR, PR, SD, or PD — that lets a panel of dozens of biologically dissimilar tumors be compared on common ground.

  • V=(L×W²)/2: Volume formula (modified ellipsoid, caliper-based)
  • 2×/week: Measurement frequency (or 3× for fast-growing lines)
  • ≤ −30%: PR threshold (PDX-adapted) (best % change from baseline)
  • ≥ +20%: PD threshold (PDX-adapted) (best % change from baseline)

From caliper to curve — measuring tumor volume in vivo

Tumor length (L, longest axis) and width (W, perpendicular axis) are measured with digital calipers, and volume is estimated by the modified ellipsoid formula V=(L×W²)/2, which approximates a prolate spheroid and correlates well with true excised tumor mass for subcutaneous flank tumors. Orthotopic or metastatic models instead rely on bioluminescence imaging (BLI) or MRI, since caliper access is impossible.

Measurements begin at randomization (day 0, baseline volume V₀) and continue 2–3 times weekly through day 21 or until a tumor reaches the institutional volume endpoint (typically 2,000 mm³ or 10% of body weight), at which point that animal is euthanized per humane endpoint policy and its last measured volume is carried forward for analysis.

Each animal's raw volumes are converted into a percent-change-from-baseline curve: %Δt = 100×(Vt−V0)/V0. The panel view stacks these curves — spaghetti-plot style — revealing at a glance which models regress, plateau, or accelerate under treatment relative to their matched vehicle-arm partner.

RECIST-like criteria adapted for xenografts

Human RECIST 1.1 defines response using unidimensional target-lesion diameters measured by CT/MRI: CR (disappearance), PR (≥30% diameter decrease), PD (≥20% increase), SD (neither). PDX pharmacology adapts this framework to tumor volume and to the compressed mouse timescale, following conventions popularized by Gao et al. (2015) and widely adopted across PDX consortia:

• Complete Response (CR): best volume change ≤ −90% to −100% from baseline (tumor becomes non-palpable) • Partial Response (PR): best volume change ≤ −30% from baseline • Stable Disease (SD): best volume change between −30% and +20% from baseline • Progressive Disease (PD): best volume change ≥ +20% from baseline, OR any new lesion

Critically, "best response" — not the endpoint value — is used, following RECIST convention: a tumor that shrinks 40% by day 14 and then regrows to −10% by day 21 is still scored PR, because transient deep regression is itself informative of drug mechanism even if resistance eventually emerges.

A complementary metric, %TGI (percent tumor growth inhibition), compares treated to control growth directly: %TGI = 100 × (1 − ΔT/ΔC), where ΔT and ΔC are the mean volume changes in treatment and vehicle arms. %TGI ≥60% is a common preclinical bar for declaring a compound "active" enough to prioritize for IND-enabling studies.

Sources of noise a panel must tolerate

Unlike a single well-powered xenograft study, a 1×1×1 panel cannot average away individual-animal noise within a model — so response calls must be interpreted with appropriate caution:

• Engraftment-site variability: subcutaneous tumors can grow asymmetrically depending on injection depth and local vascularity, inflating caliper measurement error by 10–15% • Single-animal categorization risk: with n=1 per arm per model, a CR/PD call for any individual model carries real sampling uncertainty — this is why panel-level statistics (Stage 4), not single-model calls, drive the primary conclusion • Humane endpoint censoring: models that hit volume limits early in the vehicle arm before day 21 require %TGI to be computed at the last common time point across both arms, not a fixed calendar day, to avoid bias

Response Rate Statistics — Waterfall Plots and Biomarker Stratification

Once every model in the panel has a best-response category, the panel-level question becomes statistical: what fraction of tumors respond, and can that response be predicted in advance from a genomic, transcriptomic, or proteomic biomarker? The waterfall plot — models ranked left to right by best percent change — is the field's standard visualization, instantly separating true responders from a heterogeneous background of stable and progressive disease.

  • 10–25%: Typical unselected-panel ORR (CR+PR rate, targeted agents)
  • 50–70%: Biomarker-enriched subgroup ORR (when a true predictive marker exists)
  • Fisher's exact: Statistical test for enrichment (or logistic regression on genomic features)
  • ~40–50: Minimum panel size for subgroup power (models, ≥8–10 biomarker+ cases)

Reading the waterfall plot

Each bar in a waterfall plot represents one PDX model's best percent change in tumor volume from baseline, sorted from the deepest regression (left) to the greatest growth (right). Horizontal reference lines at −30% and +20% mark the PR and PD thresholds, so a glance at how many bars cross the green line versus the red line gives an immediate visual read of objective response rate (ORR = CR+PR fraction) across the entire panel.

Because every bar is a genuinely distinct tumor biology, the shape of the waterfall itself is diagnostic: a shallow, gradually sloping waterfall suggests a modestly active but non-predictive compound (most models nudge downward a little); a bimodal waterfall — a cluster of deep responders sitting apart from a cluster of frank progressors — is the classic signature of a targeted therapy whose activity depends on a specific molecular lesion present in only a subset of models. It is exactly this bimodal pattern that motivates biomarker stratification analysis.

Stratifying responders from non-responders by biomarker

Once ORR is established for the whole panel, each model's pre-treatment molecular profile — mutation calls, copy number, RNA-seq expression, IHC protein levels — is tested for association with response category. The core analysis:

1. Define the candidate biomarker as binary (mutant/wild-type, high/low expression by a pre-specified cutoff) across all models 2. Cross-tabulate biomarker status against responder/non-responder status (CR+PR vs. SD+PD) 3. Test association using Fisher's exact test (small panels) or logistic regression with response as outcome and biomarker plus relevant covariates as predictors 4. Compute the enrichment ratio: ORR in biomarker-positive models ÷ ORR in biomarker-negative models

An enrichment ratio of 3–5× is generally considered clinically actionable — strong enough evidence to justify designing the human trial around that biomarker rather than treating an unselected population. A ratio near 1× indicates the biomarker does not meaningfully predict response, at least at the tested cutoff, and the search continues across other candidate features.

In the Novartis PDX Encyclopedia, the MEK inhibitor binimetinib showed an unselected-panel ORR under 15%, but RAS/RAF-mutant models responded at roughly 3–4× the rate of RAS/RAF-wild-type models — a signal that closely mirrored the biomarker-defined response pattern later confirmed in human MEK-inhibitor trials.

Statistical power and the limits of n=1 designs

Because each model contributes a single categorical response call, panel-level statistics depend heavily on total panel size and on how many models fall into the biomarker-positive subgroup. A panel of 60 models with a biomarker prevalence of 30% yields only ~18 biomarker-positive cases — adequate to detect a large enrichment effect (3× or more) with reasonable confidence, but underpowered to detect subtler 1.5–2× effects.

Consortium-scale efforts address this by pooling data across institutions: the PDXNet consortium and EurOPDX network aggregate response data across thousands of models internationally, allowing rarer biomarker subgroups (present in <10% of a single institution's panel) to reach adequate sample size for robust association testing before a hypothesis is carried into costly human trials.

From Panel to Protocol — Designing the Human Trial Around Preclinical Statistics

The entire purpose of a PDX panel trial is to inform decisions no purely computational or single-model experiment can make with confidence: how big should the human trial be, which patients should be enrolled, and what response rate should be considered a meaningful clinical signal. Panel-level statistics are converted directly into trial design parameters months to years before the first human patient is dosed.

  • Common: Simon two-stage design use (Phase II ORR-driven trials)
  • 40–80: Typical enrichment trial size (biomarker-selected patients)
  • Moderate–High: Preclinical-to-clinical ORR concordance (when biomarker signal is strong)
  • Rising: Basket trial adoption (biomarker-defined) (tumor-agnostic, mutation-defined cohorts)

Sizing a Phase II trial from a preclinical response rate

A PDX panel's biomarker-positive ORR becomes the working assumption for statistical trial design. Simon's two-stage design — the workhorse of single-arm Phase II oncology trials — requires exactly this kind of prior estimate: a "null" response rate (p0, the rate that would not justify further development) and an "alternative" response rate (p1, the rate that would).

If a PDX panel shows biomarker-positive ORR of ~55% against an unselected-panel background of ~15%, a trial team might set p0=20% (slightly above historical unselected response for the tumor type) and p1=45% (a conservative discount off the preclinical 55%, accounting for known preclinical-to-clinical attrition). Simon's design then computes the minimum number of biomarker-positive patients needed — often 40–55 — to distinguish these two hypotheses with standard 80% power and one-sided α=0.05, including a planned interim analysis after a first-stage cohort to allow early stopping for futility.

Without the PDX panel, a trial team would have only single-model xenograft data or in vitro cell-line screens to anchor these assumptions — far weaker predictors of the heterogeneous response actually observed in a genetically diverse patient population.

Writing patient stratification into the protocol

The biomarker enrichment ratio identified in Stage 4 directly shapes trial eligibility criteria:

• Strong enrichment (≥3×): the trial is designed as a biomarker-selected (enriched) study — only patients whose tumors carry the biomarker are eligible, maximizing the observed response rate at the cost of a smaller eligible population and a companion diagnostic requirement • Moderate enrichment (1.5–3×): an all-comers trial with mandatory pre-specified biomarker subgroup analysis is often chosen instead, preserving broader accrual while still testing the predictive hypothesis statistically • Basket / umbrella designs: when the same molecular biomarker recurs across PDX models from multiple tumor-of-origin types (e.g., a KRAS G12C mutation appearing in lung, colorectal, and pancreatic PDX models), the panel data support a tumor-agnostic basket trial — enrolling patients by mutation status rather than by organ of origin

Companion diagnostic development — the clinical-grade assay that will identify biomarker-positive patients at the bedside — is typically initiated in parallel with Phase I, using the exact biomarker definition and cutoff validated against panel response data, to avoid a costly assay-protocol mismatch discovered only after Phase II enrollment begins.

Preclinical-to-clinical translation is never guaranteed: PDX panels lack a competent immune system (by design, to permit engraftment), so immunotherapy response cannot be modeled this way, and any drug whose mechanism depends on tumor-immune interaction requires humanized-mouse or syngeneic panel alternatives instead.

What panel trials get right — and their limits

Retrospective analyses comparing PDX panel response rates to the eventual human trial outcome for the same compound show moderate-to-good concordance specifically when the preclinical biomarker signal is strong and mechanistically well understood (e.g., oncogene-addicted tumors responding to a matched targeted inhibitor). Concordance weakens for compounds with more diffuse mechanisms, for immuno-oncology agents (excluded by the immunodeficient host), and for combination regimens where PDX-host pharmacokinetic differences (mouse metabolism, dosing frequency limits) diverge meaningfully from human dosing.

Despite these caveats, the PDX mouse clinical trial remains one of the most cost-effective tools in oncology drug development for a single purpose: converting "does this drug work in principle" into "in whom, and how confidently, should we test it next" — turning a panel of individually anecdotal mouse experiments into a population-level, statistically defensible go/no-go decision before human risk is incurred.

⚙ Under the hood

A clinical trial on mice using PDX models to test the efficacy of a treatment.

CanvasBiomedicine

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)