🧩 Dose-Finding Continual Reassessment Method
A continual reassessment method (CRM) used in Phase I trials to determine the maximum tolerated dose of a drug by continuously adjusting the dose levels based on patient responses and toxicity data.
Calibrating the Dose Skeleton — Turning Clinical Judgment into a Bayesian Prior
The Continual Reassessment Method (O'Quigley, Pepe & Fisher, Biometrics 1990) replaces the rigid rule-based 3+3 algorithm with a single, continuously-updated statistical model of the dose-toxicity relationship. Everything starts with the "skeleton": six investigator-elicited guesses of P(DLT) at each planned dose level, spaced along a modified-Fibonacci escalation (100%, 100%, 67%, 54%, 50% step increases) that front-loads caution at low doses and compresses steps as toxicity risk rises.
- 1990: Original CRM publication (O'Quigley, Pepe, Fisher; Biometrics)
- Empiric power: Working model (P(DLT|d)=α^exp(β), 1-parameter)
- Lee & Cheung 2009: Skeleton spacing method (indifference-interval calibration)
- 5–8: Typical dose levels (per protocol dose-escalation scheme)
From rule-based 3+3 to model-based CRM: why the skeleton matters
The classical 3+3 design treats each dose level in isolation: escalate if 0/3 patients have a DLT, expand to 6 if 1/3, de-escalate if ≥2/3. It uses no statistical model, ignores information from prior cohorts at other doses, and its final "MTD" has no defined probability target — simulation studies (Le Tourneau et al., JNCI 2009) show 3+3 identifies the true MTD only 30–40% of the time and routinely treats patients at subtherapeutic doses.\n\nCRM instead fixes a target toxicity rate θ (commonly 0.20–0.33 in oncology, reflecting an acceptable Grade 3–4 DLT rate) and fits a monotonic working model linking dose to P(DLT):\n\nEmpiric (power) model: P(DLT|dose_i) = α_i^exp(β)\n• α_i = skeleton value at level i (fixed, elicited pre-trial)\n• β = single free parameter, given a prior (commonly N(0,1.34²) or a gamma prior)\n• Model is intentionally crude — it need only be locally correct near the MTD, not globally accurate\n\nSkeleton calibration (Lee & Cheung, Clin Trials 2009):\n• Choose skeleton values so the "indifference interval" (the dose range CRM cannot statistically distinguish from θ) has consistent width across all levels\n• Too-steep skeleton → CRM escalates too aggressively; too-flat → excessive conservatism, more patients needed\n• Simulation-guided calibration: run 1,000+ simulated trials under multiple true dose-toxicity scenarios, select skeleton minimizing % correct MTD selection variance\n\nDose spacing (modified Fibonacci):\n• Level 1→2: +100% (10→20 mg)\n• Level 2→3: +100% (20→40 mg)\n• Level 3→4: +67% (40→65 mg)\n• Level 4→5: +54% (65→100 mg)\n• Level 5→6: +50% (100→150 mg)\n• Rationale: early levels almost certainly sub-toxic (steep early escalation is safe); later steps shrink as approaching the therapeutic window\n\nRegulatory context: FDA's 2022 draft guidance "Optimizing the Dosage of Human Prescription Drugs and Biological Products for the Treatment of Oncologic Diseases" (Project Optimus) explicitly endorses model-based designs like CRM and BOIN over 3+3, citing chronic-toxicity blind spots in single-cycle DLT windows and a push toward multiple dose/schedule expansion arms rather than a single MTD.
First-in-Human Dosing — Enrolling the Starting Cohort Under the Working Model
The starting dose is conventionally the lowest skeleton level with prior P(DLT) below θ, or a level derived from allometric scaling off the preclinical no-observed-adverse-effect-level (NOAEL) divided by a safety factor (typically 1/10th the human equivalent severely toxic dose in 10% of rodents, STD10). Cohort 1 patients are dosed, observed through a full DLT assessment window, and every adverse event graded per CTCAE v5.0 before any escalation decision is made.
- 21–28 days: DLT observation window (cycle 1, per protocol)
- CTCAE v5.0: AE grading standard (NCI Common Terminology Criteria)
- 1/10 STD10: Starting dose derivation (rodent severely-toxic-dose scaling)
- 1–3 pts: Typical cohort size (single-patient cohorts common <MTD)
Defining and adjudicating a dose-limiting toxicity
A DLT is a protocol-prespecified adverse event, temporally and causally linked to study drug, occurring within the observation window, that is severe enough to preclude further dose escalation. Typical DLT criteria in an oncology protocol:\n\n• Grade 4 hematologic toxicity of any duration, or Grade 3 lasting >7 days\n• Grade 3 febrile neutropenia (ANC <1.0×10⁹/L + fever ≥38.3°C)\n• Grade 3+ non-hematologic toxicity not manageable with standard supportive care\n• Treatment delay >14 days due to unresolved toxicity\n• Any treatment-related death\n\nAdjudication: a blinded or open safety review committee (SRC) — typically the sponsor medical monitor, principal investigator, and an independent safety physician — reviews each case against the protocol definition before the outcome (DLT / no-DLT) is entered into the statistical model. Ambiguous cases (e.g., toxicity possibly related to intercurrent illness) are common sources of delay; ICH E6(R2) Good Clinical Practice requires this adjudication be documented with source-data verification before database entry.\n\nSingle-patient vs. multi-patient cohorts:\n• Many modern CRM protocols use accelerated titration: cohorts of n=1 while doses remain below the lowest skeleton-predicted toxicity, switching to standard cohorts of n=3 once any Grade ≥2 toxicity is observed or a pre-specified dose level is reached\n• This roughly halves the number of patients treated at pharmacologically inactive doses versus a fixed n=3 3+3 design\n\nReal-time enrollment holds: because CRM re-estimates the model after every cohort, protocols mandate a strict enrollment pause — no new patient may start dosing until all outcomes from the active cohort have cleared the full DLT window and the statistician has re-run the model. This "real-time" constraint is CRM's main logistical cost versus rule-based designs, and is the primary driver behind rolling six-patient or TITE-CRM (time-to-event) variants that permit partial-information enrollment.
Updating the Posterior — Bayes' Theorem Turns One Cohort into a Full Dose-Toxicity Curve
This is the mathematical heart of CRM: after every cohort's outcomes are locked, the posterior distribution of the single model parameter β is recomputed by combining the prior with the binomial likelihood of observed DLTs across all patients treated so far — at any dose level. Because all six dose levels share the one parameter, a single cohort's data reshapes the toxicity estimate at every level simultaneously, which is exactly what makes CRM statistically efficient compared to per-level rule counting.
- MCMC / conjugate: Posterior computation (dfcrm, bcrm R packages)
- Binomial: Likelihood form (per-patient DLT indicator)
- N(0, 1.34²): Parameter prior (β) (weakly-informative, common default)
- After every cohort: Re-estimation frequency (real-time model lock)
The posterior update mechanics, step by step
Given the empiric working model P(DLT|dose_i, β) = α_i^exp(β), and accumulated data D = {(dose_j, y_j)} where y_j∈{0,1} indicates DLT for patient j:\n\nLikelihood: L(β|D) = Π_j [α_{dose_j}^exp(β)]^{y_j} × [1 − α_{dose_j}^exp(β)]^{1−y_j}\n\nPosterior: π(β|D) ∝ L(β|D) × π_0(β), where π_0 is the prior (typically N(0,1.34²) or a gamma density on exp(β))\n\nComputation:\n• Conjugate/normal approximation (fast, used in early implementations): treat β posterior as approximately normal via Laplace approximation around the posterior mode — adequate once n≥6\n• MCMC (Gibbs/Metropolis-Hastings via dfcrm, bcrm, or Stan): draws 10,000+ posterior samples of β, propagates uncertainty exactly, standard in modern trial statistics units\n• Posterior mean β̂ substituted back into the working model gives updated P(DLT) at every dose level i: p̂_i = α_i^exp(β̂)\n\nWhy one parameter drives six estimates:\n• A DLT observed at dose level 2 doesn't just inform level 2 — it shifts β, which simultaneously raises the estimated toxicity at levels 3, 4, 5, 6 too (and lowers estimates at level 1)\n• This "borrowing of information" across dose levels is CRM's key efficiency gain over 3+3, which treats each level as a statistically independent bucket\n• Simulation studies (Iasonos & O'Quigley, J Clin Oncol 2014) show this typically reduces the number of patients needed to correctly identify the MTD by 20–40% versus 3+3, and roughly halves the number of patients treated at doses ≥2 levels below the true MTD\n\nModel diagnostics performed at each update:\n• Posterior 95% credible interval on p_i at the current recommended dose — narrows monotonically as n grows, used directly as a stopping criterion\n• Overdose control check (Babb, Rogatko & Zacks EWOC extension): compute P(p_i > θ + ε | D); if this exceeds a feasibility bound (commonly 0.25), that dose is excluded from consideration regardless of point estimate\n• Coherence check: verify the recommended dose never skips an untried level (no-skip escalation constraint), a rule enforced independently of the raw model output for patient safety
The Dose Recommendation Rule — Converging on the Target Toxicity Level
With an updated posterior in hand, CRM's decision rule is disarmingly simple: assign the next cohort to whichever dose level (among those already tried, or one level above the highest tried) has posterior mean P(DLT) closest to θ. Safety constraints layered on top — no-skip escalation, mandatory de-escalation after excess DLTs, and overdose-control exclusion — keep this statistically-driven recommendation from ever exposing patients to an unacceptably aggressive jump.
- Max +1 level: No-skip escalation rule (per cohort, regardless of model)
- ≥2/3 DLT rate: Mandatory de-escalation (triggers automatic step-down)
- P(overdose)<0.25: EWOC feasibility bound (Babb, Rogatko & Zacks 1998)
- 18–24 pts: Typical n to MTD (vs. 24–36 for equivalent 3+3)
Comparing CRM against 3+3, mTPI-2, and BOIN in practice
CRM is one of several model-guided or model-assisted designs now standard in Phase I protocols. Each trades statistical efficiency for operational simplicity differently:
Stopping Rules and the Formal MTD Declaration
A CRM trial does not run indefinitely — it terminates the moment any one of several pre-specified stopping rules fires: a fixed maximum sample size is reached, the posterior credible interval around the recommended dose narrows below a precision threshold, or the model recommends the same dose level for three consecutive cohorts (a practical proxy for convergence). At that point the MTD is formally declared as the dose level whose posterior mean P(DLT) is closest to θ, and this dose — together with its full posterior distribution — is written into the clinical study report.
- 20–30 pts: Typical max sample size (fixed cap, protocol-specified)
- width <0.20: CrI precision stop (95% credible interval at MTD)
- 3 consecutive cohorts: Convergence stop (same recommended level)
- ±0.09–0.14: Final CrI at declared MTD (typical at trial termination)
From statistical MTD to a defensible regulatory submission
Formal MTD declaration triggers a defined set of downstream deliverables that feed directly into the IND/CTA safety package and subsequent Phase II protocol:\n\nStatistical Analysis Plan (SAP) outputs:\n• Final posterior distribution of P(DLT) at every dose level tried, with 95% credible intervals\n• Dose-toxicity curve plot with the declared MTD marked against θ\n• Operating characteristics report: simulated calibration confirming the design's true type-I-error-equivalent (probability of declaring an incorrect MTD under the assumed scenarios) — typically targeted below 20–25%\n• CTCAE-graded DLT listing by dose level, submitted as a line listing per CDISC SDTM domain AE and DS (disposition)\n\nPharmacovigilance handoff:\n• All Grade ≥3 AEs and SAEs reported through the sponsor's safety database, coded to MedDRA preferred terms, and submitted as expedited reports (7/15-day) to FDA FAERS and, for multi-region trials, EMA EudraVigilance per ICH E2B(R3) electronic transmission format\n• Any death or life-threatening event undergoes formal causality assessment before unblinding downstream cohorts (relevant in combination-agent CRM designs)\n\nBridging to Phase II:\n• The declared MTD is not automatically the RP2D — FDA's Project Optimus guidance (2022) requires sponsors to justify RP2D selection using integrated safety, PK exposure-response, and early efficacy signal data, not toxicity alone\n• Randomized dose-optimization designs comparing 2 candidate doses (both ≤MTD) in an expansion phase are now frequently requested pre-approval for oncology small-molecule and biologic INDs
The pivotal precedent for model-based dose-finding is the 2004 registration trial for a targeted kinase inhibitor in which a rule-based 3+3 design would have declared the MTD two dose levels below the level later shown, via re-analysis with CRM operating characteristics, to be the true θ=0.20 target — a gap estimated to have cost roughly 18 months of suboptimal Phase II dosing before a dose-optimization amendment corrected the regimen. Cases like this anchored FDA's Project Optimus push away from "more is better" MTD-seeking toward θ-targeted, model-based dose-finding as the expected standard for oncology INDs filed after 2023.
Expansion Cohort — Stress-Testing the MTD Before Committing to Phase II
Once the MTD is statistically declared, an expansion cohort of 10–20 additional patients is treated at that dose to confirm the toxicity estimate holds up outside the small, closely-monitored dose-escalation population, to characterize cumulative and late-onset toxicity that a single-cycle DLT window cannot capture, and to generate the PK/PD and biomarker data that ultimately determine the recommended Phase II dose (RP2D) — which regulatory guidance now expects may differ from the statistical MTD.
- 10–20 pts: Expansion cohort size (single-dose confirmation arm)
- ≥3 cycles: Cumulative toxicity window (vs. single-cycle DLT window)
- ~35–40%: RP2D below MTD (Project Optimus era) (of oncology INDs since 2022)
- AUC/Cmax vs. target occupancy: PK exposure-response endpoint (informs RP2D refinement)
Why the statistical MTD is a starting point, not the final answer
CRM's single-cycle DLT window was designed for cytotoxic chemotherapy, where toxicity is acute and dose-limiting effects appear within the first cycle. Molecularly targeted agents, immunotherapies, and antibody-drug conjugates routinely show delayed, cumulative, or immune-related toxicity that only emerges after multiple cycles — exactly the blind spot the expansion cohort is designed to catch.\n\nExpansion-cohort deliverables:\n• Extended safety follow-up (typically through cycle 3–6) captures cumulative toxicities: peripheral neuropathy, cardiotoxicity, immune-related adverse events — none of which are visible in a 28-day DLT window\n• Steady-state PK sampling across the expansion population refines the exposure-response curve, checking whether target engagement (e.g., receptor occupancy, pathway biomarker suppression) plateaus at a dose below the MTD\n• Early efficacy signals (RECIST 1.1 response rate, duration of response) collected in parallel, particularly in expansion cohorts enriched for a specific biomarker-defined population\n\nRP2D selection logic under FDA Project Optimus:\n• If exposure-response data show pathway saturation (e.g., >90% target occupancy) at a dose 1 level below MTD, sponsors are expected to select the lower dose as RP2D, trading a small tolerability margin for better long-term adherence and cycle-1 dose-reduction avoidance\n• Multi-dose randomized expansion designs (2 doses × 20–40 patients each) are increasingly requested at End-of-Phase-1 meetings specifically to generate comparative RP2D evidence rather than relying on the single MTD arm\n• Health-economic input increasingly informs RP2D too: ICER and payer bodies scrutinize whether the MTD-driven dose is cost-effective relative to a lower, near-equivalent-efficacy dose, particularly for chronic oral targeted therapies\n\nOperationally, this expansion stage is where a CRM Phase I trial hands off to CDISC SDTM-formatted datasets for the full clinical study report, and where the accumulated dose-toxicity model — now backed by 30+ patients rather than the 18–24 used for MTD declaration — becomes the definitive safety reference cited in the Phase II protocol's dose-justification section.
A continual reassessment method (CRM) used in Phase I trials to determine the maximum tolerated dose of a drug by continuously adjusting the dose levels based on patient responses and toxicity data.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install