📈 Go/No-Go Decision Gate Criteria Simulator
This simulation evaluates the criteria for deciding whether to continue or terminate a project at critical decision points, ensuring resource efficiency and strategic alignment.
Gate Criteria Pre-Specification
A stage gate is only as trustworthy as the criteria drawn before the data existed. In R&D portfolio management, the single most important discipline at any decision gate is locking quantitative, falsifiable success criteria in writing — and getting sign-off from every function that will later sit on the review — before anyone unblinds the trial or sees the readout.
- 3–6: Typical criteria per gate (quantitative, pre-registered)
- 4+ functions: Sign-off required from (clinical, reg, commercial, mfg)
- Database lock: Charter locked before (and unblinding)
- Prohibited: Post-hoc criteria changes (by governance policy)
Why pre-specification exists at all
Confirmation bias is not a hypothetical risk in portfolio decision-making — it is the default failure mode. Once a team has spent years and tens of millions of dollars on a program, there is enormous unconscious pressure to interpret ambiguous data favorably. A trial that narrowly misses its primary endpoint but shows a promising trend in a subgroup will, absent a pre-specified rule, generate a dozen plausible-sounding reasons to call it a partial success.
Pre-specification breaks this loop by forcing the criteria-setting conversation to happen under a "veil of ignorance" — before anyone knows which way the data will cut. A team that has never seen the results has no stake in rationalizing them, so the criteria they set reflect what the program actually needs to prove, not what the data happens to show. This is the same logic that underlies pre-registration of statistical analysis plans in clinical trials, and it generalizes to every decision gate in a stage-gate portfolio process, not just Phase 2/3 go/no-go.
The test of a good gate criterion is whether it is falsifiable in advance: a reviewer reading the charter before the readout should be able to state exactly what data pattern would fail the program, not just what would pass it.
What a well-formed criterion looks like
Each pre-specified criterion typically has four components: a quantitative metric (not a vague "positive trend"), a numeric threshold, the population and timepoint it is measured on, and the analysis method that will be used to evaluate it. A criterion that says "efficacy should improve" is not usable at a gate; a criterion that says "the pre-specified primary endpoint must show a between-arm difference of at least 1.5 points on the intent-to-treat population at the primary analysis timepoint, with the 95% confidence interval excluding zero" is.
Criteria typically span three domains: efficacy (the core biological or clinical effect the program must demonstrate), safety (the ceiling on adverse events the program can tolerate, usually expressed as a rate difference or an absolute cap on a specific event type), and a translational or biomarker criterion that provides mechanistic confirmation the drug is doing what it is supposed to do at the molecular level, independent of the clinical readout. A program that hits its clinical endpoint through an unexpected mechanism, or misses it despite clean biomarker engagement, tells the committee very different things about what to do next.
Governance: who signs the charter, and why it matters
The gate charter is typically co-signed by representatives of every function that will later sit on the review committee, precisely so that no function can later claim the criteria did not reflect their concerns. A regulatory lead who signs off on a safety threshold before the readout cannot credibly argue after an unfavorable result that the bar was set in the wrong place — that conversation has to happen up front, when it is not colored by the outcome.
Many organizations also require that the pre-specified criteria be filed with an independent function (a portfolio management office or a data monitoring committee secretariat) that has no incentive to alter them later, and that any deviation from the charter after database lock requires an escalated, documented waiver — precisely because ad hoc threshold-moving after seeing unfavorable data is the single most common way a low-quality asset survives a gate it should have failed.
Data Readout at the Gate
The gate itself is a moment, not a process: the trial database locks, the statisticians unblind, and every pre-specified criterion is evaluated against the actual observed data in a single, largely mechanical pass. The interesting judgment calls should already be over by this point — this stage is about honest measurement, not interpretation.
- Simultaneously: Criteria evaluated (against locked thresholds)
- Independent stats: Analysis performed by (blinded to program politics)
- 0: Threshold changes allowed (post-unblinding)
- 1–3 weeks: Typical readout-to-committee lag (for scorecard preparation)
From locked charter to scored criterion
At readout, each pre-specified criterion is converted into a simple pass/fail (or sometimes a three-way pass/marginal/fail) score by comparing the observed statistic to the pre-registered threshold. This is deliberately mechanical: the analysis plan specifies exactly which dataset, which population, and which statistical test to use, so that the scoring step involves no discretion. A biostatistics function independent of the program team typically performs this scoring, both to avoid bias and to create an auditable record that the criteria were applied as written.
The output is a scorecard: a short table listing each criterion, its threshold, the observed value, and a pass/fail flag. This scorecard — not a narrative slide deck built after the fact — is the primary document the gate committee reviews. Building the narrative interpretation before the scorecard exists is exactly the failure mode pre-specification is designed to prevent.
Threshold stringency and the shape of risk
How aggressively a threshold is set upstream directly determines what kind of error the gate is prone to downstream. A lenient threshold (calibrated to historical base rates for similar assets) advances more programs but tolerates a higher false-positive rate — assets that will ultimately fail in a larger, more expensive trial get funded to that next stage. A strict threshold rejects more programs early, which protects capital but risks discarding assets that would have worked with a larger sample size or a better-selected population.
Neither setting is inherently correct; the right stringency is a function of the cost asymmetry at that specific gate. Early discovery gates, where the cost of being wrong is low, typically tolerate lenient thresholds designed to keep optionality open. Late-stage gates immediately preceding a pivotal trial — where a false Go can cost hundreds of millions of dollars — are set deliberately strict, sometimes intentionally above the historical base rate needed for eventual regulatory success, precisely to build in a margin against the optimism that tends to accumulate in a program's own data.
A single-arm, uncontrolled readout can hit almost any threshold you like by chance alone if enough post-hoc subgroups are examined — which is why the pre-specified analysis population and single primary comparison, not the most favorable cut of the data, are what the gate is scored against.
Sample gate criteria scorecard (Standard stringency)
| Product | Indication | Trial Design | Key Result |
|---|---|---|---|
| Primary Efficacy Endpoint | Threshold: ≥52 | Observed: 61 (ITT population, primary timepoint) | PASS |
| Safety Margin | Threshold: ≥57 | Observed: 64 (100 − Grade≥3 AE rate) | PASS |
| Biomarker Response | Threshold: ≥49 | Observed: 41 (% target-engaged responders) | FAIL |
| PK Target Attainment | Threshold: ≥55 | Observed: 58 (% time above MIC) | PASS |
| Patient-Reported Outcome | Threshold: ≥50 | Observed: 44 (PRO composite delta) | FAIL |
Cross-Functional Gate Committee Review
A scorecard tells the committee what the data says; it does not tell the committee what to do. That judgment is deliberately assigned to a cross-functional committee — clinical, regulatory, commercial, and manufacturing leads — because a program that clears every statistical threshold can still be the wrong bet, and a program that misses one can still be worth continuing.
- 4+: Standing committee functions (clinical, reg, commercial, mfg)
- Before discussion: Independent scoring (to avoid anchoring)
- Data + strategic fit: Dimensions scored (not criteria alone)
- Standing monthly board: Typical review cadence (plus ad hoc gate sessions)
Why the committee is cross-functional at all
A trial statistician can tell you whether a p-value crossed a pre-specified line. Only a regulatory lead can tell you whether the specific pattern of misses is the kind that a health authority will accept with a confirmatory study versus the kind that predicts an outright rejection. Only a commercial lead can tell you whether the observed effect size, even if statistically significant, is large enough to support a differentiated label against a competitor asset that will likely reach the market first. Only a manufacturing lead can tell you whether the process that produced clinical-trial material is scalable to commercial volumes, or whether a technically successful trial was run on batches that cannot be reproduced at scale.
No single function has visibility into all of these considerations, which is exactly why the review is structured as an independent, cross-functional vote rather than a single decision-maker's call. Each function is expected to bring a perspective the others structurally cannot.
Independent scoring before group discussion
The best-run gate committees require every member to submit an independent go/no-go/hold assessment before the group discussion begins, precisely to prevent anchoring: if the most senior or most vocal person in the room states their view first, subsequent speakers tend to converge toward it regardless of their private assessment, a well-documented group dynamic that silently collapses the diversity of judgment the cross-functional structure was designed to preserve.
Independent pre-scores are then aggregated — sometimes as a simple vote count, sometimes as a weighted composite that accounts for the fact that a safety concern from the clinical lead should carry more weight than a commercial preference — and the resulting spread (unanimous vs. split) is itself informative. A program with a 4-0 committee vote and a program that passed 3-1 on a contested manufacturing objection are not equally "Go" decisions, even if both technically cross the same bar.
A split committee vote is not a procedural failure to be smoothed over — it is signal. Documenting the dissent and the specific reservation that produced it is often more useful to future decision-makers than the headline decision itself.
Strategic fit is a real, separate dimension
Strategic fit asks a different question than "did the trial succeed": does this asset still belong in the portfolio, given everything the organization has learned since the program was funded? A program can pass every pre-specified criterion and still be a weak candidate for continued investment if a competitor has meanwhile reached the market with a superior profile, if the target population has shrunk due to a diagnostic or standard-of-care shift, or if the organization's strategic priorities have moved toward a different modality or therapeutic area entirely.
Conversely, a program that narrowly misses a criterion may still merit continued investment if it is first-in-class for an area of high unmet need with no credible competitive threat, if the miss is plausibly attributable to a fixable trial-execution issue rather than a biology failure, or if it anchors a broader platform whose value extends beyond the single asset. This is precisely the judgment a scorecard alone cannot render, and precisely why the cross-functional review sits downstream of, not instead of, the quantitative gate.
Go / No-Go / Hold Decision
The committee's independent scores and the criteria scorecard are synthesized into one of three outcomes. Framing the decision as strictly binary — advance or kill — is a common governance mistake; the third option, Hold, exists precisely for the frequent case where the data is genuinely ambiguous rather than clearly favorable or clearly unfavorable.
- 3: Decision outcomes (Go / No-Go / Hold)
- ~30–50%: Typical Go rate at late gates (varies heavily by therapeutic area)
- Mandatory: Hold re-review deadline (defined at the time of the Hold)
- Portfolio system of record: Decision documented in (with rationale, not just outcome)
Three outcomes, three very different actions
A Go decision authorizes the program to advance into its next phase of investment: budget is committed, headcount is allocated or retained, and typically a new set of pre-specified criteria is drafted for the next gate before the current one even closes. A No-Go decision terminates further investment in the program in its current form — this does not necessarily mean the underlying science is discarded (a killed clinical program can still yield a valuable biomarker, a manufacturing process, or a mechanistic insight that informs a different asset), but the specific investment thesis being tested at this gate is closed.
A Hold decision defers the choice: it neither commits new resources at full scale nor releases the program's people and budget back to the portfolio. Instead it specifies a narrow, defined path — usually one or two additional pieces of evidence — that would resolve the ambiguity, along with a hard deadline for producing that evidence and returning to committee.
The Hold trap: how a temporary pause becomes a zombie project
Hold is the most operationally dangerous of the three outcomes, because without strict governance it quietly becomes a fourth, unintended category: an underfunded program that is neither reviewed rigorously nor formally killed, consuming a trickle of resources indefinitely while contributing nothing decisive to the portfolio. This is the "zombie project" failure mode, and it is extremely common in organizations that treat Hold as a way to avoid an uncomfortable No-Go conversation rather than as a genuine, time-boxed evidence-gathering pause.
Well-governed portfolios prevent this by attaching two things to every Hold decision at the moment it is made, not afterward: a specific, falsifiable re-review criterion (what new evidence resolves the ambiguity, and what threshold applied to that evidence will produce a Go or a No-Go next time) and a hard calendar deadline after which the program automatically escalates to a forced decision, with a default outcome of No-Go if the deadline passes without the additional evidence being produced. Absent both, "Hold" functions as a permanent stay of execution and defeats the entire purpose of gated portfolio governance.
A Hold without a pre-committed re-review date and a pre-committed default outcome is not a decision at all — it is a deferred decision, and deferred decisions are how weak programs quietly consume years of budget without ever being formally evaluated.
Post-Decision Portfolio Impact
A gate decision is not the end of the process — it is the trigger for a resource-allocation cascade across the entire portfolio. How disciplined an organization is about actually executing that cascade, rather than letting it happen slowly or incompletely, is one of the clearest signals of whether its stage-gate system is real governance or theater.
- Reserved: Go → next-phase budget (at time of decision, not after)
- Within one cycle: No-Go → resources released (to the free portfolio pool)
- Frozen: Hold → resources (pending fixed re-review date)
- Highest-ranked backlog asset: Recycled capital destination (not the nearest open program)
A Go reserves capacity — and creates the next charter
When a program receives a Go, the immediate operational task is converting that decision into a committed budget line and a staffed team for the next phase, ideally before the enthusiasm of a successful readout fades into ordinary competition for scarce resources. Mature portfolio processes require the next gate's criteria to be drafted and circulated for sign-off within a fixed window of the Go decision — for the same reason criteria are pre-specified in the first place: setting the next bar while still close to a position of relative objectivity, before the team has accumulated a fresh set of favorable assumptions about the program's trajectory.
A No-Go's real value is the capital and people it frees
The financial discipline of a portfolio is tested far more by what happens after a No-Go than by the decision itself. Killing a program is not primarily valuable because it stops a specific line item of spend — it is valuable because it releases scarce, specialized capacity (a clinical operations team, a manufacturing slot, a regulatory affairs headcount allocation) back into a pool where it can be redeployed to the organization's highest-ranked remaining asset, rather than left idle or, worse, quietly reabsorbed by the very program that was just killed under a different budget code.
Organizations that are disciplined about recycling — releasing the freed capital and headcount within a defined cycle and reallocating it via the same ranked-backlog process used for all other capital requests, rather than informally to whichever team asks first — extract substantially more portfolio value per dollar than organizations that treat a No-Go purely as a stopping decision and neglect the redeployment half of the equation.
The discipline of resource recycling after a No-Go is arguably a bigger driver of overall portfolio return than the accuracy of any individual gate decision, because it determines whether freed capital compounds into the next-best opportunity or simply evaporates into organizational overhead.
A Hold freezes, but freezing is not free
Even a well-governed Hold has a real opportunity cost: the specialized people and reserved capacity attached to a held program are typically not fully redeployable to other work during the hold period, because pulling them away would undermine the very re-review the hold is meant to enable. This is why the hard deadline attached to every Hold matters financially, not just procedurally — every month a program sits in Hold beyond its intended window is a month of capacity that is neither advancing the held program at full pace nor available to the rest of the portfolio, a cost that compounds silently if Hold deadlines are allowed to slip without consequence.
This simulation evaluates the criteria for deciding whether to continue or terminate a project at critical decision points, ensuring resource efficiency and strategic alignment.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install