🖼 Companion Algorithm for IHC Biomarker Quantification
This tool assists pathologists in selecting the most appropriate therapy by quantitatively assessing IHC biomarker expression levels.
Immunohistochemistry — Turning a Molecular Target Into a Visible Signal
Companion diagnostic biomarker testing begins with immunohistochemistry (IHC), a decades-old technique that remains the clinical gold standard for visualizing protein expression directly in tissue context. A formalin-fixed, paraffin-embedded (FFPE) tumor section is stained so that the target protein — HER2 on the breast/gastric cancer cell membrane, or PD-L1 on tumor and immune cells — becomes visible as a brown chromogen deposit under brightfield microscopy, with intensity roughly proportional to protein abundance.
- DAB chromogen: Detection method (3,3'-diaminobenzidine, brown precipitate)
- 4B5 / 22C3 / SP263: Antibody clones (examples) (HER2 and PD-L1 pDx clones)
- 4–8 hours: Turnaround time (fixation to stained slide)
- 0.25 µm/px: Slide scan resolution (40× equivalent whole-slide scan)
The IHC staining chemistry pipeline
A companion diagnostic IHC assay follows a tightly controlled, FDA-cleared staining protocol executed on automated staining platforms (Ventana BenchMark, Dako/Agilent Autostainer Link) to minimize the run-to-run variability that would otherwise undermine a quantitative test:
1. Tissue fixation and sectioning: tumor tissue is fixed in 10% neutral-buffered formalin for a validated window (typically 6–72 hours — under- or over-fixation both distort antigen detectability), embedded in paraffin, and cut into 4–5 µm sections mounted on adhesive slides
2. Deparaffinization and antigen retrieval: paraffin is removed with xylene/alcohol series; heat-induced epitope retrieval (HIER) in citrate or EDTA buffer at ~95–100°C reverses formalin-induced protein cross-linking that would otherwise mask the antibody-binding epitope
3. Primary antibody incubation: a validated monoclonal antibody (e.g., clone 4B5 for HER2, clone 22C3 or SP263 for PD-L1) binds specifically to the target protein epitope. Companion diagnostic assays use a single, regulatory-locked antibody clone, lot-controlled and paired to a specific staining platform — substituting a different "equivalent" antibody voids the companion diagnostic claim
4. Signal amplification and detection: a secondary antibody or polymer-based detection system conjugated to horseradish peroxidase (HRP) binds the primary antibody; multiple HRP molecules cluster at each binding site for signal amplification
5. Chromogen development: DAB (3,3'-diaminobenzidine) substrate is oxidized by HRP to form an insoluble brown precipitate exactly where the target antigen is located — critically, for membrane-bound targets like HER2, the precipitate should trace the cell membrane outline rather than diffusely filling the cytoplasm
6. Counterstaining: hematoxylin lightly stains all nuclei blue, providing morphological context for the pathologist and algorithm to distinguish tumor cells from stroma and immune infiltrate
7. Digital scanning: the stained slide is digitized on a validated whole-slide scanner at 0.25 µm/pixel (40× equivalent), producing the gigapixel image the quantification algorithm will analyze
Because DAB intensity is what the entire downstream scoring pipeline depends on, even small deviations in antibody concentration, incubation time, or HRP amplification kinetics can shift a borderline case from one score category to another — which is precisely why companion diagnostic assays lock the antibody clone, staining platform, and protocol together as a single regulated unit rather than certifying the antibody alone.
Computer Vision Quantification — Measuring Intensity and Completeness Per Cell
Where a pathologist visually estimates staining pattern across a slide in seconds, the companion algorithm performs the same assessment computationally and exhaustively: every individual tumor cell membrane is segmented, and its staining intensity and circumferential completeness are measured as continuous, reproducible numerical quantities before any categorical score is assigned.
- 1,000–10,000+: Cells segmented per slide (tumor cell nuclei/membranes)
- 0–100%: Membrane completeness metric (circumference stained)
- Optical density (OD): Intensity metric (DAB absorbance, continuous scale)
- ~2–5 min: Processing time per slide (GPU-accelerated pipeline)
Cell segmentation and membrane feature extraction
The quantification pipeline operates in two computational stages, mirroring the way a trained pathologist reads a slide but with pixel-level precision:
Stage A — Nuclear and cell segmentation: • A deep convolutional network (typically a U-Net or Mask R-CNN variant trained on thousands of pathologist-annotated nuclei) detects and segments individual hematoxylin-stained nuclei • A membrane boundary is estimated around each nucleus using a combination of learned cell-shape priors and local DAB-signal gradients, since the cell membrane itself is often only faintly visible except where stained • Tumor cells are distinguished from stromal cells, lymphocytes, and normal epithelium using morphological and staining-context features — critical because only tumor cell membrane staining counts toward the diagnostic score, and misclassifying stroma as tumor would corrupt the result
Stage B — Per-cell staining quantification: • Color deconvolution (Ruifrok & Johnston method) separates the composite RGB image into pure hematoxylin and pure DAB optical density channels, since the two stains overlap spectrally in the raw image • For each segmented cell membrane, the algorithm measures: - Intensity: mean DAB optical density along the membrane contour, on a continuous scale later mapped to none/weak/moderate/strong - Completeness: the percentage of the membrane's circumference that shows continuous, connected staining above a detection threshold, versus staining that is punctate or restricted to one side of the cell • These two continuous measurements — intensity and completeness — are exactly the two dimensions a pathologist evaluates visually under ASCO/CAP guidelines, but here quantified per cell rather than impressionistically across the whole field
Because the algorithm evaluates every tumor cell on the slide rather than a pathologist's sampled visual fields, its output is a full population distribution of per-cell scores rather than a single holistic impression — this distribution is what feeds the histogram used in score-category assignment in Stage 3.
Binning Cells Into 0 / 1+ / 2+ / 3+ and Checking Pathologist Concordance
Continuous per-cell intensity and completeness measurements must ultimately collapse into the discrete score categories that clinicians act on. The ASCO/CAP (American Society of Clinical Oncology / College of American Pathologists) HER2 testing guidelines — and analogous tumor-proportion-score frameworks for PD-L1 — define exactly how membrane staining pattern translates into a 0, 1+, 2+, or 3+ call, and the algorithm must reproduce this logic with high fidelity to a trained pathologist's independent read.
- No staining: Score 0 (or membrane staining in <10% of cells)
- Faint/barely perceptible: Score 1+ (incomplete membrane, >10% of cells)
- Weak-to-moderate complete: Score 2+ (membrane staining, >10% of cells — reflex FISH)
- Strong complete: Score 3+ (membrane staining, >10% of cells)
ASCO/CAP scoring criteria and the 2+ reflex-testing rule
The ASCO/CAP HER2 IHC scoring system (as applied, for example, to guide trastuzumab eligibility) defines four categories based jointly on staining intensity and the completeness/percentage of positive tumor cells:
• Score 0: no staining observed, or incomplete, faint/barely perceptible membrane staining in ≤10% of tumor cells • Score 1+: incomplete, faint/barely perceptible membrane staining in >10% of tumor cells — considered HER2-negative • Score 2+: weak-to-moderate complete membrane staining in >10% of tumor cells (equivocal) — OR circumferential, intense staining in ≤10% of cells; this category cannot be called definitively positive or negative from IHC alone • Score 3+: circumferential, complete, intense ("strong") membrane staining in >10% of tumor cells — considered HER2-positive, sufficient on its own for therapy eligibility
The critical clinical decision point is the 2+ category: it is deliberately equivocal by design, and ASCO/CAP guidelines mandate reflex testing by fluorescence in situ hybridization (FISH) to directly assess HER2 gene copy number/amplification ratio before a definitive positive/negative call can be made. This two-tier system (IHC as primary screen, FISH as tiebreaker) exists precisely because IHC intensity alone cannot always reliably distinguish moderate overexpression from true gene amplification.
For PD-L1 companion diagnostics (e.g., the 22C3 pharmDx assay paired with pembrolizumab), the analogous quantity is the Tumor Proportion Score (TPS) — the percentage of viable tumor cells showing any membrane staining at any intensity — with clinically actionable cutoffs typically at TPS ≥1% and TPS ≥50%, rather than the four-tier HER2 scale.
Algorithm-to-pathologist concordance as the core validation metric
An algorithm's slide-level score is derived by aggregating its per-cell measurements according to the same percentage-and-pattern logic pathologists use (e.g., "what fraction of cells show complete, strong membrane staining"), then compared against the independent score assigned by a board-certified pathologist reading the same slide under standard bright-field microscopy, blinded to the algorithm's output.
Concordance is typically reported as: • Overall percent agreement (OPA): fraction of cases where algorithm and pathologist assign the identical 0/1+/2+/3+ category • Positive/negative percent agreement (PPA/NPA): agreement specifically on the clinically actionable positive (3+) versus negative (0/1+) calls, since 2+ cases are already routed to FISH regardless • Cohen's kappa: chance-corrected agreement statistic, with κ>0.8 generally considered excellent inter-rater agreement in pathology
Regulatory-grade companion diagnostic algorithms typically target and achieve OPA in the low-to-mid 90% range against expert pathologist consensus reads, with the majority of residual discordance concentrated at score-category boundaries (e.g., a case genuinely intermediate between 1+ and 2+) rather than gross misclassification — reflecting genuine biological/staining ambiguity that both human and algorithmic readers struggle with equally.
Analytical Validation — Proving the Algorithm Gives the Same Answer Every Time
Before an algorithm can be trusted to guide a treatment decision, it must demonstrate that its output is not merely accurate on average but reproducible — that the same tissue, scored on different scanners, by different laboratory operators, on different days, at different sites, yields the same clinical call. This is analytical validation, and it is evaluated independently of whether the score correlates with clinical outcome.
- <5%: Intra-run precision (CV%) (repeated runs, same slide/scanner)
- >90%: Inter-scanner concordance (category-level agreement)
- >85%: Inter-site reproducibility (multi-laboratory study)
- 3+ platforms: Scanners validated (typical) (e.g., Aperio, Philips, Hamamatsu)
Precision and reproducibility study design
Analytical validation for a digital pathology companion algorithm follows a structured precision study design closely modeled on clinical chemistry assay validation, adapted for image-based measurement:
• Repeatability (intra-run precision): the identical digitized slide image is re-analyzed by the algorithm multiple times to confirm the algorithm itself is deterministic (or, for any stochastic components, that variability is below a defined tolerance) — this isolates pure computational reproducibility from any physical/optical source of variation
• Intermediate precision (same site, different day/operator): the same physical slide is re-scanned on different days, sometimes by different laboratory technicians operating the scanner, to capture normal day-to-day laboratory variability in focus, illumination calibration, and slide handling
• Reproducibility (inter-site, inter-scanner): a shared slide set is distributed to multiple laboratory sites, each digitizing on their own scanner hardware (which may be a different manufacturer/model — e.g., Aperio GT450, Philips IntelliSite, Hamamatsu NanoZoomer) — this is the most stringent test, since scanner optics, color calibration, and z-stack focusing algorithms genuinely differ between platforms and can shift measured staining intensity
• Operator variability: different histotechnologists performing the physical staining protocol (even on an automated stainer) introduce subtle reagent-handling and timing variation that precision studies must capture
Results are reported as coefficient of variation (CV%) for continuous intermediate measurements, and as categorical percent agreement for the final 0/1+/2+/3+ calls — since a clinician ultimately acts on the category, not the underlying continuous score, a well-validated algorithm can tolerate modest continuous-score CV while still delivering high categorical reproducibility, provided variability does not cluster near a decision threshold.
FDA guidance for digital pathology companion diagnostics generally expects inter-site categorical concordance above 85–90% and intra-laboratory CV below 10% on continuous intensity metrics before an algorithm is considered analytically validated — thresholds modeled on, but distinct from, the precision requirements applied to quantitative clinical chemistry assays.
Linking Algorithm Score to Real Patient Outcomes
Analytical reproducibility alone does not establish clinical usefulness — a perfectly reproducible score is worthless if it does not predict which patients actually benefit from the paired therapy. Clinical validation closes this gap by correlating algorithm-derived scores, applied retrospectively or prospectively to trial-cohort slides, with the objective response rate and progression-free survival data collected during the therapeutic drug's pivotal clinical trial.
- 500–3,000+: Trial cohort size (typical pDx) (patients with matched slides)
- ORR / PFS / OS: Endpoint correlated (objective response, survival)
- ROC analysis: Cutoff optimization method (sensitivity/specificity trade-off)
- ~0.4–0.6: Hazard ratio at optimal cutoff (biomarker-positive vs. negative)
Establishing the clinically actionable cutoff
Clinical validation of a companion diagnostic algorithm answers a specific regulatory question: at what biomarker score threshold does a patient meaningfully benefit from the paired therapy, versus a comparator or standard of care?
The validation process typically proceeds as:
1. Retrospective slide re-analysis: archived tumor slides from a completed or ongoing pivotal trial of the therapeutic drug are digitized and scored by the algorithm, blinded to the trial's treatment-arm assignment and clinical outcome data
2. Outcome correlation: algorithm scores are statistically correlated against the trial's primary and secondary endpoints — objective response rate (ORR), progression-free survival (PFS), overall survival (OS) — stratified by treatment arm
3. Cutoff optimization: receiver operating characteristic (ROC) curve analysis identifies the score threshold that best separates patients who respond to therapy from those who do not, balancing sensitivity (capturing true responders) against specificity (excluding true non-responders who would face therapy toxicity without benefit)
4. Hazard ratio confirmation: a Cox proportional-hazards model confirms that patients above the chosen cutoff show a statistically and clinically meaningful hazard ratio for progression or death relative to patients below it — the biomarker-positive population must show a treatment effect the biomarker-negative population does not
This is precisely the validation pathway followed historically for real companion diagnostics — HER2 IHC/FISH testing establishing trastuzumab eligibility in the original Herceptin pivotal trials, and PD-L1 22C3 IHC testing establishing pembrolizumab eligibility using the Tumor Proportion Score cutoffs of ≥1% and ≥50% — in each case, the diagnostic cutoff was derived directly from, and co-validated against, the therapeutic trial's outcome data rather than chosen independently.
Why clinical and analytical validation must both hold
A companion diagnostic algorithm must satisfy both validation domains simultaneously, and neither substitutes for the other:
• An analytically perfect but clinically uncorrelated score would reproducibly measure something — but not necessarily therapy benefit • A clinically correlated but analytically unreliable score might show impressive retrospective statistics on curated trial data, yet fail to reproduce that correlation when deployed across the heterogeneous scanners, staining lots, and operators of real-world clinical laboratories
Regulatory submissions therefore present both validation packages jointly: the analytical validation study (Stage 4) establishing that the measurement itself is trustworthy, and the clinical validation study (this stage) establishing that the measurement, once trustworthy, actually identifies patients who benefit — together forming the complete evidentiary basis for a companion diagnostic claim.
The Regulatory Pathway — Co-Approval With the Paired Therapeutic
A companion diagnostic is not approved as a standalone device — it is reviewed and cleared by the FDA in tandem with the specific therapeutic drug it is designed to guide, typically through the Premarket Approval (PMA) pathway, with the diagnostic's label explicitly cross-referencing the drug's label and vice versa. This co-dependent regulatory structure is what legally links a biomarker test to a treatment decision.
- PMA: Regulatory pathway (Premarket Approval, Class III device)
- Drug + Diagnostic: Co-review target (concurrent FDA divisions)
- IHC/FISH + trastuzumab: Real precedent (HER2) (Herceptin, 1998)
- 22C3 pDx + pembrolizumab: Real precedent (PD-L1) (Keytruda, 2015)
The FDA companion diagnostic co-approval framework
The FDA formalized companion diagnostic regulation through guidance issued in 2014 ("In Vitro Companion Diagnostic Devices"), building on precedent set by earlier drug-diagnostic pairs. The core regulatory principle: if safe and effective use of a therapeutic depends on results from a specific diagnostic test, that diagnostic must be reviewed and approved essentially simultaneously with the drug, and the drug label should identify the specific, named companion diagnostic (not "any validated HER2 test") required for prescribing.
The review pathway proceeds through several structural elements:
• Premarket Approval (PMA): companion diagnostics are Class III devices, the FDA's highest-risk category, requiring the most rigorous premarket review — full analytical and clinical validation data packages (as built in Stages 4 and 5) are submitted, alongside manufacturing quality-system documentation
• Concurrent review coordination: FDA's device center (CDRH) and drug center (CDER) coordinate review timelines so the companion diagnostic and therapeutic can be approved on the same day or in close sequence — historically a major logistical challenge, since diagnostic and drug development timelines rarely align naturally
• Label cross-referencing: the drug's package insert specifies the named, cleared companion diagnostic test (by manufacturer, assay name, and antibody clone) required to identify eligible patients; the diagnostic's label states the specific therapeutic and score cutoff it supports
• Real-world precedents: HER2 testing (originally HercepTest IHC and PathVysion FISH) was co-approved alongside trastuzumab (Herceptin) in 1998, establishing the template for the entire companion diagnostic category; PD-L1 IHC 22C3 pharmDx was co-approved alongside pembrolizumab (Keytruda) in 2015 for non-small-cell lung cancer, later expanded to additional PD-L1-dependent indications with the same assay
Once a digital pathology algorithm is intended to replace or assist a pathologist's manual IHC read as part of a companion diagnostic claim, it must clear the same PMA bar as the underlying assay itself — including full analytical and clinical validation — because the algorithm's output, not just the stain, is now the regulated diagnostic result that determines treatment eligibility.
This tool assists pathologists in selecting the most appropriate therapy by quantitatively assessing IHC biomarker expression levels.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install