Multi-gene pharmacogenomic genotyping performed once, before any drug is ever prescribed — CYP2D6, CYP2C19, CYP2C9, VKORC1, TPMT/NUDT15, DPYD, SLCO1B1, HLA-B — filed in the EHR for lifelong point-of-care use
The defining feature of preemptive pharmacogenomic (PGx) testing is timing: DNA is collected at a routine primary-care or pediatric well visit, during a hospital admission unrelated to drug therapy, or even prenatally in some pilot programs — never at the moment a clinician is trying to decide what dose of clopidogrel or codeine to prescribe. This decouples genotyping turnaround time from the urgency of the clinical decision, so results are already sitting in the chart, fully annotated, the first time they are needed.
Traditional pharmacogenomic testing is reactive: a patient has an adverse drug reaction, or a clinician suspects one is likely, and orders a single-gene test — CYP2C19 before clopidogrel, TPMT before azathioprine — that takes 3–10 days to return. By the time the result is available, the drug decision has usually already been made empirically.
Preemptive panel testing inverts this sequence entirely: • Testing occurs once, at a health-system touchpoint unrelated to any specific drug (annual physical, hospital admission, pediatric visit, or direct-to-consumer/health-system partnership programs) • A single multi-gene panel (not one gene) is run, covering essentially all CPIC Level A gene-drug pairs simultaneously • Results are stored as structured, computable data in the electronic health record (EHR) — not a scanned PDF • The genotype is queried automatically by clinical decision support (CDS) the moment any relevant drug is ordered, at any point over the following decades
Because DNA sequence does not change over a lifetime (with rare exceptions such as allogeneic bone marrow transplant, which can alter blood-derived genotype), a single test performed in childhood remains clinically valid into old age — turning a one-time $150–250 laboratory cost into a resource reused dozens of times across a lifetime of prescribing.
Two collection modalities dominate large preemptive programs:
• Buccal swab: non-invasive, patient-collected or clinician-administered, yields 2–20 µg of DNA of variable quality; preferred for pediatric and outpatient screening programs (e.g., St. Jude PG4KDS) • Peripheral blood (2 mL EDTA tube): higher DNA yield and purity, preferred when the panel is added onto an existing clinical blood draw to avoid an extra patient encounter
Extraction is performed on automated magnetic-bead platforms (QIAsymphony, Chemagic 360, MagMAX) processing 96 or 384 samples per batch: • Cell lysis: proteinase K + chaotropic salt buffer disrupts membranes and denatures nucleases • Magnetic silica bead binding: DNA adsorbs to bead surface under high-salt conditions • Wash steps: remove protein, salt, and PCR inhibitors while DNA remains bound • Elution: low-salt buffer releases purified genomic DNA, typically 200 ng–2 µg per swab, up to 10–20 µg per blood tube
Quality control gates every sample before it proceeds to genotyping: • NanoDrop A260/280 ratio 1.8–2.0 (protein contamination flagged outside this range) • Qubit dsDNA fluorometric quantification (more accurate than absorbance for low-yield samples) • Gel or capillary electrophoresis fragment sizing — degraded DNA (<10 kb median fragment) is often re-collected, since structural-variant calling for CYP2D6 requires reasonably intact template
Rather than ordering a separate laboratory-developed test for each gene-drug pair, preemptive programs run a single multiplexed assay that simultaneously interrogates every CPIC/DPWG actionable pharmacogene. Two technology platforms dominate: SNP microarrays purpose-built for pharmacogenomics, and targeted next-generation sequencing panels — each with different strengths for the structurally complex loci that make PGx genotyping harder than routine SNP genotyping.
SNP microarrays remain the workhorse of large preemptive programs because a single chip can be manufactured with tens of thousands of pharmacogenomically relevant probes and run for a marginal cost of $40–100 per sample when batched at scale:
• Thermo Fisher PharmacoScan: 4,627 variants across 1,191 genes, including dense coverage of all CPIC Level A/B genes plus ADME (absorption/distribution/metabolism/excretion) gene families • Illumina Infinium Global Screening Array + PGx content: combines genome-wide SNP backbone with curated pharmacogene add-on content, useful when PGx is bundled with broader genomic screening • Allele-specific hybridization probes distinguish single-nucleotide variants with high confidence (>99.9% concordance to sequencing for well-behaved biallelic SNPs) • Copy-number probes at closely spaced intervals across CYP2D6 allow relative signal-intensity comparison to reference samples, enabling CNV/deletion/duplication calling directly from array intensity data
Limitation: arrays genotype only pre-selected, known variants — a rare or novel allele not represented on the probe set will be silently missed or miscalled as the nearest reference allele.
Targeted next-generation sequencing panels (PGRNseq, VeriDose Core Panel, or custom amplicon/hybrid-capture designs) sequence the full genomic region of each pharmacogene at 100–150× depth, rather than only interrogating pre-selected SNP positions:
• Captures novel and rare variants missed by fixed-content arrays • Provides base-level read data needed for haplotype phasing (which variants sit on the same physical chromosome) • Essential for CYP2D6, the single hardest pharmacogene to genotype accurately: • CYP2D6 sits adjacent to a highly homologous pseudogene, CYP2D7, and frequent unequal crossover between them generates gene deletions (*5), duplications (up to 4+ copies), and CYP2D6/CYP2D7 hybrid alleles (*36, *68) that standard variant callers systematically misclassify • Long-read or read-depth-based CNV algorithms (Cyrius, Stargazer) are required to resolve true copy number before star alleles can even be assigned • HLA-A and HLA-B genotyping for abacavir hypersensitivity (HLA-B*57:01) and carbamazepine-induced Stevens-Johnson syndrome (HLA-B*15:02) requires 4-digit resolution typing algorithms (e.g., HLA-LA, OptiType) distinct from ordinary SNP genotyping, because of the extreme polymorphism and sequence similarity among HLA alleles
Most production pipelines combine both technologies: an array for cost-effective genome-wide SNP calling, layered with targeted deep sequencing specifically for CYP2D6 and HLA loci.
Raw genotype or sequencing-read data is clinically meaningless until it is translated into the standardized "star allele" (haplotype) nomenclature that pharmacogenomics guidelines are written against. Specialized calling software — not generic variant callers — resolves phased haplotypes, structural variants, and combinations of variants into named alleles such as CYP2D6*4 or CYP2C19*17, then pairs the two inherited haplotypes into a diplotype.
A diploid human carries two copies of every autosomal gene, one from each parent. Genotyping alone tells you which variant positions are heterozygous, but not which combinations of variants sit on the same physical chromosome (haplotype, "in cis") versus opposite chromosomes ("in trans") — and this distinction changes the clinical interpretation completely.
Example: if a patient is heterozygous at two CYP2D6 positions that individually define *4 and *10, the true diplotype could be *4/*10 (each haplotype carries one variant) or *1/*4+*10-in-cis (both variants on one chromosome, forming a rare combined allele, with a fully functional *1 on the other) — these imply very different enzyme activity.
Phasing methods: • Statistical phasing: population haplotype reference panels (1000 Genomes, HRC) combined with algorithms (Beagle, SHAPEIT) infer the most probable phase from population linkage patterns — fast and cheap but occasionally wrong for rare combinations • Read-based phasing: for NGS data, sequencing reads or read-pairs that physically span two variant positions directly observe which alleles co-occur — the gold standard when read length/insert size is sufficient • Long-read sequencing (PacBio HiFi, Oxford Nanopore) increasingly used as an orthogonal confirmation method for ambiguous CYP2D6 structural calls, since a single long read can span the entire ~4.3 kb gene plus flanking CYP2D7 homology region
General-purpose variant callers (GATK HaplotypeCaller, DeepVariant) are excellent for typical SNPs and small indels but perform poorly on the structurally complex loci that dominate clinical pharmacogenomics. Dedicated PGx callers close this gap:
• Stargazer: combines SNP/indel genotypes with read-depth-based CNV detection to call star alleles for CYP2D6 and >50 other pharmacogenes; explicitly models gene deletions, duplications, and CYP2D6/CYP2D7 hybrid structures • Aldy: uses integer linear programming to find the most parsimonious combination of star alleles explaining observed sequencing reads, robust to complex structural rearrangements • PyPGx: open-source Python pipeline supporting both array and NGS input, with published concordance >99% against validated reference samples for core CYP genes • Astrolabe / Constellation: rules-engine callers mapping curated variant combinations directly to PharmVar-defined star alleles
Every caller is benchmarked against Genetic Testing Reference Material (GeTRM) samples characterized by the CDC and Association for Molecular Pathology, and clinical laboratories participating in College of American Pathologists (CAP) proficiency testing must demonstrate concordant star-allele calls on blinded reference samples before reporting patient results.
A diplotype such as CYP2D6*1/*4 is not directly actionable at the bedside; it must first be translated into a functional phenotype — Poor, Intermediate, Normal, Rapid, or Ultrarapid Metabolizer — using standardized activity-score tables published and continuously updated by the Clinical Pharmacogenetics Implementation Consortium (CPIC) and the Dutch Pharmacogenetics Working Group (DPWG). This translation step is what makes a genotype clinically interpretable by a prescriber with no genetics training.
CPIC assigns each individual star allele a function value based on the aggregate evidence for how much residual enzyme activity it confers relative to a fully functional reference allele:
• 0 — no function (e.g., CYP2D6*4, a common splice-defect null allele) • 0.25 or 0.5 — decreased function (e.g., CYP2D6*10, *41 — reduced but non-zero activity) • 1.0 — normal function (e.g., CYP2D6*1, *2) • 1.25–2.0+ — increased function, typically from gene duplication of a normal-function allele
The Activity Score (AS) is the sum of the two inherited allele values. For CYP2D6: • AS = 0 → Poor Metabolizer (PM) — e.g., *4/*4 • AS = 0.25–1.0 → Intermediate Metabolizer (IM) — e.g., *4/*10 • AS = 1.25–2.25 → Normal Metabolizer (NM) — e.g., *1/*2 • AS > 2.25, or ≥3 functional gene copies from duplication → Ultrarapid Metabolizer (UM)
Clinical consequence flips depending on whether the drug is a prodrug or an active drug cleared by the enzyme: • Codeine (prodrug, activated by CYP2D6 to morphine): PM patients get no analgesia; UM patients risk life-threatening morphine toxicity, including reported pediatric deaths after standard codeine doses • Tamoxifen (prodrug, activated to endoxifen by CYP2D6): PM patients have reduced breast-cancer treatment efficacy • Most other CYP2D6 substrates (active drug cleared by the enzyme): PM patients risk drug accumulation and toxicity at standard doses
Not every actionable pharmacogene fits a metabolic-activity model — several report as simple categorical risk instead:
• VKORC1: a promoter SNP (-1639G>A) alters warfarin sensitivity by changing vitamin K epoxide reductase expression, combined with CYP2C9 genotype in validated warfarin dosing algorithms (e.g., the International Warfarin Pharmacogenetics Consortium equation) rather than a standalone metabolizer class • TPMT / NUDT15: decreased-function alleles predict severe, potentially fatal myelosuppression from standard-dose thiopurines (azathioprine, mercaptopurine, thioguanine); reported as Normal/Intermediate/Poor Metabolizer analogous to CYP genes • DPYD: decreased-function alleles predict severe, sometimes fatal toxicity from fluoropyrimidine chemotherapy (5-fluorouracil, capecitabine); CPIC recommends 25–50% dose reduction or alternative therapy for intermediate metabolizers • SLCO1B1: the *5 (rs4149056) decreased-function allele impairs hepatic uptake of statins, raising the risk of simvastatin-induced myopathy — reported as Normal/Decreased/Poor Function rather than a metabolizer class • HLA-B*57:01 and HLA-B*15:02: not a metabolic enzyme at all — a T-cell-mediated hypersensitivity risk marker, reported simply as positive or negative, with a positive result representing an absolute contraindication to abacavir or carbamazepine respectively
The entire value proposition of preemptive pharmacogenomic testing depends on this final step working correctly, potentially decades after the sample was collected: structured genomic results must be discoverable by clinical decision support (CDS) the instant any actionable drug is ordered, regardless of which clinician, which department, or which decade of the patient's life it happens in. Get this integration right, and one $150–250 test quietly prevents adverse drug reactions across a lifetime of prescriptions.
For a PGx result to function years after it was generated, it cannot live as a scanned PDF in the document-imaging tab of the chart — the prescribing clinician will never see it. Modern implementations store results as discrete, computable data:
• HL7 FHIR Genomics resources (Observation, DiagnosticReport with genomic profiles) encode diplotype, phenotype, and evidence level in a structured, machine-queryable format • CDS Hooks standard allows the EHR's medication-ordering workflow to fire a real-time service call at the moment a drug is selected, querying the patient's stored genomic observations without the prescriber taking any extra action • Interruptive alert: when a clinician orders clopidogrel for a patient with a CYP2C19 poor/intermediate metabolizer genotype, the order screen displays a hard-stop or interruptive warning with the specific CPIC-recommended alternative (e.g., prasugrel or ticagrelor) and the supporting diplotype • Passive/informational alert: for lower-severity gene-drug pairs, a non-interruptive banner or sidebar notification avoids alert fatigue while still documenting the genotype was considered
Large academic implementations — Vanderbilt's PREDICT program, St. Jude's PG4KDS, the University of Chicago 1200 Patients Project, and the Mayo Clinic RIGHT protocol — have collectively demonstrated that this architecture scales to tens of thousands of patients with alert-firing latency under one second.
Consider a patient panel-tested at age 34 during a routine physical, found to be CYP2C19 Ultrarapid Metabolizer (via *17/*17), CYP2D6 Intermediate Metabolizer, SLCO1B1 decreased function, and HLA-B*57:01 negative:
• Age 34 (testing): result filed silently in the EHR; no clinical action required at this visit • Age 41 (ST-elevation MI, cardiac catheterization): cardiologist orders clopidogrel post-stenting — CDS alert does NOT fire, since this patient is actually a rapid CYP2C19 metabolizer with adequate clopidogrel activation (illustrating that alerts fire only for genotypes that change the recommended action) • Age 47 (elective hip surgery, postoperative pain): surgeon orders codeine — interruptive alert fires because of the CYP2D6 Intermediate Metabolizer result, recommending a non-CYP2D6-dependent analgesic instead • Age 55 (hyperlipidemia): primary care orders simvastatin 40 mg — alert recommends a lower starting dose or alternative statin given SLCO1B1 decreased function, reducing myopathy risk • Age 58 (new HIV diagnosis): infectious disease physician orders an antiretroviral regimen — CDS confirms HLA-B*57:01 negative, meaning abacavir remains a safe option with no additional testing needed
Four separate prescribing decisions across 24 years, each informed instantly by a single sample collected once — this is the operational logic that makes preemptive testing more cost-effective at a population level than reactive single-gene testing, despite testing many patients who will never need most of the genes on the panel.
A pragmatic randomized trial embedding preemptive CYP2C19 genotyping into post-PCI antiplatelet selection (TAILOR-PCI, published in JAMA 2020, n=5,302) showed a 34% relative reduction in major adverse cardiovascular events among CYP2C19 loss-of-function carriers whose therapy was genotype-guided versus those on standard clopidogrel — a benefit only realizable because the genotype was already available at the moment of the prescribing decision. Broader implementation cohorts across CPIC gene-drug pairs report 30–45% reductions in genotype-related adverse drug reactions, reinforcing that the clinical value of PGx testing is captured almost entirely at this final integration step, not at the sequencing bench.
| Product | Indication | Trial Design | Key Result |
|---|---|---|---|
| CYP2D6 | Codeine, tramadol, tamoxifen | Star-allele + CNV calling resolves activity score 0–≥3 | Prevents opioid toxicity/failure & tamoxifen under-activation |
| CYP2C19 | Clopidogrel, voriconazole, PPIs | *2/*3 (loss) vs. *17 (gain) diplotyping | Routes LoF carriers to prasugrel/ticagrelor post-PCI |
| TPMT / NUDT15 | Azathioprine, mercaptopurine | Decreased-function diplotype → dose reduction table | Averts life-threatening myelosuppression |
| HLA-B | Abacavir, carbamazepine | 4-digit HLA typing, categorical positive/negative | Absolute contraindication flag before first dose |