From saliva chip to prescribing decision — where DTC pharmacogenomic reports earn CPIC-guideline confidence, and where they quietly don't
Direct-to-consumer pharmacogenomic testing almost never sequences a gene from end to end. Companies like 23andMe and AncestryDNA run saliva DNA on a fixed genotyping microarray — a silicon chip pre-loaded with probes for a defined set of known single-nucleotide positions. This distinction between genotyping and sequencing is the single most important — and least understood — fact about how these reports are built.
A genotyping microarray works by hybridization: millions of short DNA probes, each complementary to a specific known variant position, are fixed to a chip surface. Saliva-derived DNA is fragmented, labeled, and washed across the chip; wherever a fragment binds its matching probe, a fluorescent signal reports which allele (A/C/G/T) is present at that exact, pre-chosen coordinate.
This means the chip can only "see" positions someone already decided to look for. A truly novel variant, a rare private mutation, or a large structural rearrangement that doesn't match any probe simply produces no signal — it is not flagged as absent, it is invisible. Clinical pharmacogenomic testing, by contrast, typically uses targeted PCR-based genotyping panels or long-read sequencing designed specifically to resolve the handful of pharmacogenes with known structural complexity.
Because only a few hundred thousand of the roughly 4–5 million common variants in a genome are directly genotyped, DTC pipelines routinely use statistical imputation — inferring likely genotypes at un-probed positions from population reference panels (e.g. 1000 Genomes, HRC) based on linkage disequilibrium patterns.
Imputation works well for common variants in well-represented populations, but accuracy drops sharply for rare alleles and for individuals of non-European ancestry, who are historically underrepresented in reference panels. A pharmacogene variant common in an East Asian or Sub-Saharan African population but rare in the training data can be imputed with far lower confidence — a documented equity gap in consumer genomics.
Every downstream step — diplotype calling, CPIC tiering, the final actionability score — inherits whatever uncertainty was baked in at the genotyping stage. A miscalled or un-imputed SNP doesn't announce itself; it silently propagates into a diplotype that looks exactly as authoritative as a correctly called one.
This is why FDA-authorized DTC pharmacogenetic reports (23andMe's 2018 authorization being the main US example) carry mandatory label language stating results are not diagnostic and must be confirmed with independent clinical-grade testing before any medication decision is made.
23andMe's FDA-authorized Pharmacogenetic Reports explicitly exclude CYP2D6 from their output — despite CYP2D6 governing metabolism of roughly 25% of all commonly prescribed drugs — specifically "because of the complexity of this gene." The chip that can tell you your ancestry composition to a fraction of a percent cannot reliably tell you your CYP2D6 status.
Once genotype calls exist, software translates them into "star alleles" — standardized names (CYP2C19*2, CYP2D6*4) catalogued by PharmVar for each pharmacogene — then pairs two star alleles into a diplotype, the genetic basis for predicting whether someone is a poor, intermediate, normal, rapid, or ultrarapid drug metabolizer. For most genes this works reasonably well. For CYP2D6 it is genuinely hard.
CYP2D6 sits on chromosome 22 immediately adjacent to CYP2D7, a highly similar non-functional pseudogene. Unequal crossover between the two during meiosis routinely produces structural variants: whole-gene deletions (*5, giving zero enzyme activity), whole-gene duplications and multiplications (*1xN, *2xN, giving amplified activity), and hybrid CYP2D6/CYP2D7 fusion alleles.
A SNP array measures signal intensity and allele ratio at fixed points — it was not designed to count gene copy number. A person carrying three copies of a functional CYP2D6 allele can look, on raw array output, indistinguishable from someone with two normal copies, because the array reports "what base is here," not "how many copies of this whole gene do you have."
Clinical pharmacogenomic labs resolve CYP2D6 structural variation using dedicated copy-number assays (TaqMan CNV assays, droplet digital PCR, or long-read sequencing that can span the full gene and its breakpoints), often combined with long-range PCR to separate CYP2D6 from CYP2D7 sequence before genotyping. This is fundamentally different laboratory work from a consumer genotyping chip, and it is why clinical-grade PGx panels cost several hundred dollars while a DTC PGx add-on can cost nothing.
Misclassifying a duplication carrier as "normal metabolizer" instead of "ultrarapid metabolizer" is not a rounding error: for a prodrug like codeine, an ultrarapid metabolizer converts codeine to morphine far faster than expected, risking toxicity at a "standard" dose.
A diplotype also requires phasing: knowing whether two heterozygous variants sit on the same physical chromosome copy (cis) or opposite copies (trans), since this changes which star allele combination actually exists. Consumer pipelines infer phase statistically from population haplotype frequencies rather than measuring it directly (e.g. via long-read sequencing or family trio data), which is usually accurate for common haplotypes but can fail for rare or novel combinations — another silent source of diplotype error that never appears in the final printed result.
The Clinical Pharmacogenetics Implementation Consortium (CPIC), founded in 2009 and affiliated with PharmGKB and the NIH Pharmacogenomics Research Network, publishes peer-reviewed guidelines that grade drug-gene evidence by strength — from "genetic information should change prescribing" down to "insufficient evidence to act." A DTC report can technically report a variant honestly while still obscuring exactly where it falls on that ladder.
CPIC assigns each guideline a level reflecting the strength and consistency of evidence: Level A pairs have data strong enough that genetic results should change prescribing (e.g. CYP2C19–clopidogrel, CYP2D6–codeine, SLCO1B1–simvastatin, TPMT/NUDT15–thiopurines, DPYD–fluoropyrimidines). Level B pairs have moderate evidence where genetic information could reasonably inform prescribing. Level C pairs have weaker or preliminary evidence not yet sufficient to recommend a prescribing action. PharmGKB layers a parallel four-tier clinical annotation scale (1A highest, 4 lowest) built from curated literature review.
This tiering exists precisely because pharmacogenomic research spans an enormous range of rigor — from large multi-cohort replicated studies to single small candidate-gene papers that never replicated.
A consumer report page showing "CYP2C19: Rapid Metabolizer" next to "Caffeine Metabolism: Fast" next to "Muscle Composition: Likely more power, less endurance" presents all three with the same typography, the same confident tone, and often the same colored badge — even though only the first rests on CPIC Level A clinical guideline evidence, while the caffeine and muscle-composition associations derive from population-level GWAS trait studies never intended to guide any decision.
The FDA's 2018 authorization order for 23andMe explicitly required disclaimers distinguishing "reports on genetic variants that may be associated with medication metabolism" from any implication about clinical dosing, precisely to counter this flattening effect.
The FDA independently maintains a public "Table of Pharmacogenomic Associations" listing gene-drug pairs with evidence sufficient to support dosing or safety label language — a narrower, drug-label-anchored list than CPIC's broader guideline set. Where CPIC asks "is there enough evidence to guide prescribing," the FDA table asks "is there enough evidence to change this drug's approved label." A pair can appear in one list and not the other; a DTC report rarely explains which standard it is using for any given claim.
| Product | Indication | Trial Design | Key Result |
|---|---|---|---|
| CPIC Level A | CYP2C19 – clopidogrel | Multiple replicated cohorts; FDA boxed warning; guideline dosing algorithm exists | Colored badge, same visual weight as lower tiers |
| CPIC Level B | CYP2C9/VKORC1 – warfarin | Moderate, fairly consistent evidence; informs but doesn't mandate prescribing | Often unlabeled as "moderate" on the page |
| PharmGKB 2B–3 | COMT – opioid response | Small or mixed-replication candidate-gene studies | Presented as a discrete trait, tier rarely disclosed |
| No CPIC guideline | "Caffeine metabolism," "weight response" | GWAS trait associations, not designed for clinical action | Marketed as a headline wellness insight |
To make a report legible, DTC platforms compress dozens of variant calls into a small set of headline findings and, increasingly, a composite "actionability" or "impact" score. Compression is not inherently deceptive — but collapsing a CPIC Level A finding and a speculative wellness trait into visually equivalent tiles erases the exact distinction a prescriber would need before acting on either one.
In November 2013 the FDA sent 23andMe a warning letter ordering it to immediately stop marketing its Saliva Collection Kit and Personal Genome Service for health purposes, citing lack of analytical and clinical validation for the company's disease-risk and drug-response claims — the device had never received FDA marketing authorization or clearance. 23andMe suspended health reporting entirely for roughly two years while it pursued validation studies.
The company returned to health reporting gradually, culminating in the October 2018 De Novo authorization for its Pharmacogenetic Reports — the first FDA marketing authorization for a DTC test reporting on how a person's genetics may affect their metabolism of certain medications. That authorization came with binding label conditions: reports must state they are not intended to diagnose a disease, must not be used alone to determine any medication or dose, and must direct users to consult a healthcare professional before making any medication changes.
A well-built score can legitimately encode the CPIC tier of the underlying evidence, the confidence of the genotype call, and whether the associated drug is commonly prescribed — all of which are computable from the pipeline stages before it. What it structurally cannot encode without external clinical context is the patient's actual medication list, other diagnoses, drug-drug interactions, renal or hepatic function, or pregnancy status — all of which a prescribing clinician would weigh before changing therapy.
This is the core actionability gap: even a perfectly validated CYP2C19 poor-metabolizer call, correctly reported at CPIC Level A, is not itself a prescribing decision. It is one input a clinician combines with everything else they know about the patient.
Many "wellness" trait reports bundled alongside pharmacogenomic findings — caffeine sensitivity, sleep quality, dietary response — are marketed using language reminiscent of the Dietary Supplement Health and Education Act (DSHEA, 1994) framework: structure/function claims that avoid the FDA's drug-efficacy evidentiary bar because they don't claim to treat or diagnose disease. That is a legitimate regulatory category, but placed next to an FDA-authorized pharmacogenetic finding on the same report page, the visual and rhetorical distinction between the two evidentiary standards can disappear entirely for the reader.
The FDA's 2018 authorization order is unusually explicit: it requires that 23andMe's Pharmacogenetic Reports state outright that the test "does not describe if a person will or will not respond to a specific medication" — a direct regulatory acknowledgment that a genotype, even correctly called and correctly tiered, is not a response prediction.
Every layer of this pipeline — array genotyping, diplotype calling, CPIC tiering, score compression — converges on one final, decisive gap: a DTC report reaches the consumer directly, with no prescribing clinician required to interpret it before the consumer acts. CPIC, PharmGKB, and FDA guidance are unanimous that a clinical decision should never rest on an uninterpreted consumer report alone.
CYP2D6 ultrarapid metabolizers convert codeine to morphine unusually fast; a case that shaped US drug labeling involved an infant death from morphine toxicity via breast milk after the nursing mother, an unrecognized CYP2D6 ultrarapid metabolizer, took standard-dose codeine. The FDA subsequently issued a boxed warning (2013) contraindicating codeine in breastfeeding mothers and in children after tonsillectomy/adenoidectomy.
For clopidogrel (Plavix), CYP2C19 poor metabolizers convert less of the prodrug to its active antiplatelet form; the FDA added a boxed warning in 2010 noting reduced effectiveness in poor metabolizers, particularly relevant after coronary stenting, where inadequate platelet inhibition raises stent-thrombosis risk. For simvastatin, the SEARCH Collaborative Group (2008, NEJM) found SLCO1B1 reduced-function carriers face roughly 4.5x higher myopathy risk per allele copy, up to ~17x for two copies at high dose — a CPIC Level A pair informing dose selection, not drug avoidance.
CLIA (Clinical Laboratory Improvement Amendments, 1988) certification governs any laboratory performing testing used to diagnose, prevent, or treat disease; results intended purely for consumer education can be produced in a CLIA-certified lab yet still be labeled "not for diagnostic use," a distinction that confuses many consumers who reasonably assume a certified lab implies clinical-grade reliability for every purpose.
CPIC guidelines and FDA authorization conditions both specify that a DTC pharmacogenetic finding — even a correctly tiered Level A result — should be confirmed with independent, clinically validated pharmacogenetic testing before it informs an actual prescribing change, precisely because of the array-genotyping and structural-variant limitations covered in Stage 1 and Stage 2.
The regulatory environment around the underlying genetic data adds a second layer of risk. The Genetic Information Nondiscrimination Act (GINA, 2008) bars health insurers and employers from using genetic information adversely, but explicitly does not cover life, disability, or long-term-care insurance — a gap consumers are frequently unaware of when they test. HIPAA, meanwhile, generally does not apply to DTC genetic-testing companies at all, since they are not "covered entities" under the statute; the FTC Act's prohibition on unfair or deceptive practices has become the default enforcement tool instead, as in the FTC's 2023 settlement with genetic-testing company 1Health.io (Vitagene) over deceptive data-security and retention claims.
The October 2023 23andMe breach — attackers used credential stuffing to access roughly 6.9 million users' profile and ancestry data via the "DNA Relatives" feature — and the company's subsequent March 2025 Chapter 11 bankruptcy filing, which put millions of consumers' genetic records up for sale as a company asset, together illustrate that the actionability question extends beyond any single report: it is also a question of who controls the underlying genome over time.
A patient who unilaterally stops a CYP2C19-metabolized antiplatelet drug, or adjusts a statin dose, based solely on a DTC report — without the confirmatory clinical testing and clinician interpretation every CPIC guideline and the FDA authorization itself require — has converted a genuinely useful piece of evidence into an uncontrolled real-world experiment on their own prescription.