HomeDirect-to-Consumer Genomics TestingRaw Genetic Data Third-Party Reanalysis Tool

🧬 Raw Genetic Data Third-Party Reanalysis Tool

This tool enables users to submit raw genetic data for reanalysis by a third-party service, providing additional insights and interpretations of the data.

Direct-to-Consumer Genomics Testing2DModerate60 FPS
raw-genetic-data-reanalysis ↗ Open standalone

What "Raw Data" Actually Is — Microarray Genotyping, Not Genome Sequencing

Every direct-to-consumer (DTC) ancestry or health test — 23andMe, AncestryDNA, MyHeritage DNA, FamilyTreeDNA — is built on a SNP genotyping microarray, not whole-genome sequencing. Consumers who download their "raw data" receive a plain-text file of a few hundred thousand pre-selected genotype calls, sampling well under 0.1% of the ~3.2 billion base-pair genome. Understanding this distinction is the single most important fact for interpreting any third-party reanalysis correctly.

  • 630–700k: SNPs genotyped (typical) (Illumina GSA-class chip)
  • <0.02%: Genome fraction sampled (of ~3.2 Gb genome)
  • 0×: Whole-genome sequencing depth (arrays do not sequence)
  • >14M: 23andMe users (cumulative) (genotyped since 2007)

Microarray genotyping vs. clinical sequencing

DTC platforms use fixed-content SNP microarrays — most recently variants of the Illumina Global Screening Array (GSA) — where hundreds of thousands of short DNA probes are chemically bound to a chip. Each probe is designed to hybridize to one specific, pre-selected genomic position and report which of two alleles is present (AA, AG, or GG, for example). This is fundamentally a "multiple choice" measurement: the chip can only report on positions it was manufactured to look for.

Clinical whole-exome sequencing (WES) reads ~85 Mb across ~20,000 genes base-by-base; whole-genome sequencing (WGS) reads essentially the entire 3.2 Gb genome. A consumer array, by contrast, genotypes roughly 630,000–700,000 positions selected for ancestry informativeness, common trait associations, and a curated set of pharmacogenomic and disease-risk markers — it cannot detect novel mutations, most indels, structural variants, or the vast majority of the ~4–5 million single-nucleotide variants an individual actually carries relative to the reference genome.

The raw data file itself (a .txt or .csv, roughly 15–25 MB) lists four columns per line: rsID, chromosome, position, and genotype — with no interpretation, no annotation, and no clinical review attached.

Why the file looks authoritative but is not diagnostic

Because the export comes directly from the testing company's own account portal, users often assume it carries the same rigor as their curated health reports. It does not. 23andMe's FDA-authorized (2017/2018) Health Predisposition Reports interpret only a narrow, validated subset of positions on the chip — for example, just 3 specific Ashkenazi-founder BRCA1/BRCA2 variants out of more than 1,000 known pathogenic variants across those two genes. The remaining ~630,000+ raw calls on the same file were never validated for individual clinical reporting; they were selected and genotyped for population-scale research, ancestry inference, and trait-association purposes.

When a third-party tool reprocesses the entire raw file, it is running an unfiltered pass over positions the original company deliberately chose not to report on clinically — which is precisely what creates the discordance explored in later stages.

Third-Party Interpretation Services — Promethease, SNPedia, and the Upload Ecosystem

Once exported, the raw file can be uploaded to services entirely unaffiliated with the original testing company. The best known is Promethease, built on SNPedia — a crowdsourced genetics wiki founded in 2006 by Mike Cariaso and Greg Lennon. These tools re-parse the same underlying genotype calls, cross-referencing every rsID against a literature-derived knowledgebase to generate a new, often far larger, report — with no new laboratory work performed.

  • 2006: SNPedia founded (Cariaso & Lennon)
  • >100,000: SNPedia catalogued SNPs (annotated rsIDs (peak))
  • thousands: Promethease report size (of matched genotype "snippets")
  • 2019: Promethease → MyHeritage (acquisition of SNPedia/Promethease)

How an upload-based reanalysis tool works

The user's raw genotype file is uploaded (historically for a small one-time fee, around $12 for Promethease) to a server that runs a purely computational matching process: for every line of the file, the tool looks up the rsID in its reference database and returns any known annotation for that specific genotype — for example, "rs429358(C;C) — increased Alzheimer's risk via APOE ε4/ε4."

No blood draw, no new chip, no sequencing occurs. The signal quality of the output is entirely bounded by (a) the accuracy of the original array call and (b) the quality and currency of the literature the third-party database has indexed. This is fundamentally different from a clinical genetic test, where a laboratory director reviews evidence tiers (per ACMG/AMP guidelines) before a variant is reported as pathogenic, benign, or of uncertain significance.

The 2023–2024 security disruption — a real-world case study

The upload ecosystem's dependence on centralizing sensitive raw genetic files became a live liability after the October 2023 23andMe data breach, in which attackers used credential-stuffing (reused passwords from other breaches) to access accounts, then exploited the "DNA Relatives" opt-in feature to scrape profile data from roughly 6.9 million linked accounts — including ancestry and, for a subset, health-related information.

In direct response, MyHeritage — which had acquired Promethease/SNPedia in 2019 — suspended new Promethease raw-data uploads in December 2023, citing data-security risk from processing highly identifying, non-reversible genetic files on top of an already-battered trust environment in the DTC sector. The episode illustrates a structural risk unique to reanalysis tools: every additional upload destination is another location where an irrevocable, immutable dataset (your genome cannot be "reset" like a password) is stored.

Genetic data cannot be revoked or reissued the way a password or credit card can. A raw genotype file leaked once from any of the (potentially several) services a consumer has uploaded it to remains exploitable for that person's entire life — and, because relatives share DNA, for their family members' lives as well.

The regulatory gap: DTC raw-data uploads generally fall outside HIPAA

A persistent misconception is that HIPAA protects genetic data uploaded to these tools. HIPAA's Privacy Rule applies only to "covered entities" — health plans, healthcare clearinghouses, and providers who transmit health information electronically in connection with certain transactions — plus their business associates. Direct-to-consumer testing companies and independent reanalysis tools are typically neither, because the consumer purchases the test directly rather than through a covered healthcare provider.

Instead, this class of data is governed by a patchwork: the FTC's general unfair-and-deceptive-practices authority (Section 5 of the FTC Act), state genetic privacy laws (for example, Illinois' Genetic Information Privacy Act and California's CGDPA/CCPA "genetic information" category), and, for EU residents, the GDPR, which explicitly classifies genetic data as a "special category" requiring heightened protection and explicit consent. The FTC's 2023 enforcement action against Vitagene/1Health.io — settled for mishandling consumer genetic and health data, including allegedly leaving data exposed and failing honor deletion requests — is a concrete example of enforcement filling the gap where HIPAA does not reach.

Matching Genotypes to SNPedia and ClinVar — Two Very Different Kinds of Evidence

The interpretive core of any reanalysis tool is a lookup step: each of the consumer's ~630,000+ genotype calls is checked against one or more reference knowledgebases. SNPedia and ClinVar are the two most commonly used — and they differ enormously in curation rigor, which matters more than most users realize once a "hit" appears in a report.

  • >3M: ClinVar submitted records (variant-condition assertions)
  • ~9–17%: ClinVar conflicting interpretations (of multi-submitter variants)
  • wiki: SNPedia curation model (community-edited, uneven review)
  • thousands: Typical hits per raw file (many of low/no clinical weight)

SNPedia "magnitude" is not a clinical significance grade

SNPedia assigns each annotated genotype a "magnitude" score, loosely from 0 to 10, meant to convey how interesting or significant a finding is. This score is editorially assigned by wiki contributors and is explicitly not equivalent to the ACMG/AMP five-tier system used in clinical genetics (Pathogenic, Likely Pathogenic, Uncertain Significance, Likely Benign, Benign). A magnitude of 8 might reflect a well-replicated pharmacogenomic association (e.g., CYP2C19 poor-metabolizer status affecting clopidogrel efficacy) or, in other entries, a single small study that was never replicated.

Because the underlying source articles range from large GWAS meta-analyses to individual case reports, two variants with identical-looking "high magnitude" flags in a Promethease-style report can carry wildly different actual evidentiary weight — a distinction the report format does not make visually obvious to a lay reader.

ClinVar is stronger evidence, but still not free of conflict

ClinVar, maintained by the NCBI, aggregates variant-to-condition interpretations submitted directly by clinical testing laboratories, research groups, and expert panels, along with the assertion criteria and evidence each submitter used. It is far closer to clinical-grade evidence than a wiki, and “reviewed by expert panel” or multi-lab “no conflicts” consensus records carry real diagnostic weight.

However, ClinVar is a submission archive, not an adjudicated verdict: independent studies auditing ClinVar have found conflicting classifications (e.g., one lab calling a variant pathogenic while another calls the same variant a variant of uncertain significance, VUS) for roughly one in ten to one in six variants with multiple submissions. A raw-data reanalysis tool that simply surfaces "this variant has a ClinVar entry" without displaying which submitters agree, their review status, and their evidence level can therefore present contested science as settled fact.

The matching step inherits every upstream measurement error

Cross-referencing is purely a lookup operation — it cannot correct an incorrect genotype call. If the original array mis-called a base (see Stage 4), the wrong call will be matched confidently against the correct database entry for the wrong genotype, producing an internally consistent but factually wrong finding. This is why reanalysis tools that advertise "instant" health insights from an existing raw file are, at best, exactly as reliable as the underlying array call — and, at worst, propagate array-specific errors into clinical-sounding language the array manufacturer never intended for individual diagnostic use.

Where Third-Party Reports Disagree With the Original Company's Report

When a third-party reanalysis surfaces findings the original DTC company never reported — or contradicts it outright — the cause is rarely a database error. It is usually the genotyping array itself: research-grade microarrays carry measurable error and no-call rates that clinical sequencing pipelines are specifically engineered to minimize.

  • 0.1–0.5%: Array genotype error rate (per-SNP miscall, typical GSA chip)
  • 1–2%: No-call rate (dropped SNPs) (low signal / probe failure)
  • ~40%: False "pathogenic" flags in DTC raw data (Tandy-Connor et al. 2018, Genet Med)
  • >1,000: BRCA1/2 pathogenic variants known (vs. 3 tested by 23andMe FDA report)

Why a 0.1–0.5% error rate is not reassuring at genome scale

A per-SNP error rate that sounds small becomes large in aggregate. Applied across ~630,000 genotyped positions, even a conservative 0.1% miscall rate implies roughly 600+ individual genotype calls on a typical raw-data file are simply wrong — flipped alleles, miscalled heterozygotes, or no-calls silently filled with a default. Error rates rise further in specific failure modes: probes that overlap segmental duplications or processed pseudogenes (for example, the PMS2 gene relevant to Lynch syndrome has multiple closely related pseudogenes) are prone to cross-hybridization, where the probe binds a near-identical but functionally different sequence and reports a genotype for the wrong locus entirely.

Arrays are also poorly suited to certain variant classes outright: large insertions/deletions, copy-number variants, and repeat expansions are frequently missed or miscalled by design, because the hybridization chemistry assumes a short, single-base difference at a known position.

The landmark discordance study

A widely cited 2018 study by Tandy-Connor and colleagues (Genetics in Medicine, the official journal of the American College of Medical Genetics and Genomics) sent DTC raw-data files that had generated "pathogenic" or "increased risk" flags through third-party interpretation tools to a CLIA-certified clinical laboratory (Ambry Genetics) for confirmatory testing. Of the variants flagged as clinically significant by the raw-data reanalysis process, approximately 40% did not validate on clinical-grade confirmatory testing — they were false positives, most attributable to genotyping platform error rather than a true underlying mutation.

The same study documented cases where a consumer received a third-party report suggesting a serious pathogenic finding (e.g., in a cancer-predisposition gene) that a certified lab subsequently determined was simply an incorrect array call — with the accompanying real-world consequence of significant anxiety and, in some reported cases, unnecessary follow-up medical procedures before the error was caught.

In the Tandy-Connor study, roughly 40% of "pathogenic" findings generated by third-party raw-data interpretation failed to replicate on CLIA-certified clinical confirmatory testing — meaning two out of every five alarming results examined were array noise, not real mutations.

Discordance is compounded by scope mismatch, not just error

A second, distinct source of discordance is scope: the original company's clinical health report and the third-party reanalysis are answering different questions from the same file. 23andMe's FDA-authorized BRCA1/BRCA2 report tests exactly 3 founder variants common in people of Ashkenazi Jewish descent; it explicitly does not screen the full coding sequence of either gene. A third-party tool processing the same raw file may surface a "hit" at a different BRCA1/2 position elsewhere on the array — a position the array was never validated to genotype reliably for individual clinical use, and one the original company deliberately excluded from its cleared report for exactly that reason.

A lay user comparing the two reports experiences this as "the two reports disagree" — when the more precise description is that one report answered a narrow, validated question and the other attempted to answer a much broader, unvalidated one using the same raw measurement.

Array genotyping vs. clinical-grade sequencing

ProductIndicationTrial DesignKey Result
DTC SNP microarray (raw data source)~630k–700k fixed positionsProbe hybridization; reports 2-allele genotype at pre-selected sites onlyCheap ($99–$200), fast, ideal for ancestry/traits at scale
Third-party reanalysis (Promethease-style)Any rsID present in raw fileDatabase lookup only; no new measurement; inherits array errorBroad, low-cost secondary interpretation of existing data
Clinical NGS panel / exome sequencingSpecific gene(s) or exome, base-by-baseSequencing with orthogonal confirmation (Sanger) of flagged variantsCLIA/CAP-validated, ACMG-tiered, actionable for medical decisions
"Discordant" raw-data flagN/A — an artifact, not a locusArray miscall, no-call default, or out-of-scope position matched to DBNone — signals need for confirmatory testing before acting

Why Nothing From a Reanalysis Report Should Drive a Medical Decision Unconfirmed

The appropriate endpoint of a third-party raw-data reanalysis is not a diagnosis — it is a hypothesis to be tested. Clinical genetics has a specific, well-established mechanism for turning an uncertain array-based flag into an actionable result: confirmatory testing in a CLIA-certified laboratory, ideally paired with pre- and post-test genetic counseling.

  • Yes: CLIA labs required for diagnosis (raw-data tools are not CLIA-certified)
  • confirm before acting: ACMG guidance on DTC raw data (2019 points-to-consider statement)
  • employment + health insurance only: GINA protection scope (2008 federal law)
  • life, disability, LTC insurance: GINA gaps (not covered by GINA)

The confirmatory testing pathway

Professional genetics bodies, including the American College of Medical Genetics and Genomics (ACMG), have issued formal guidance stating that findings from DTC raw-data third-party interpretation should not be used for medical decision-making without confirmation through a CLIA-certified clinical laboratory using a validated methodology (typically targeted Sanger sequencing of the specific position, or a clinical NGS panel).

This confirmation step exists precisely because it corrects the two failure modes explored in Stage 4: it re-measures the actual DNA sequence at the position in question (catching array miscalls and no-call artifacts) and it applies ACMG/AMP evidence-tiering rather than a wiki magnitude score (catching scope and evidence-quality mismatches). A board-certified genetic counselor or clinical geneticist can also place any confirmed finding in the context of family history, penetrance data, and available risk-reduction options — context a static report, however detailed, cannot provide.

Legal protections are narrower than most consumers assume

The Genetic Information Nondiscrimination Act (GINA, 2008) is the primary U.S. federal statute addressing genetic discrimination. GINA prohibits health insurers from using genetic information (including DTC and reanalysis results) to set premiums, deny coverage, or require genetic testing, and prohibits most employers from using genetic information in hiring, firing, or promotion decisions.

Critically, GINA does not extend to life insurance, disability insurance, or long-term care insurance — insurers in those markets can, in most states, request and use genetic test results, including raw-data reanalysis findings, in underwriting decisions. This gap is precisely why unconfirmed, alarming "hits" generated by an unvalidated third-party reanalysis carry downstream financial risk beyond the immediate emotional impact of an unconfirmed result: a false-positive pathogenic flag disclosed on a life-insurance application, even if later shown to be array noise, has already entered that decision-making record.

Practical guardrails for anyone using a reanalysis tool

Genetics professionals converge on a consistent set of practices for engaging with third-party raw-data reanalysis responsibly:

• Treat every "high magnitude" or "pathogenic" flag as a lead, not a result — confirmatory testing is required before any clinical action • Prefer ClinVar entries with "reviewed by expert panel" or multi-lab, no-conflict status over single-submitter or SNPedia-only annotations • Seek pre-test and post-test genetic counseling for any flag touching a medically actionable gene (the ACMG maintains a list of ~80 genes recommended for secondary-finding reporting in clinical sequencing) • Understand that deleting an account does not necessarily delete an already-downloaded raw file sitting on a third-party server or a personal hard drive — data minimization has to happen upstream, at the decision to upload at all • Recognize that relatives share a meaningful fraction of your genome — a decision to upload raw data has privacy implications for family members who never consented

The gap between "a data file says pathogenic" and "a certified laboratory confirms pathogenic" is not a technicality — in the Tandy-Connor dataset it was the difference between a real finding and a false alarm roughly 40% of the time. Confirmatory testing is not optional caution; it is the step that makes a reanalysis result meaningful at all.
⚙ Under the hood

This tool enables users to submit raw genetic data for reanalysis by a third-party service, providing additional insights and interpretations of the data.

CanvasBiomedicine

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)