Why a consumer DNA test can call your eye color with confidence but only whispers a probability about your height
A genome-wide association study (GWAS) tests millions of single-nucleotide polymorphisms (SNPs) one at a time for statistical association with a trait or disease across a huge cohort. Consumer DNA testing companies build their polygenic predictions on top of publicly deposited GWAS summary statistics — the raw material for every "trait report" a customer receives.
A GWAS regresses a phenotype (height in cm, presence/absence of baldness, diagnosed type 2 diabetes) against the genotype at each SNP — typically encoded as 0, 1, or 2 copies of a reference allele — while statistically controlling for age, sex, genotyping batch, and principal components of ancestry.
Because hundreds of thousands to millions of SNPs are tested simultaneously, the significance threshold must be extremely strict to avoid false positives: p<5×10⁻⁸ is the field-standard "genome-wide significance" line, derived from a Bonferroni correction for roughly one million independent tests (accounting for linkage disequilibrium between neighboring SNPs).
The output is a Manhattan plot — chromosome position on the x-axis, −log₁₀(p-value) on the y-axis — where each significant "peak" marks a genomic region containing a causal or (more often) a tagging variant in linkage disequilibrium with the true causal variant.
The 2022 Yengo et al. height GWAS (Nature) combined data from 5.4 million individuals across the GIANT consortium and 23andMe — at the time the largest GWAS ever conducted for any human trait — and still explained less than half of the heritability twin studies had already established a century earlier.
Statistical power to detect a SNP's effect scales with cohort size and with the square of the effect size. Because most complex-trait SNPs have tiny individual effects (often shifting height by a fraction of a millimeter), discovering them at all requires enormous cohorts.
Early height GWAS in the mid-2000s (n≈5,000–30,000) found only a handful of loci. By 2014, the GIANT consortium (n≈250,000) had found 697 loci. By 2022 (n≈5.4M), that number exceeded 12,000 independent loci — and the curve has not yet saturated: each doubling of sample size continues to reveal new, smaller-effect variants.
Consumer testing companies like 23andMe leverage their own genotyped customer base (>14 million as of 2024, most of whom consent to research) as a GWAS discovery cohort in its own right, often in meta-analysis with UK Biobank and academic consortia.
If a GWAS cohort mixes ancestral subpopulations with different allele frequencies AND different average trait values for reasons unrelated to those alleles (e.g. cultural, dietary, or socioeconomic differences that correlate with ancestry), spurious associations appear. This is population stratification.
Modern GWAS correct for it using principal component analysis (PCA) on genome-wide genotypes and/or linear mixed models (e.g. BOLT-LMM, REGENIE) that model relatedness directly. Even so, residual stratification can inflate polygenic score performance when discovery and target cohorts share subtle ancestry or even social/environmental structure correlated with genotype — a major reason PGS trained in one population transfer poorly to others.
A polygenic score (PGS) is nothing more than a weighted sum: for every included SNP, multiply the number of trait-increasing alleles a person carries (0, 1, or 2) by that SNP's estimated effect size, then add them all up. The statistical art lies in choosing which SNPs to include and how to weight them without double-counting correlated signals.
If a PGS only used the ~12,000 genome-wide significant SNPs for height, it would leave substantial predictive information on the table. Many true causal variants have effects too small to cross p<5×10⁻⁸ even in million-person cohorts, but their aggregate contribution is real and measurable.
Modern PGS methods (LDpred2, PRS-CS, SBayesR) use a Bayesian framework: they shrink noisy, small-effect SNP estimates toward zero while preserving genuine signal, using the whole genome-wide distribution of association statistics rather than a hard significance cutoff. This is why a state-of-the-art height PGS can include over a million SNPs — nearly the entire common-variant genome — each contributing a minuscule, but non-zero, weighted nudge.
Neighboring SNPs on a chromosome are often inherited together as a block (linkage disequilibrium), so a naive sum of all "significant-looking" SNPs would massively over-count a single true signal that happens to be tagged by 50 correlated markers.
LD-aware methods use reference panels (1000 Genomes, or UK Biobank's own LD structure) to estimate the correlation between nearby SNPs and re-weight or "clump" them so that each independent genomic signal contributes proportionally once, not fifty times. Getting this step wrong is one of the most common causes of inflated or miscalibrated PGS performance in poorly validated consumer products.
A polygenic score is a probability nudge, not a genetic diagnosis: for height, moving from the 10th to the 90th percentile of the PGS distribution shifts expected adult height by roughly 6–7 cm on average — but individual outcomes still vary by several centimeters around that expectation due to the unexplained 60% of heritability plus environment.
Because >80% of GWAS discovery participants to date have been of European ancestry, PGS accuracy drops substantially when applied to individuals of African, East Asian, South Asian, or admixed American ancestry — driven by differing allele frequencies, differing LD patterns, and differing causal-variant effect sizes across populations.
Martin et al. 2019 (Nature Genetics) showed height and other trait PGS can lose 50–80% of their predictive accuracy when transferred from European-ancestry discovery cohorts to African-ancestry target samples. Consumer DNA companies now disclose ancestry-specific confidence intervals for exactly this reason, and multi-ancestry GWAS consortia (e.g. the All of Us Research Program, ~45% non-European participants) are actively working to close this gap.
Not all traits are equally predictable from DNA. The deciding factor is genetic architecture: how many loci contribute, how large their individual effects are, and how much of the trait's total variance is heritable at all versus shaped by environment, chance, or diagnostic criteria. Comparing four traits side-by-side shows the full spectrum consumer genomics has to work with.
Eye color is often cited as a genetics success story, and largely deserves it — but only for the blue-versus-brown axis. A single intronic SNP, rs12913832, located near the HERC2 gene (which regulates expression of the neighboring OCA2 pigmentation gene), explains roughly 74% of the variance in blue/brown eye color by itself (Sturm et al. 2008, American Journal of Human Genetics).
The HIrisPlex-S forensic/consumer prediction system adds a handful of additional SNPs (from genes including OCA2, SLC24A4, TYR, and IRF4) to reach AUC 0.91–0.94 for correctly classifying blue versus brown eyes. But prediction accuracy collapses for intermediate categories — green and hazel eyes involve more complex, less-understood combinatorial genetics, and reported accuracy for these categories often falls to 30–50%, far closer to chance.
Height is the canonical model of a complex, highly polygenic trait: thousands of independent loci across every chromosome each contribute a tiny effect, and no single variant explains more than a fraction of a percent of variance. The 2022 Yengo et al. GIANT/23andMe meta-analysis (n=5.4 million) identified 12,111 approximately independent genome-wide-significant SNPs, together explaining ~40–45% of height variance in individuals of European ancestry — the current practical ceiling for common-variant PGS.
Compare that to twin and family studies, which have consistently estimated height heritability at 80% since Francis Galton's original 19th-century studies of parent-offspring height correlation. The 40-percentage-point gap between what twins tell us is heritable and what current PGS can capture is the "missing heritability" problem explored in the next stage.
Male-pattern baldness sits between the two extremes. It is strongly heritable (twin studies estimate ~80% heritability) and strongly polygenic (624 loci identified by Yap et al. 2018 in ~205,327 UK Biobank men), including a large-effect locus on the X chromosome near the androgen receptor gene. A PGS built from these loci reaches AUC ≈0.78–0.80 for predicting severe (Norwood-Hamilton stage 4+) baldness by midlife — respectable, but well short of eye color's near-deterministic accuracy.
Type 2 diabetes sits at the far end: it is a disease outcome shaped heavily by diet, body weight, physical activity, and age, on top of a substantial but diffuse genetic component. The largest published PGS for type 2 diabetes (Mahajan et al. 2018, DIAGRAM consortium, >898,000 individuals, 338 genome-wide significant loci) achieves AUC of only ≈0.66–0.70 — modestly better than chance (0.50), and clinically useful mainly as one input alongside BMI, age, and family history rather than as a standalone predictor.
| Product | Indication | Trial Design | Key Result |
|---|---|---|---|
| Eye color (blue/brown) | ~6 SNPs (HERC2/OCA2 dominant) | Near-oligogenic; one variant explains ~74% of variance | AUC 0.91–0.94 |
| Height | 12,111+ SNPs, up to >1M in full PGS | Maximally polygenic; every locus tiny effect | R² ≈ 40% (ceiling 80%) |
| Male-pattern baldness | 624 loci incl. X-chromosome AR locus | Highly polygenic + strong single locus | AUC ≈ 0.78–0.80 |
| Type 2 diabetes | 338+ loci, DIAGRAM consortium | Polygenic + dominant environmental drivers | AUC ≈ 0.66–0.70 |
Twin and family studies have measured the heritability of height, baldness, and countless other traits for over a century, using a method entirely independent of modern genotyping: comparing trait similarity between identical (monozygotic) twins who share 100% of their DNA against fraternal (dizygotic) twins who share ~50%. The gap between that classical estimate and what current PGS technology can explain is one of genetics' most persistent open problems.
It helps to separate three distinct heritability estimates that are often conflated in press coverage:
1. Twin/family (broad-sense) heritability — derived from resemblance between relatives with known degrees of genetic relatedness. For height, this has reliably come out around 80% across dozens of studies spanning a century, in many different populations and environments.
2. SNP-heritability — estimated using methods like GCTA-GREML (Yang et al. 2010) that measure how much phenotypic variance is explained by the additive effects of all common SNPs simultaneously, using genome-wide genetic relatedness between unrelated individuals rather than known family structure. For height this lands around 50–60% — already lower than twin heritability, because it only captures common variants tagged on genotyping arrays.
3. PGS-explained variance — the practical, "out of sample" predictive R² achieved by an actual polygenic score built from GWAS summary statistics. For height this is currently ~40%, lower still than SNP-heritability because of statistical noise, imperfect effect-size estimation, and finite discovery sample size.
Manolio et al.'s landmark 2009 Nature paper "Finding the Missing Heritability of Complex Diseases" formally named this gap and catalyzed a decade of methodological innovation — larger cohorts, better statistical models, and whole-genome sequencing — that has narrowed but not closed it for almost every complex trait studied.
Leading explanations, likely all contributing simultaneously:
• Rare variants: variants with minor allele frequency below ~1% are poorly captured by standard genotyping arrays and imputation, yet can carry larger individual effects than common variants — whole-genome sequencing studies are steadily recovering some of this signal.
• Structural variation: copy-number variants, insertions, deletions, and repeat-length polymorphisms are invisible to SNP arrays entirely but can meaningfully affect gene expression and trait variance.
• Non-additive effects: dominance and epistasis (gene-gene interaction) are mathematically excluded from standard additive PGS models, even though they contribute to broad-sense twin heritability.
• Imperfect LD tagging: even common causal variants may not be perfectly correlated with any genotyped SNP, diluting the observable association signal.
• Gene-environment correlation inflating twin estimates: some classical twin-study heritability may itself be modestly inflated by identical twins sharing more similar environments than fraternal twins (the "equal environments assumption"), though multiple validation studies suggest this effect is real but small for height specifically.
Unlike a hard physical limit, the PGS/twin-heritability gap has narrowed steadily and predictably as GWAS sample sizes have grown: 5% variance explained (2008, small early GWAS) → 16% (2014, GIANT n≈250k) → ~25% (2018, n≈700k) → ~40% (2022, n≈5.4M). Extrapolating this trend, some researchers project common-variant PGS could approach the ~50–60% SNP-heritability ceiling within another decade of cohort growth, though closing the remaining gap to the full ~80% twin estimate will likely require sequencing-based approaches that capture rare and structural variation directly.
Even a perfectly calibrated polygenic score, built from a maximally powered GWAS, would still leave a substantial share of any trait unexplained — because a meaningful fraction of human variation is not genetic at all. Nutrition, illness, chance developmental events, socioeconomic conditions, and simple biological noise all shape the final, observed trait. This is the deepest reason consumer reports must speak in percentiles and probabilities.
If height were purely genetic, average population height would be stable across generations absent large-scale migration or selection. It is not. Average male height in the Netherlands rose by roughly 20 cm between the mid-19th century and the early 21st century — a change far too fast for allele frequencies to have shifted meaningfully, and instead driven overwhelmingly by improved childhood nutrition, reduced infectious disease burden, and better prenatal care.
This "secular trend," documented across many industrializing nations, is direct empirical proof that genetics sets a range of potential outcomes, not a fixed destiny — the same DNA, raised in different nutritional environments a century apart, produces substantially different adult heights.
Because GWAS are observational, any factor that correlates with both ancestry-linked allele frequencies and the trait of interest — regional diet, healthcare access, cultural practices, even subtle social stratification within a single country — can inflate or distort estimated SNP effects if not carefully modeled. Height GWAS have historically been especially vulnerable to this: a famous 2019 reanalysis (Sohail et al., eLife) showed that apparent genetic signals for height differences between northern and southern Europeans were substantially attenuated once additional population-structure corrections were applied, suggesting earlier estimates had partly captured subtle stratification rather than pure biology.
This is precisely why direct-to-consumer genomics companies present trait predictions as probability distributions or percentile ranges rather than single deterministic values: "your genetics predict you are likely taller than about 65% of people with similar ancestry" — not "you will be exactly 178 cm tall."
For high-heritability, oligogenic traits like blue/brown eye color, that probability distribution is narrow and the prediction reads almost like a fact. For genuinely polygenic, environmentally-modulated traits like height, or disease outcomes like type 2 diabetes where environment dominates, the honest distribution is wide — and responsible reporting should say so explicitly, with confidence intervals, rather than implying false precision.
Understanding a PGS report therefore requires understanding all four preceding stages at once: how many SNPs were discovered, how they were aggregated, how that specific trait's genetic architecture compares to others, and how much of the remaining variance is heritable at all versus shaped by the life a person actually lives.
A 2018 study in Genetics in Medicine found that a majority of consumers misunderstood probabilistic DNA trait and disease-risk reports as deterministic facts — underscoring that the statistical honesty of "R²=40%, not 100%" is only useful if it is actually communicated, and understood, at the point where a customer reads their results.