HomeArticlesBiology

Population Genetics and Gene Flow

Mathematical and empirical study of genetic variation within and between populations

mysimulator teamUpdated June 2026≈ 9 min read▶ Open the simulation

Introduction to Population Genetics

Population genetics is the mathematical and empirical study of genetic variation within and between populations—describing how allele frequencies change over time through the evolutionary forces of mutation, genetic drift, natural selection, and gene flow. The foundations were laid by Sewall Wright, J.B.S. Haldane, and Ronald Fisher in the 1920s-30s unifying Mendelian genetics with Darwinian evolution (the Modern Synthesis). The Hardy-Weinberg equilibrium (HWE) principle provides the null model: in an idealised infinite, randomly-mating, non-selected population, allele frequencies remain constant across generations (p2 + 2pq + q2 = 1 for any two alleles p and q). Deviation from HWE in observed populations reveals natural selection, non-random mating, population structure, or genotyping errors.

Population genetics is fundamental to evolutionary biology, human genetics, conservation genetics, and evolutionary medicine. Genomic tools—whole genome sequencing of thousands of individuals and population-scale biobanks (UK Biobank with 500,000 genomes; gnomAD with 730,000 exomes and genomes)—have transformed population genetics from theoretical modelling to empirical genome-wide analyses. Population genetic theory predicts that beneficial mutations spread rapidly, neutral mutations drift, and deleterious mutations are removed by purifying selection—determining the relative abundances of different variant types observed in population sequencing databases and constraining genetic interpretation in clinical genomics.

Forces of Evolution

Genetic Drift

Genetic drift is the stochastic change in allele frequencies due to random sampling in finite populations—the dominant force in small populations. The effective population size (Ne) determines drift magnitude: neutral allele fixation probability equals its current frequency (1/2Ne for a new mutation); fixation time averages 4Ne generations. Population bottlenecks (severe reduction in population size—founder effect, disease, warfare) dramatically reduce Ne, causing random fixation or loss of alleles unrelated to their adaptive value. Ashkenazi Jewish, Finnish, and Amish populations show distinctive disease gene enrichments from founder effects. Island populations show accelerated genetic drift. Selection-drift balance determines the likelihood of slightly deleterious mutations accumulating in small populations—relevant for understanding purifying selection limits and species extinction risk.

Natural Selection

Natural selection—differential survival and reproduction based on genotype—changes allele frequencies directionally. Positive (directional) selection increases beneficial allele frequency; negative (purifying) selection removes deleterious alleles; balancing selection maintains multiple alleles (heterozygote advantage, negative frequency-dependent selection). Classic positive selection signature: selective sweep—a rapidly spreading beneficial allele carries surrounding neutral variants to high frequency, creating a region of reduced polymorphism and characteristic long haplotype blocks detectable by elevated LD and Fst statistics. Population differentiation for traits under local adaptation (skin pigmentation responding to UV, lactase persistence in dairy-farming populations, high-altitude haemoglobin variants in Andean and Tibetan populations) represents documented positive selection in humans detectable by GWAS and genome scans.

жива демонстрація · пов'язана симуляція● LIVE

Population Structure and Gene Flow

Principal Component Analysis

Principal component analysis (PCA) of genome-wide SNP data visualises population structure—individuals genetically cluster by continental ancestry with PC1-PC2 axes often separating European, African, East Asian, and South Asian ancestry groups reflecting historical divergence and migration. Patterson's EIGENSOFT was the first widely-used PC-based population stratification tool; GENESIS and PLINK2 are modern implementations. PCA is essential in GWAS to control for population stratification (systematic allele frequency differences between cases and controls due to ancestry rather than disease association creating false positive associations). Admixture analysis (STRUCTURE, ADMIXTURE programs) infers ancestral population fractions for each individual—revealing recent admixture events detectable through characteristic patterns in chromosome-length haplotype ancestry segments.

Gene Flow and Migration

Gene flow (migration) homogenises allele frequencies between populations, opposing local adaptation by genetic drift. Ancient DNA analysis transformed understanding of human population history: the peopling of the Americas ~15,000 years ago from Siberian population via Beringia; multiple waves of West Eurasian migration into Europe (Anatolian farmers ~8000 years ago, Yamnaya steppe pastoralists ~5000 years ago) displacing hunter-gatherers; Denisovan introgression in Papua New Guinean and Aboriginal Australian populations contributing adaptive immune variants. The Fst statistic (fixation index) measures population differentiation across allele frequencies: Fst = 0 indicates identical populations; Fst = 1 indicates complete fixation of different alleles. Global human Fst (~0.12-0.15) reflects modest differentiation consistent with recent common ancestry of all living humans.

Linkage Disequilibrium and Haplotypes

Linkage disequilibrium (LD)—the non-random association between alleles at different loci—provides the basis for GWAS (typing tag SNPs to impute surrounding associated variants) and reflects historical recombination, selection, and demographic history. LD measured by D' (normalized disequilibrium) and r2 (correlation coefficient between alleles) decays with physical distance and population age—older populations (African) show faster LD decay (shorter haplotype blocks ~20 kb) than younger populations (European ~80 kb, isolated island populations ~200+ kb). 1000 Genomes Project, gnomAD, and UK Biobank provide continent-specific LD reference panels enabling trans-ethnic GWAS and increasing fine-mapping resolution by exploiting variable LD structure across ancestries to narrow credible variant sets.

Examples and Applications

Example 1: Sickle Cell Disease and Balancing Selection

Sickle cell trait (HbAS heterozygotes) provides ~70% protection against severe malaria by causing infected RBCs to sickle prematurely, impairing Plasmodium falciparum growth and RBC adherence. The heterozygote advantage maintains HbS (rs334, beta-globin E6V) at high frequency in malaria-endemic Africa, Middle East, and Mediterranean despite HbSS homozygotes bearing severe haematological disease—a classic overdominant balancing selection example. Geographic distribution of HbS frequency correlates closely with historical malaria endemicity—the frequency correlates with estimated malaria selection pressure over millennia. Multiple independent origins of the same HbS mutation on different haplotypes (Benin, Senegal, Bantu, Cameroon haplotypes) provide textbook evidence for repeated independent selection of the same beneficial mutation.

Example 2: Human Population Bottlenecks

Genomic evidence reveals a severe human bottleneck approximately 900,000-800,000 years ago (Lu et al. 2023, Science)—reducing the ancestral human population to approximately 1,280 individuals for ~117,000 years—explaining a near-complete gap in hominin fossils and the loss of genetic diversity in comparison to great ape relatives. The Out-of-Africa bottleneck (~50,000-70,000 years ago) further reduced diversity of non-African populations to ~60-80% of African populations' diversity. Each bottleneck left lasting signatures: reduced heterozygosity, longer haplotype blocks, higher LD, elevated frequency of slightly deleterious alleles that escaped purifying selection, and higher frequency of rare recessive disease alleles in isolated populations descending from the bottleneck.

Example 3: Selective Sweeps in Modern Humans

Genomic scans for positive selection identify regions with elevated Fst (differentiation between populations), extended haplotype homozygosity (EHH—long haplotypes carrying the selected allele to high frequency faster than recombination can break them down), and reduced nucleotide diversity. Classic sweeps in humans: the LCT region (lactase persistence—13910*T allele in Europeans, −14010*G in East Africa) shows the strongest selective sweep signals in the genome in European and East African populations respectively; EPAS1 region in Tibetans (conferring high-altitude adaptation through reduced haemoglobin synthesis preventing polycythaemia) carries a Denisovan-introgressed haplotype and shows largest Fst population differentiation of any human gene; SLC24A5 for skin pigmentation; FOXP2 for speech and language in human lineages; AMY1 for salivary amylase adaptation to starchy diets.

Example 4: Ancient DNA Revolution

Ancient DNA extraction from prehistoric skeletal remains, teeth, and sediments combined with whole genome sequencing has transformed understanding of human and faunal population history. Key findings: archaic human introgression—Neanderthal contribution to non-African genomes (1-4%) identified by Pääbo's group (2010 Nature), earning a 2022 Nobel Prize; Denisovan DNA in South/East Asian and Oceanian populations (3-6%); ghost archaic populations with signal in sub-Saharan Africa with no sampled ancient or modern genomes. Ancient pathogen DNA (Yersinia pestis in medieval plague victims, SARS-CoV-2 ancestral strains, hepatitis B virus in Neolithic Europeans) reveals past epidemics and evolutionary origins. Johannes Krause, David Reich, and Eske Willerslev lead laboratories transforming global prehistory through palaeogenomics.

Example 5: Consanguinity and Inbreeding

Inbreeding—mating between relatives—increases homozygosity across the genome (runs of homozygosity, ROH). ROH burden measured from genome-wide SNP arrays quantifies inbreeding coefficient F—the probability a random position is autozygous (both alleles inherited from a common ancestor). First-cousin marriages increase F by 1/16 (0.0625); in populations routinely practising consanguineous marriage (Middle East, South Asia—affecting 30-50% of marriages), inbreeding-related recessive disease is substantially elevated. ROH analysis identifes new autosomal recessive disease genes by homozygosity mapping: affected individuals from consanguineous families share a single autozygous ROH region (often <5 Mb) containing the disease gene. Recessive conditions enriched in consanguineous populations include many metabolic, neurological, and eye diseases—providing a human genetics discovery approach complementing linkage and exome sequencing.

Example 6: Conservation Genetics

Population genetics tools preserve endangered species by characterising genetic diversity, identifying distinct conservation units, assessing inbreeding depression, and guiding captive breeding to maintain genetic diversity. Cheetah (Acinonyx jubatus) has exceptionally low genetic diversity (~100x less than domestic cats) from a severe bottleneck during the last glacial maximum—making it vulnerable to infectious disease outbreaks. Florida panthers—highly inbred sub-population—benefited from genetic rescue: introductions of 8 Texas puma females restored genetic diversity, reducing kinked tails and cardiac defects from inbreeding depression, increasing kitten survival and population growth. Genomic tools enable identifying evolutionarily distinct population segments (ESUs) for conservation management and tracking founder contributions in captive breeding programmes.

Example 7: Polygenic Risk Scores

Polygenic risk scores (PRS) aggregate thousands to millions of GWAS-associated SNPs weighted by their effect sizes into a single genome-wide score for complex trait prediction. PRS for CAD, T2DM, breast cancer, and schizophrenia identify individuals at 3-5x elevated risk in the top centile relative to average—comparable to or exceeding monogenic risk variants in predictive value for common diseases. PRS portability across ancestries is limited—PRS trained in European GWAS cohorts performs poorly in African or South Asian ancestry due to differences in LD, allele frequencies, and causal variant architecture. Ancestry-specific or multi-ancestry PRS training in diverse cohorts (Million Veteran Programme, All of Us, H3Africa) improves predictive performance across global populations. Clinical PRS implementation (Genomics England, NHS Genomics, UK 100,000 Genomes) for cancer and cardiovascular risk stratification is advancing toward routine screening.

Example 8: Mutation Rate Estimation

The human germline de novo mutation rate—spontaneous new mutations arising per generation—is approximately 1.2×10^-8 per base pair per generation (38 de novo mutations per diploid newborn), estimated from parent-offspring whole genome sequencing in large trios. The paternal mutation rate is ~4x higher than maternal—correlating with the greater number of cell divisions in male spermatogenesis (~30/year from puberty). De novo mutations are the dominant source of new mutations in human populations and disproportionately contribute to dominant Mendelian disease (85% of severe intellectual disability de novo). The mutation rate calibrates the molecular clock linking genetic divergence to time—enabling human-chimpanzee divergence (~6 million years ago) and within-species timeline estimates from sequence divergence to coalescence theory demographic modelling.

Try it live

Everything above runs in your browser — open Michaelis-Menten Kinetics and change the parameters while it is running. Nothing is installed, nothing is uploaded, the whole model lives in one tab.

▶ Open Michaelis-Menten Kinetics simulation

What did you find?

Add reproduction steps (optional)