🔍 DNA Match Random Match Probability Simulator
This simulation calculates the statistical probability of a random DNA match between profiles.
DNA Profile Obtained From STR Loci
A lab reads allele pairs at each STR marker.
- 20: CODIS core loci (Standard FBI panel used nationwide.)
- 2: Alleles per locus (One copy from each parent.)
- ~80%: Typical heterozygosity (Most loci show two distinct peaks.)
- ~2 hrs: Profile turnaround (PCR then capillary electrophoresis.)
What an STR locus is
STRs are short repeated DNA sequences.
Repeat counts vary widely between people.
Each locus sits at a known chromosome position.
Building the profile
PCR amplifies each targeted locus.
Capillary electrophoresis sorts fragments by size.
A genotype (two alleles) is called per locus.
CODIS uses 20 core loci plus a sex marker.
Single-Locus Match Probability
Each locus has its own population match frequency.
- 1 in 10–100: Typical locus probability (Depends on allele diversity.)
- Lower probability: Rarer alleles (Uncommon repeats are more distinctive.)
- Population database: Allele frequency source (Frequencies vary by reference group.)
- p² + 2pq: Hardy-Weinberg use (Combines two allele frequencies.)
Where probabilities come from
Population studies measure allele frequencies.
Genotype probability follows Hardy-Weinberg equilibrium.
Common alleles give higher match chances.
Independence between loci
Loci sit on different chromosomes or far apart.
This makes their probabilities statistically independent.
Independence is what enables multiplication later.
Independence is the assumption the whole method relies on.
Loci Multiplication — The Product Rule
Independent locus probabilities multiply together.
- P₁ × P₂ × … × Pₙ: Rule applied (Standard forensic statistics method.)
- Order-of-magnitude drop: Effect per added locus (Each locus shrinks the odds fast.)
- ~1 in 10¹³–10¹⁵: 13 loci (FBI legacy set) (Astronomically rare by design.)
- Statistical independence: Assumption required (No linkage between chosen markers.)
Why multiply at all
For independent events, joint probability multiplies.
Each extra locus compounds the rarity.
Few loci already yield very small numbers.
Sensitivity to loci count
More loci compared means lower combined probability.
Fewer loci leave more room for coincidence.
Modern kits use 20+ loci routinely.
Doubling loci roughly squares the rarity of the match.
Random Match Probability (RMP)
One number now expresses the entire profile's rarity.
- P(match by chance): RMP definition (Across all loci compared.)
- ~1 in 10¹⁸+: Standard full-profile RMP (Common with 20-locus kits.)
- Scientific notation: Reported as (Numbers too large for plain digits.)
- Weight of evidence: Used in court as (Not absolute proof of identity.)
Reading the RMP
RMP answers one narrow question.
How likely is an unrelated match by chance.
It is not the chance of innocence.
Reporting conventions
Labs report RMP in scientific notation.
Jurors see phrases like "1 in a trillion."
Smaller RMP means stronger identification support.
A vanishing RMP does not by itself prove guilt.
Statistical Interpretation & Confidence
Extremely low RMP values support confident identification.
- ~8.1 billion: World population (2026 est.) (Reference scale for comparison.)
- ~7.5×10¹⁸: Grains of sand on Earth (Common rarity benchmark used.)
- RMP invalid: Identical twin caveat (Twins share full STR profiles.)
- Adjust for search size: Database search caveat (More comparisons raise coincidence odds.)
Putting RMP in context
Comparing RMP to world population clarifies scale.
Many profiles are rarer than the population itself.
This is why full profiles are so persuasive.
Limits of the statistic
Relatives share more alleles than strangers.
Lab error rates matter more than tiny RMPs.
Database trawls need statistical correction.
RMP measures coincidence, not laboratory or relative risk.
This simulation calculates the statistical probability of a random DNA match between profiles.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install