From tumor shedding to deep-sequencing variant calls — detecting circulating tumor DNA against a background of normal cell-free DNA
Every solid tumor sheds a constant, low-level stream of its own genomic DNA into the circulation — a byproduct of ordinary cell turnover (apoptosis, necrosis, and to a lesser extent active secretion) that gives liquid biopsy its biological foundation, but also its fundamental sensitivity limit: you can only detect what the tumor sheds.
Cell-free DNA does not float freely in plasma — it is protected from rapid nuclease degradation by wrapping around a histone octamer as a nucleosome, and cfDNA fragmentation patterns closely mirror nucleosome positioning at the moment of cell death, producing a characteristic ~166 bp modal fragment length (147 bp of nucleosomal DNA plus ~20 bp of linker).
Tumor-derived ctDNA fragments frequently show a subtly shorter modal length (often peaking closer to 145 bp) than normal cfDNA, likely reflecting altered chromatin packaging and nuclease accessibility in cancer cells — this fragmentomics signal is now actively exploited as an independent, sequence-free enrichment strategy: simply size-selecting shorter fragments during library prep can enrich the relative proportion of tumor-derived molecules before sequencing even begins.
The pre-analytical handling of a liquid biopsy sample is unusually unforgiving: because the analyte of interest is a tiny minority fraction of an already dilute molecule, any contamination from lysed white blood cells releasing their own high-molecular-weight genomic DNA can overwhelm and mask the ctDNA signal entirely.
Standard EDTA blood tubes are unsuitable for delayed liquid biopsy processing because white blood cells begin lysing within hours, flooding the plasma with genomic DNA from normal leukocytes that dilutes the already-scarce ctDNA fraction by orders of magnitude — this is why cfDNA-specific stabilizing tubes (containing a fixative that keeps white cells intact) are standard for any sample that will not be processed within a few hours.
After collection, plasma is separated by double centrifugation (a low-speed spin to pellet cells, followed by a high-speed spin on the supernatant to remove residual cellular debris and platelets) before cfDNA extraction, typically via magnetic bead-based silica columns optimized for short-fragment recovery — recovery efficiency directly determines the effective input material available for the downstream sequencing assay, and losses here cannot be recovered by any amount of downstream signal processing.
Standard PCR-based sequencing cannot distinguish a true rare mutation present in the original DNA from a PCR or sequencing error introduced during library preparation — at the ultra-low allele frequencies relevant to early ctDNA detection, these error rates are actually higher than the signal being sought. Unique Molecular Identifiers solve this by molecular bookkeeping.
Before PCR amplification, each individual cfDNA molecule is ligated to an adapter carrying a short random UMI sequence — because these barcodes are randomly generated, the probability that two different starting molecules receive the identical UMI is vanishingly small, meaning every read that shares a UMI (a "UMI family") can be traced back to PCR copies of one single original DNA molecule.
After sequencing, all reads sharing a UMI are collapsed into a single consensus call: if 20 PCR-duplicate reads from one UMI family agree on a base, that call is treated as a true reflection of the original molecule; if only 1 of 20 reads shows a variant, that is recognized as a PCR or sequencing error and discarded rather than called as a mutation. Duplex sequencing extends this further by independently tagging and comparing both DNA strands of the original double-stranded molecule, since true biological mutations should appear on both strands (or complementary positions) while most chemical damage and enzymatic errors are strand-specific — this pushes achievable error rates below one error per 10⁷–10⁸ base pairs, several orders of magnitude better than raw sequencing accuracy.
Without UMI-based consensus calling, standard NGS error rates (~0.1–1% per base) would generate false-positive "mutations" far more often than a real ctDNA variant at 0.01–0.1% VAF actually appears — UMI technology is not an optimization, it is the precondition that makes ultra-sensitive ctDNA detection possible at all.
Detecting a mutant allele present in only 1 in 1,000 or 1 in 10,000 DNA molecules is fundamentally a sampling statistics problem: you need enough independent molecular observations at that genomic position to have a realistic chance of capturing even a single copy of the rare mutant allele, then enough redundant coverage of that same molecule to confirm it is real rather than error.
Sequencing depth alone cannot rescue a sample with too little input material — if only 1,000 unique genome-equivalents of cfDNA were recovered from a blood draw and the true ctDNA VAF is 0.05%, there is roughly only a 40% chance any given draw even physically contains one copy of the mutant molecule, no matter how many times you re-sequence the PCR-amplified copies of what you started with (which are, by definition, technical replicates of the same limited input, not independent biological samples).
This is why ultra-high-depth sequencing (30,000×+) is deployed specifically for narrow, clinically actionable gene panels rather than genome-wide — concentrating a fixed budget of sequencing reads onto a small number of high-value positions maximizes the chance of both capturing and confirming a rare variant at those specific loci, at the cost of not surveying the rest of the genome. Whole-genome or whole-exome ctDNA approaches (used for tumor fraction estimation or minimal residual disease monitoring via personalized variant panels) instead spread coverage thinly across many positions, trading per-locus sensitivity for breadth.
Every technical advance in liquid biopsy — better UMIs, deeper sequencing, cleverer bioinformatics error suppression — ultimately runs into the same hard biological ceiling: a small, early-stage tumor simply does not shed enough DNA into circulation for any assay, however sensitive, to reliably detect it from a standard blood draw volume.
ctDNA shedding scales roughly with total tumor cell mass and tumor vascularization/necrosis rate — a small, well-contained Stage I tumor of a few cubic centimeters sheds proportionally far less DNA than a large, necrotic, or metastatic Stage IV tumor, which is why sensitivity for early-stage cancer detection remains meaningfully lower (roughly 50–70% for even leading multi-cancer early detection assays) than for late-stage or minimal-residual-disease monitoring in a patient with known prior tumor genomic profile.
Because a standard blood draw yields only a finite, fixed number of cfDNA molecules (on the order of a few thousand genome-equivalents from ~10 mL of blood), there is a hard information-theoretic floor beneath which no amount of assay sensitivity improvement can help — if a tumor sheds so little DNA that a typical blood draw contains zero copies of the mutant allele on average, that sample will read negative regardless of sequencing depth or error suppression quality. This is why minimal residual disease monitoring assays after curative-intent surgery often use serial sampling over time and patient-specific variant panels derived from the resected tumor tissue, rather than trying to solve early single-draw detection sensitivity through sequencing technology alone.
The fundamental sensitivity ceiling of liquid biopsy is set by tumor biology (how much DNA a tumor sheds) and blood draw volume (how many molecules can physically be sampled) — not by sequencing technology, which is why serial monitoring and larger blood volumes are active areas of clinical protocol development alongside continued assay refinement.