Slow Off-rate Modified Aptamer (SOMAmer) reagents multiplexed against a DNA microarray — quantifying thousands of plasma proteins from 55 µL of sample in a single run
The human plasma proteome is famously difficult to survey because protein concentrations range across roughly eight orders of magnitude — from albumin at ~40 mg/mL down to cytokines such as IL-6 circulating at low pg/mL. SOMAscan solves this dynamic-range problem not by depleting abundant proteins (as older 2D-gel and mass-spec workflows did) but by running the same 55 µL plasma or serum sample at three parallel dilutions, each optimized for a different abundance tier, then recombining the data computationally.
Classical plasma proteomics workflows physically remove the ~10–20 most abundant proteins (albumin, immunoglobulins, transferrin, haptoglobin) using immunoaffinity depletion columns before mass-spectrometry analysis, because these proteins otherwise swamp the detector. Depletion is lossy and variable: it co-depletes proteins bound to albumin's hydrophobic pockets, and recovery is not perfectly reproducible plate-to-plate.
SOMAscan instead exploits the fact that each SOMAmer is a discrete, independent binding reaction — abundant proteins do not compete with rare proteins for signal the way they do on a mass spectrometer's detector. The dynamic-range problem is solved purely by running three dilutions of the same sample in parallel:
• 40% dilution — captures low-abundance analytes (cytokines, chemokines, growth factors) in the pg/mL–ng/mL range • 1% dilution — captures mid-abundance analytes (most signaling proteins, enzymes) in the ng/mL–µg/mL range • 0.005% dilution — captures high-abundance analytes (acute-phase proteins, complement, clotting factors) in the µg/mL–mg/mL range
Each SOMAmer in the 7,289-member library is pre-assigned to the dilution bin that gives it the cleanest quantitative window, based on the median concentration of its target protein in reference plasma pools.
Because SOMAscan is run at CLIA-certified commercial labs (SomaLogic, Boulder CO) and in academic core facilities on 96-well or 384-well plates, pre-analytical handling is tightly controlled:
• Collection tubes: EDTA or heparin plasma preferred over serum for most panels (less ex vivo platelet activation artifact) • Time to processing: centrifugation and freezing within 2 hours recommended to avoid coagulation-cascade drift • Freeze-thaw cycles: limited to ≤2; repeated freeze-thaw measurably degrades ~5–8% of labile analytes • Hemolysis index: samples above threshold flagged, since hemolysis releases intracellular proteins (e.g., hemoglobin fragments) that create false elevations • Batch design: samples from a single clinical cohort are randomized across plates to avoid confounding batch effects with disease status — a mistake that has invalidated more than one early biomarker discovery paper
A SOMAmer is not a natural DNA aptamer. It is a single-stranded oligonucleotide in which select thymidine bases are replaced with modified deoxyuridine analogs carrying bulky, hydrophobic, protein-like side chains — benzyl, naphthylmethyl, tryptaminocarbonyl, or isobutyl groups. These side chains let the folded aptamer present chemistry that mimics antibody CDR loops, achieving affinities and specificities that plain unmodified DNA or RNA aptamers cannot reach, with equilibrium dissociation constants commonly in the picomolar-to-low-nanomolar range.
SOMAmers are discovered by an enhanced form of SELEX (Systematic Evolution of Ligands by EXponential enrichment):
1. A randomized library of ~10¹⁴ modified-DNA sequences (40 random positions flanked by fixed PCR primer sites) is incubated with the immobilized target protein. 2. Binders are partitioned from non-binders, PCR-amplified using polymerases engineered to incorporate the modified dU bases, and the library is re-challenged for 8–12 rounds. 3. Critically, a slow-off-rate partitioning step is added: after equilibrium binding, the target-aptamer mixture is challenged with an excess of unmodified competitor DNA for a defined dwell time. Only complexes with slow dissociation kinetics (long residence time) survive this wash and are recovered — this is precisely what gives the reagent its name. 4. The surviving sequence pool is sequenced (high-throughput sequencing of the SELEX output), individual candidate sequences are synthesized and validated for affinity (Kd), specificity (cross-reactivity against related family members), and reproducibility across lots.
The result is a reagent class with off-rates (koff) an order of magnitude slower than typical unmodified aptamers, translating directly into higher effective affinity and a much lower nonspecific background — the property that later makes the kinetic-challenge wash step in Stage 3 so effective.
All 7,289 SOMAmers for a given dilution bin are pooled and added simultaneously to the diluted plasma sample in a single well — the entire binding reaction, all thousands of independent protein-aptamer equilibria, occurs in the same 3.5-hour, room-temperature incubation:
• Binding buffer contains blocking agents and a fixed ionic strength to standardize electrostatic contributions to binding across all 7,289 reagents • Each SOMAmer carries a 5′ biotin tag, allowing later capture of the whole aptamer-protein complex on a streptavidin-coated surface regardless of which protein it has bound • Because binding is at true thermodynamic equilibrium (not a fixed endpoint kinetic snapshot), signal intensity for a given SOMAmer is a monotonic function of its target's free plasma concentration — this is what allows the assay to be quantitative, not just qualitative • Cross-reactivity is empirically characterized for every SOMAmer against structurally related paralogs (e.g., IL-6 family cytokines) during reagent validation, and any SOMAmer with unacceptable cross-reactivity is excluded from the production panel
The defining engineering trick of the SOMAscan platform is a bead-capture-and-release cascade that progressively strips away everything except a clean, protein-concentration-proportional pool of free SOMAmer DNA — converting an analog protein-binding measurement into a digital, PCR/hybridization-compatible nucleic-acid signal that can be read on a DNA microarray instead of requiring 7,289 separate antibody-based detection reactions.
After equilibrium binding (Stage 2), unbound and weakly-bound material is removed through a sequence of physically and chemically distinct steps:
1. Primary capture: the entire binding reaction is mixed with streptavidin-coated beads. Every biotinylated SOMAmer — whether bound to its target protein or still free — attaches to a bead. Unbound plasma proteins that never engaged a SOMAmer are washed away at this point.
2. Kinetic (polyanionic) challenge: beads are resuspended in a solution containing a large excess of dextran sulfate, a highly negatively charged polyanion that competes with any protein-SOMAmer interaction driven mainly by nonspecific electrostatic attraction. Because SOMAmers were originally selected (Stage 2) for slow dissociation kinetics with their true target, genuine complexes survive this wash while nonspecific low-affinity complexes are stripped off within the wash dwell time — this single step is responsible for most of the platform's specificity.
3. Photocleavable release: each SOMAmer carries a photocleavable linker between its biotin tag and the oligonucleotide body. Illumination at ~300 nm cleaves this linker, releasing the SOMAmer — with its bound protein still attached — from the bead into solution, while a fresh biotin (present on the protein side of the cleaved linker) remains behind for the next step.
The kinetic-challenge wash is the single largest contributor to SOMAscan specificity. Because SOMAmers were explicitly selected during SELEX for slow off-rates against their cognate target, a brief dextran-sulfate competition wash removes the vast majority of nonspecific complexes (>10-fold signal reduction for random protein-aptamer pairs) while leaving true target complexes essentially intact — this is what lets a single well tolerate 7,289 simultaneous independent binding reactions without cross-talk.
After photocleavable release, the mixture still contains a small fraction of SOMAmers that never found their target protein. A second capture step removes them:
• The newly exposed biotin (now on the released protein, not the original bead-linked tag) is used to recapture only complexes where a protein is actually present — i.e., only SOMAmers that are genuinely bound to a target protein get pulled onto a fresh streptavidin surface a second time • Free SOMAmer that never bound protein has no biotin left after photocleavage and is washed away • The protein is then denatured and removed, releasing pure SOMAmer DNA — now present in solution at a concentration directly proportional to how much of its target protein was originally in the plasma sample • This two-capture design is what gives SOMAscan its low background: a spurious signal requires a SOMAmer to survive kinetic challenge AND retain a protein through the second capture, a vanishingly unlikely combination for a non-specific interaction
Once purified, each of the 7,289 SOMAmer DNA sequences is present at a concentration proportional to its target protein's original plasma abundance. Because the readout is now DNA, not protein, it can be quantified the same way a gene-expression microarray is — by hybridization to a custom-synthesized array of complementary capture probes, followed by confocal fluorescence scanning. This is the step that gives SOMAscan its throughput: thousands of independent protein measurements are read out in one scan instead of one immunoassay well per protein.
The purified SOMAmer pool is applied to a custom microarray manufactured with the same in-situ oligonucleotide synthesis technology used for gene-expression arrays (Agilent SurePrint), except here each array feature carries a probe sequence complementary to exactly one SOMAmer's unique 40-mer variable region:
• Each of the 7,289 SOMAmers has its own unique sequence "barcode," so hybridization is highly specific — cross-hybridization between different SOMAmer barcodes is negligible by design • Hybridization proceeds for several hours at controlled temperature and stringency (salt concentration, formamide where used) to ensure only fully complementary SOMAmer-probe pairs remain duplexed • A wash step removes any SOMAmer that hybridized to a mismatched address • Built-in hybridization control SOMAmers (spiked in at fixed, known concentrations) occupy dedicated array features and are used downstream to normalize for well-to-well and array-to-array hybridization efficiency differences
The hybridized array is imaged on a confocal laser microarray scanner (e.g., an Agilent SureScan-class instrument), exciting the Cy3 fluorophore and capturing emission through a photomultiplier tube at high spatial resolution:
• Scan resolution of roughly 3 µm per pixel resolves each array feature (spot) cleanly from its neighbors • Raw image is converted to a 16-bit TIFF; feature-extraction software grids the image, identifies each of the 7,289+ feature positions, and integrates pixel intensity within each spot to produce a single raw Relative Fluorescence Unit (RFU) value per SOMAmer per sample • Local background (inter-feature signal) is subtracted per spot • Flagging rules exclude spots with saturation, high local-background variance, or physical array defects (scratches, dust, bubbles) before the value is passed downstream • Typical raw RFU values for a single run span roughly 10¹ to 10⁵ across the full concentration range of the panel — this huge per-feature dynamic range is why the three-dilution-bin design (Stage 1) is needed in the first place: even within a single dilution bin, RFU output must stay within the scanner's linear response range
Raw fluorescence values are not directly comparable across wells, plates, or studies until they pass a defined normalization and quality-control pipeline (SomaLogic's ADAT processing). Only after this step does the data become suitable for differential-abundance testing and, ultimately, for building and validating clinical biomarker panels — the whole point of running a 7,000-plex proteomic assay in the first place.
SomaLogic's standard ADAT (Analyte Data) processing pipeline applies normalization in a fixed order, each stage correcting a different source of technical variance:
1. Hybridization normalization: each sample's signal is scaled against the built-in hybridization control SOMAmers spiked at known concentration, correcting for well-to-well differences in hybridization efficiency on the array.
2. Median signal normalization: within each dilution bin, every sample's SOMAmer signals are scaled so the median RFU across all analytes matches a reference value, correcting for systematic differences in total protein input, pipetting, or overall assay efficiency between samples.
3. Plate-scale (calibration) normalization: a panel of calibrator samples run on every plate is compared against a fixed reference standard, generating a per-analyte scale factor that corrects for plate-to-plate and lot-to-lot reagent drift — this is what allows data generated months apart, or at different sites, to be pooled statistically.
After all three steps, remaining technical coefficient of variation is typically under 5% within a plate and under 10% between plates for the majority of the 7,289 analytes — precision that rivals a well-optimized singleplex immunoassay, despite measuring thousands of proteins simultaneously.
Normalized, log2-transformed RFU values are treated as a high-dimensional expression matrix, analyzed with the same statistical toolkit used for transcriptomics:
• Differential abundance: linear models (limma-style) or nonparametric tests compare case vs. control groups per analyte; Benjamini-Hochberg false-discovery-rate correction is applied across all ~7,289 simultaneous tests to control for multiple comparisons • Effect-size filtering: candidates are typically required to clear both a significance threshold (q<0.05) and a minimum fold-change, to avoid chasing statistically significant but clinically trivial differences • Panel building: multivariate models (LASSO-penalized logistic regression, random forests, Cox proportional-hazards for survival endpoints) combine dozens of individual SOMAmer signals into a single composite biomarker score, evaluated by ROC-AUC or C-statistic on held-out validation cohorts • Orthogonal confirmation: top candidate proteins are re-measured by an independent method — Olink proximity extension assay (PEA), targeted mass spectrometry (PRM-MS), or conventional ELISA — before a biomarker claim is considered validated, since aptamer-based and antibody-based platforms occasionally disagree for individual analytes due to epitope or conformational differences
SOMAscan-derived biomarker panels have moved beyond discovery into clinical products: SomaLogic's SomaSignal tests combine dozens to hundreds of individual SOMAmer measurements into single composite scores for outcomes such as cardiovascular event risk, kidney function (eGFR), liver fibrosis, and diabetes risk, typically achieving ROC-AUC values of 0.85–0.95 in validation cohorts — rivaling or exceeding traditional clinical risk scores while requiring only a single 55 µL blood draw and no disease-specific assay redevelopment.