Hypothesis-free detection of metabolite features from high-resolution LC-MS runs — peak picking, cross-sample alignment, and spectral annotation
Untargeted metabolomics begins with the deceptively simple act of drawing blood or collecting urine — but every downstream feature depends on how that biofluid is handled in the first sixty seconds. Because the goal is to detect the entire small-molecule complement of a sample without any prior hypothesis about which metabolites matter, extraction protocols must be maximally unbiased, highly reproducible, and rigorously quality-controlled from the very first pipette step.
The choice of biofluid dictates which metabolic compartments are visible:
• Plasma/serum: reflects systemic circulating metabolism; EDTA plasma preferred over serum for LC-MS because clotting alters lipid and amino acid profiles • Urine: captures renal clearance products and highly polar, water-soluble metabolites; naturally normalized by creatinine or osmolality • Cerebrospinal fluid: low-volume, low-abundance matrix used for neurometabolic studies • Tissue biopsies: require homogenization and yield compartment-specific (intracellular) metabolite pools
Pre-analytical variability is the single largest source of unwanted variance in metabolomics studies — larger than most true biological effects:
• Time-to-freeze: metabolite concentrations (especially unstable species like lactate, ATP, and lysophospholipids) shift measurably within 30–60 minutes at room temperature • Fasting state: 8–12 h fasting is standard for clinical cohorts to remove diet-driven variance in bile acids, amino acids, and lipids • Hemolysis: ruptured red blood cells release large quantities of purines, nucleotides, and glutathione that swamp true plasma signal • Freeze-thaw cycles: >2 cycles measurably degrade labile lipids and oxidize polyunsaturated fatty acids
Standard operating procedures (SOPs) following the Metabolomics Standards Initiative (MSI) and Consortium standards mandate a documented, harmonized collection tube type, centrifugation protocol (typically 1,300×g, 10 min, 4°C), and time-to-freeze target of <30 minutes.
Cold-solvent protein precipitation is the most widely used untargeted extraction method because it is broadly unbiased across chemical classes:
• Solvent: methanol:water (80:20 v/v) or acetonitrile:water, chilled to −20°C, added at 4:1 solvent:plasma ratio • Mechanism: denatures and precipitates >95% of plasma proteins (albumin, globulins) while polar and moderately non-polar metabolites remain soluble in the supernatant • Centrifugation: 14,000×g, 15 min, 4°C — pellet discarded, supernatant collected for LC-MS • Internal standards (isotope-labeled amino acids, lipids) are spiked in before extraction to monitor recovery and correct for extraction efficiency
Pooled quality-control (QC) samples are the backbone of untargeted metabolomics data quality:
• Construction: equal aliquots (5–20 μL) taken from every single study sample and pooled into one large QC stock • Injected: every 5th–10th sample throughout the acquisition sequence (plus at the start to equilibrate the column) • Function: provides a technical-replicate signal used later for (a) retention-time alignment anchoring, (b) signal drift correction, and (c) feature-level reproducibility filtering (features with QC coefficient of variation >30% are discarded) • Without QC pooling, a study cannot distinguish an interesting biological feature from an analytical artifact — QC design is not optional in any publishable untargeted workflow.
Each biofluid extract is injected onto an ultra-high-performance liquid chromatography (UHPLC) system coupled to a high-resolution mass spectrometer — either an Orbitrap (Thermo Fisher) or a quadrupole time-of-flight (Q-TOF, Agilent/Sciex/Waters) analyzer. In a single 15–20 minute run, tens of thousands of ion signals are recorded across the full mass range, with no pre-selection of target compounds.
No single chromatography mode captures the full chemical diversity of the metabolome, so complementary methods are run in parallel:
• Reversed-phase C18 (RP-LC): retains moderately non-polar to non-polar metabolites (lipids, steroids, fatty acids, many xenobiotics) — sub-2μm particle columns (e.g., Waters BEH C18, 1.7 μm) operated at 0.3–0.5 mL/min with a water/acetonitrile or water/methanol gradient with 0.1% formic acid • Hydrophilic interaction chromatography (HILIC): retains highly polar metabolites poorly captured by RP-LC — amino acids, sugars, nucleotides, organic acids — using a bare-silica or amide-bonded column with a high-organic mobile phase • Ion-pairing or GC-MS: sometimes added for extremely polar species (phosphorylated sugars) or volatile metabolites
Polarity switching: positive-mode ESI (protonated [M+H]⁺ ions, favors basic/neutral metabolites) and negative-mode ESI ([M−H]⁻ ions, favors acidic metabolites) are run as separate injections or alternated scan-to-scan, together roughly doubling metabolome coverage.
The mass analyzer is operated in full-scan (MS1-only) mode across the acquisition, recording every ion reaching the detector without preselecting targets — this is what makes the method "untargeted":
• Resolving power: 60,000–140,000 FWHM (full width at half maximum) at m/z 200 — sufficient to resolve isobaric metabolites differing by only a few mDa (e.g., distinguishing C₃H₄O₃ from C₂H₄N₂O₂ at the same nominal mass) • Mass accuracy: internal or external lock-mass calibration achieves <5 ppm mass error, which is the single most important parameter for confident downstream formula assignment • Scan rate: 8–15 Hz across the LC peak width (typically 3–6 seconds at baseline) ensures ≥10 data points are collected across each chromatographic peak, a requirement for accurate peak-shape-based integration • Data-dependent acquisition (DDA): top-N most intense precursor ions per cycle are selected for MS2 fragmentation, generating spectral evidence used later for annotation • Data-independent acquisition (DIA, e.g., SWATH, AIF): fragments all ions in sequential wide m/z windows, trading spectral purity for complete MS2 coverage of every detectable feature
A single 18-minute run in full-scan mode at 10 Hz across ~1,000 m/z channels produces on the order of 40,000–60,000 individually resolvable ion traces before any peak-picking filtering is applied.
Raw LC-MS data is a three-dimensional surface of intensity across retention time and m/z — far too large and noisy to interpret by eye. Peak-picking algorithms such as centWave (implemented in XCMS and MZmine) scan this surface computationally, using a continuous wavelet transform to identify regions of the chromatogram that have the characteristic Gaussian-like shape of a real chromatographic peak, converting raw signal into a clean, structured feature table.
centWave (Tautenhahn et al., BMC Bioinformatics 2008) is the dominant peak-picking algorithm in untargeted metabolomics because it works directly on high-resolution centroid data without requiring binning into unit-mass bins:
1. Region-of-interest (ROI) detection: the algorithm scans m/z space and groups consecutive scans whose m/z values fall within a user-defined ppm tolerance (typically 2.5–10 ppm) into candidate regions of interest — this exploits the mass accuracy of Orbitrap/Q-TOF data directly 2. Continuous wavelet transform (CWT): each ROI's intensity-vs-retention-time profile is convolved with a Mexican-hat (Ricker) wavelet at multiple scales, which is highly sensitive to the peak-shaped bump of a real chromatographic feature while suppressing single-scan noise spikes 3. Peak boundary refinement: wavelet coefficient maxima mark candidate peak apexes; boundaries are refined by descending the signal until it falls below the noise threshold or a signal-to-noise ratio (SNR) cutoff (commonly 3–10×) is crossed 4. Peak integration: the area under the refined boundaries is integrated to give the reported peak area (proxy for ion abundance), and the intensity-weighted mean m/z and apex retention time are recorded
Key tunable parameters: • ppm — mass tolerance for grouping scans into the same ROI (tighter for higher-resolution instruments) • peakwidth — expected chromatographic peak width range in seconds (e.g., 5–20 s for UHPLC) • snthresh — minimum signal-to-noise ratio for peak acceptance • prefilter — minimum number of consecutive scans above a minimum intensity, used to reject isolated noise spikes before wavelet analysis even runs
Each detected peak becomes one row in a "peak table": accurate m/z, retention time, peak apex intensity, integrated area, and peak width — but within a single sample this table is not yet a "feature table" until it is reconciled across samples (Stage 4).
Sources of false positives that peak pickers must filter out: • Electronic/chemical noise spikes: single-scan intensity blips lacking a coherent peak shape across ≥3 consecutive scans • Isotope peaks: the ¹³C, ³⁴S, and other natural-abundance isotopologues of a real metabolite ion generate their own "features" unless explicitly grouped with the monoisotopic peak (handled by isotope annotation, e.g., CAMERA or the xcms::groupFWHM function) • Adducts and in-source fragments: a single metabolite can generate [M+H]⁺, [M+Na]⁺, [M+NH₄]⁺, and in-source fragment ions, each appearing as a distinct "feature" at a different m/z but identical retention time — annotation tools cluster these back into one compound • Column bleed and background contaminants: plasticizers (e.g., phthalates), polymer additives, and mobile-phase contaminants generate highly reproducible peaks that must be flagged against solvent blank injections
On a typical 120-sample cohort, raw peak picking across all samples can generate 15,000–25,000 candidate peak groups before quality filtering; after removing features present in <50% of QC injections, features with QC CV >30%, and blank-associated contaminants, a robust untargeted study typically retains 3,000–8,000 high-confidence chemical features for statistical analysis.
A metabolite detected at retention time 4.52 minutes in sample 1 may elute at 4.61 minutes in sample 87 due to subtle chromatographic drift — column aging, temperature fluctuation, or mobile-phase batch differences. Before any statistical comparison is possible, every sample's peak table must be non-linearly warped onto a common retention-time axis and every feature's intensity must be corrected for systematic instrument signal drift across the acquisition sequence.
Retention-time alignment operates in two linked steps within tools like XCMS:
1. Initial feature grouping (density-based): peaks from all samples are pooled and grouped into candidate "feature groups" if their m/z values fall within a tolerance window and their retention times fall within a rough bandwidth — using a kernel density estimate across the RT dimension to find regions where many samples have a peak
2. Nonlinear retention-time correction (obiwarp / loess): using either the QC pooled-sample injections or a subset of well-behaved feature groups as anchor points, each sample's entire retention-time axis is nonlinearly warped (dynamic time warping, "obiwarp") onto a shared reference chromatogram. Unlike a simple linear shift, this corrects for the fact that early-eluting and late-eluting compounds can drift by different amounts within the same run.
3. Re-grouping after correction: feature grouping is repeated on the RT-corrected data, now producing tight, well-aligned feature groups — each one representing the same chemical entity measured consistently across all 120 samples
Alignment quality is assessed directly on the QC pooled injections, which — because they are technical replicates of the identical pooled matrix — should show near-zero retention time variance for every feature after correction (target: median RT deviation <5 seconds across an 18-minute run, versus 15–30 seconds of raw uncorrected drift in a long batch).
Even after alignment, raw peak intensities are confounded by instrument sensitivity drift — electrospray source fouling and gradual ion-source contamination cause measured intensity for the same compound to slowly decline (or occasionally spike) across a multi-hour or multi-day acquisition batch.
QC-based robust LOESS signal correction (QC-RLSC): • For each feature independently, a locally weighted regression (LOESS) curve is fit through that feature's intensity values in the QC injections only, as a function of injection order • This curve models the systematic instrument drift trend for that specific feature • Every sample's intensity for that feature is then divided by the interpolated QC trend value at its injection position, removing the drift while preserving true biological variance • Batch effects (if samples were run across multiple days/batches) are handled with ComBat or similar empirical Bayes methods anchored again on QC replicates
Missing value handling: • "True zeros" (metabolite genuinely absent/below detection) must be distinguished from "missing not at random" (present but below the peak-picking SNR threshold in some samples) • Typical untargeted feature tables have 15–30% missingness; common strategies include k-nearest-neighbor imputation, half-minimum value imputation, or explicit left-censored statistical modeling • Features missing in >50% of samples within every study group are typically dropped entirely rather than imputed, since imputation of majority-missing data introduces more noise than signal
After alignment and correction, the median QC coefficient of variation across retained features should fall below 20–30% — the generally accepted quality bar (Broadhurst et al., Metabolomics 2018 consensus guidelines) for a feature table suitable for biological interpretation.
A cleaned, aligned feature table is still just thousands of anonymous m/z–retention time coordinates. The final and most labor-intensive stage of untargeted metabolomics converts these numeric features into named biological molecules through MS/MS spectral matching, and then applies multivariate and univariate statistics to determine which of those identified — or still-unidentified — features actually differ between the phenotypes under study.
Not all "identifications" carry equal certainty — the field uses a standardized 4-tier confidence framework (Sumner et al., Metabolomics 2007, MSI guidelines) to prevent over-claiming:
• Level 1 — Confirmed identification: retention time, accurate mass, AND MS/MS fragmentation spectrum all match an authentic chemical standard run on the same instrument under identical conditions. This is the only level that constitutes a true "identification." • Level 2 — Putatively annotated compound: MS/MS spectrum matches a reference library (METLIN, mzCloud, MassBank, GNPS, NIST) or in-silico prediction with high confidence, but no authentic standard was run in-house • Level 3 — Putatively characterized compound class: only partial structural information available (e.g., "this feature is a lysophosphatidylcholine" based on characteristic fragment losses, but the exact acyl chain is uncertain) • Level 4 — Unknown: a statistically significant, reproducible feature with an accurate mass and MS/MS spectrum but no match to any database — these "unknown unknowns" are often the most biologically novel findings and are prioritized for further structural elucidation (NMR, synthesis, ion mobility CCS matching)
In a typical human plasma untargeted study, roughly 5–15% of features reach Level 1, 15–25% reach Level 2, and the majority remain Level 3/4 — a stark illustration of how much of the human metabolome is still chemically dark matter.
Once the feature table is annotated (or left as anonymous features), statistical analysis identifies which features actually distinguish the biological groups of interest:
Unsupervised exploration: • Principal Component Analysis (PCA): projects the full feature table into 2–3 dimensions capturing maximum variance; used first as a quality-control check — QC pooled samples should cluster tightly at the center, and no single sample should appear as a gross outlier
Supervised modeling: • OPLS-DA (orthogonal partial least squares discriminant analysis): separates variance related to the group of interest (e.g., disease vs. control) from irrelevant "orthogonal" variance; validated by permutation testing (typically 1,000 permutations) and 7-fold cross-validation to guard against overfitting, given that feature counts (thousands) vastly exceed sample counts (tens to hundreds)
Univariate testing and multiple-testing correction: • Each feature is tested individually (t-test, Wilcoxon, or linear model with covariates) between groups • Because thousands of features are tested simultaneously, raw p-values must be corrected for multiple comparisons — the Benjamini-Hochberg false discovery rate (FDR) procedure controlling q<0.05 is standard • Fold-change thresholds (commonly |log2FC|>1) are combined with the FDR-corrected significance in a volcano plot to nominate candidate features
Biomarker panel construction: • Statistically significant, biologically annotated features are combined into multivariate classifier panels (logistic regression, random forest, or PLS-DA) and evaluated by receiver-operating-characteristic (ROC) analysis on a held-out validation cohort • A well-powered untargeted study achieving a validated panel AUC >0.85–0.90 is considered strong evidence for clinical translational potential — for example, bile acid and acylcarnitine panels have reached this bar in NAFLD and cardiovascular risk stratification cohorts.
Untargeted metabolomics inverts the classical hypothesis-driven paradigm: instead of measuring a handful of pre-selected metabolites, the entire small-molecule complement of a biofluid is measured first, and biological meaning is assigned afterward. This is precisely why annotation remains the rate-limiting step of the field — the Human Metabolome Database (HMDB) currently catalogs over 220,000 entries, yet a majority of reproducible, statistically significant LC-MS features detected in any given untargeted study still cannot be matched to a known structure, representing a vast reservoir of undiscovered human and microbiome-derived chemistry.