Introduction to Proteomics
Proteomics is the large-scale study of proteins—their identities, abundances, modifications, interactions, structures, and functions—at the global level. While genomics provides a static genetic blueprint, the proteome is dynamic: it varies between cell types, developmental stages, and physiological and pathological conditions; it is shaped by post-translational modifications (phosphorylation, ubiquitination, acetylation, glycosylation) that dramatically alter protein function; and protein levels correlate only roughly with mRNA levels due to translational regulation and protein stability differences. The proteome is thus the functional executor of cellular processes, making it biologically and clinically essential to characterise.
Modern proteomics is driven by mass spectrometry (MS) technology—particularly high-resolution liquid chromatography-tandem MS (LC-MS/MS)—which can identify and quantify thousands of proteins from complex mixtures. Bioinformatic database searching matches experimental mass spectra to theoretically predicted peptide fragment masses. A single LC-MS/MS experiment can quantify over 10,000 proteins in a cell lysate; sensitivity continues improving toward single-cell proteomics. Structural proteomics—mapping protein three-dimensional structures at proteome scale using cryo-EM and AlphaFold-assisted interpretation—adds conformational information to abundance data.
Mass Spectrometry in Proteomics
Shotgun Proteomics
Bottom-up (shotgun) proteomics digests protein mixtures with trypsin generating peptides, separates them by reverse-phase liquid chromatography, and analyses them by tandem MS. MS1 scans measure peptide m/z values (precursor masses); MS2 fragmentation spectra identify amino acid sequences. Database searches (Mascot, Sequest, MaxQuant) match spectra to protein databases. Label-free quantification uses peptide peak intensities or spectral counts for relative abundance; stable isotope labelling (SILAC incorporating heavy lysine/arginine in cell culture; TMT chemical labelling) enables multiplexed quantitative comparison of up to 18 samples simultaneously. Sensitivity advances enable quantitative proteomics from as few as 100-1000 cells.
Post-Translational Modification Proteomics
Over 600 post-translational modification (PTM) types are known; phosphorylation is the most studied. Phosphoproteomics enriches phosphopeptides using metal oxide affinity chromatography (MOAC) or immunoprecipitation with anti-phosphotyrosine antibodies before LC-MS/MS. Large-scale phosphoproteomics identifies thousands of regulated phosphorylation sites across cell signalling perturbations—kinase-substrate networks, drug mechanism-of-action, and pathway activation states. Ubiquitinomics uses anti-K-epsilon-GG antibodies to enrich ubiquitination sites mapping protein degradation regulation. Proteomics can empirically characterise the dynamic PTM landscape at scale impossible by candidate-gene approaches.
Clinical Proteomics
Biomarker Discovery
MS-based proteomics discovers disease biomarkers more comprehensively than antibody arrays, because it is unbiased and can discover unexpected proteins. The Human Protein Atlas project systematically measured protein expression across 44 human tissue and cell types by immunohistochemistry and transcriptomics—identifying tissue-specific proteins and cancer biomarkers. Plasma proteomics using proximity extension assay (Olink) or SomaScan (aptamer-based) quantifies thousands of circulating proteins in large epidemiological cohorts identifying disease biomarkers, drug target proteins, and protein QTLs (pQTLs) for Mendelian randomisation analysis. Proximity-based methods enable 7000+ protein quantification from microliters of plasma at population scale.
Cancer Proteomics
NCI Cancer Proteogenomics (CPTAC) consortium performed comprehensive proteomics of TCGA tumour cohorts, complementing genomic data with protein abundance and phosphorylation data. Key findings: many highly expressed cancer driver genes show protein-level amplification disconnected from DNA copy number due to compensatory mechanisms; phosphoproteomics identifies activated kinase pathways providing actionable targets independent of mutations; protein co-expression clusters reveal functional tumour subtypes orthogonal to mRNA subtypes. Comparing transcriptomics with proteomics at scale reveals post-transcriptional regulatory patterns—certain cancers downregulate specific proteins relative to their mRNA through translational regulation.
Structural and Interaction Proteomics
Cross-linking mass spectrometry (XL-MS) covalently links proximally located lysine residues in protein complexes, providing spatial constraints for structural modelling. Chemical proteomics uses reactive chemical probes to map protein-ligand binding sites (activity-based protein profiling, ABPP)—identifying all proteins in a proteome binding a specific chemotype. Drug target identification for phenotypically active compounds uses thermal proteome profiling (TPP)—measuring proteome-wide thermal stability shifts upon drug binding in intact cells. Proximity labelling (BioID, TurboID) using biotin ligase fused to a bait protein biotinylates proximal proteins in living cells, capturing transient protein interactions in situ impossible by conventional pull-down approaches.
Examples and Applications
Example 1: Human Proteome Draft Maps
Two papers published in 2014 presented draft human proteome maps (Human Proteome Project, HPP). Combined data from tissues and cell lines confidently identified proteins from over 17,000 of ~20,000 protein-coding genes—achieving 84% gene-centric coverage. Approximately 193 proteins were lacking any MS evidence, suggesting they may not be expressed or were below detection. These reference maps enable cross-study comparison, biomarker validation, and systematic gap-filling towards a complete proteome. The HPP chromosome-centric initiative ensures comprehensive coverage of every chromosome's protein products with orthogonal evidence types.
Example 2: Phosphoproteomics and Drug Mechanisms
Phosphoproteomics following imatinib (BCR-ABL inhibitor) treatment of CML cells identified hundreds of downregulated phosphosites across multiple signalling pathways downstream of BCR-ABL, confirming on-target inhibition and revealing unexpected off-target effects on other kinases. Importantly, phosphoproteomics identified compensatory pathway activation in resistant cells (Src family kinase activation providing alternative survival signals) guiding combination therapy development with dasatinib. Such global phosphoproteomic characterisation of drug-treated cells is routinely used in drug development to understand mechanism of action and resistance mechanisms systemically.
Example 3: Single-Cell Proteomics
Single-cell proteomics is emerging from SCoPE-MS (single cell ProtEomics by MS) and related methods using multiplexed tandem mass tags labelling individual cells. Current methods quantify 1000-2000 proteins per single cell—far less than single-cell RNA-seq but providing direct protein measurement. The technology is improving rapidly with nanowell sample preparation, improved LC sensitivity, and faster MS duty cycles. Single-cell proteomics can identify protein-level cell type heterogeneity invisible to transcriptomics due to post-transcriptional regulation, and reveal proteomic states of specific rare cell populations like circulating tumour cells from patient blood.
Example 4: Proteomics in Drug Discovery
Quantitative proteomics is standard in pharmaceutical drug discovery—phenotypic screen hits are characterised by thermal proteome profiling to identify all protein targets in intact cells; lead compounds are profiled against the entire proteome (selectivity proteomics) to identify off-target liabilities; mechanism of action is established by phosphoproteomics. ABPP competitive profiling with phenotypic hits identifies the specific enzymatic pocket by competition with known activity-based probes. These approaches transform drug discovery from single-target hypothesis testing to unbiased proteome-wide target identification and selectivity profiling at unprecedented scale.
Example 5: AlphaFold and Structural Proteomics
AlphaFold2 predicts accurate protein three-dimensional structures from sequence for essentially the entire proteome of any organism. The AlphaFold database contains predicted structures for virtually all UniProt proteins (~200 million). These structural predictions enable rational drug design for previously undruggable proteins by identifying cryptic ligand-binding pockets, assist interpretation of missense variants (is the affected residue in a functional site?), and enable protein-protein interaction modelling (AlphaFold-Multimer). Combining AlphaFold structures with experimental MS-based crosslinks and cryo-EM maps creates integrated structural models of protein complexes—structural proteomics at proteome scale.
Example 6: Secretome and Plasma Proteomics
The plasma proteome contains proteins secreted from virtually all tissues (secretome) making it a minimally invasive window into health and disease. Albumin and immunoglobulins dominate plasma at high abundance, but depletion strategies and enrichment allow detection of low-abundance tissue-derived proteins. Large-scale plasma proteomics in UK Biobank (>50,000 participants) quantified 2923 proteins discovering thousands of protein-disease associations and causal relationships through pQTL-based Mendelian randomisation. Proteomic ageing clocks predict biological age from plasma proteins, identifying ageing mechanisms and potential anti-ageing intervention targets at population scale.
Example 7: Glycoproteomics
Over 50% of human proteins are glycosylated—N-linked or O-linked sugars attached to asparagine or serine/threonine residues. Glycosylation profoundly affects protein folding, stability, interactions, and function. Glycoproteomics enriches glycopeptides and characterises both the protein site and glycan structure by MS. Cancer glycosylation is extensively altered: mucin-type O-glycans are truncated (Tn antigen) in many carcinomas; PSA (prostate-specific antigen) glycoforms distinguish prostate cancer from benign hypertrophy; increased sialylation promotes immunosuppression and metastasis. Cancer-specific glycoforms provide biomarker and therapeutic target opportunities exploiting the unique glycosylation patterns of tumour cells.
Example 8: Proteomics in COVID-19 Research
Proteomic analysis of COVID-19 patients and plasma from different disease severity groups rapidly identified severity predictors and mechanistic insights. Huang et al. 2021 identified plasma proteomics signatures distinguishing severe from mild COVID-19—elevated complement activation proteins, coagulation factors, and APRs in severe disease. SARS-CoV-2-infected cell proteomics revealed host protein interactions of viral proteins—NSP7-8-12 polymerase complex, ORF3a interacting with VPS39 disrupting lysosomal trafficking. Proteomics identified ACE2 expression patterns, dexamethasone targets, and antibody blocking epitopes, demonstrating proteomics' speed and breadth in pandemic emergency research.
Try it live
Everything above runs in your browser — open Proteomics: Mass Spectrometry Protein Identifier and change the parameters while it is running. Nothing is installed, nothing is uploaded, the whole model lives in one tab.
▶ Open Proteomics: Mass Spectrometry Protein Identifier simulation