Pipelines
Proteogenomic pipelines begin with high-throughput RNA sequencing (RNA-seq) data to comprehensively capture the transcriptome, identifying all expressed transcripts within a sample.
Mass spectrometry (MS) analysis is then employed to identify and quantify proteins, typically utilizing peptide mapping techniques to generate unique protein identifications based on their amino acid sequences.
Furthermore, these pipelines incorporate variant peptides – reflecting genetic variations – and novel junction identification to account for the impact of mutations on protein expression and function.
Applications
Neoantigen discovery, isoform resolution, and pathway insights.
Examples
Example: Tumor Proteogenomics Workflow
Generate RNA-seq and proteomics data.
Create custom DB; search and filter.
Prioritize neoantigens for validation.
Frequently asked questions
Database size?
The size of the database is carefully balanced to maximize coverage while maintaining stringent control over false discovery rates. Larger databases provide more comprehensive protein identification but require increased computational resources and careful FDR management during analysis.
FDR control?
False Discovery Rate (FDR) control is a critical component of proteogenomic workflows, typically achieved through target-decoy strategies where both true proteins and random sequences are searched. Stringent FDR thresholds are applied to minimize the number of spurious protein identifications.
Post-translational mods?
Search strategies for post-translational modifications (PTMs) are continuously refined with controlled expansion, incorporating known PTMs and novel modification sites identified through mass spectrometry. This allows researchers to capture the full complexity of protein expression and regulation.
Sample prep?
Consistent sample preparation protocols and rigorous quality control (QC) measures are essential for reliable proteogenomic analysis. Variations in sample preparation can significantly impact both RNA-seq and MS data, leading to inaccurate results if not carefully controlled.
Validation?
Validation of identified neoantigens or protein modifications is crucial for ensuring the accuracy of proteogenomic findings. Orthogonal assays, such as peptide sequencing, and targeted mass spectrometry experiments are commonly employed to confirm initial identifications.
Computation?
Efficient search algorithms and optimized data storage solutions are vital for managing the large datasets generated in proteogenomic workflows. Utilizing high-performance computing resources can significantly accelerate analysis times and improve computational efficiency.
Annotation?
Gene models are continuously updated with evidence derived from proteogenomic data, incorporating newly identified protein isoforms and modifications. This iterative process ensures that gene annotations accurately reflect the complex interplay between genes and proteins within a biological system.
Neoantigens?
HLA binding prediction algorithms are used to identify neoantigens – tumor-specific peptides presented on MHC molecules – which are key targets for immunotherapy. Furthermore, immunogenicity assessments are conducted to determine the potential of these neoantigens to elicit an immune response.
Standards?
FAIR data principles (Findable, Accessible, Interoperable, and Reusable) are followed in the generation and sharing of proteogenomic datasets. Detailed reporting protocols ensure transparency and reproducibility of research findings, facilitating collaboration and validation efforts.
Ethics?
Ethical considerations surrounding data collection, storage, and use are paramount in proteogenomics research. Informed consent is obtained from participants, and robust data protection measures are implemented to safeguard patient privacy and ensure responsible scientific practices.
Try it live
Everything above runs in your browser — open Proteogenomics: DNA-to-Peptide Neoantigen Explorer and change the parameters while it is running. Nothing is installed, nothing is uploaded, the whole model lives in one tab.
▶ Open Proteogenomics: DNA-to-Peptide Neoantigen Explorer simulation