Approach
Custom databases from RNA-seq
Mass spectrometry identification
Variant peptides and splice forms
Example
Example: Cancer Variant Peptides
Construct sample database.
Search MS data for variants.
Validate and interpret.
Frequently asked questions
FDR?
False Discovery Rate (FDR) control methods are utilized to account for potential false positives when identifying variant peptides. These controls, often implemented with target-decoy strategies, ensure that the reported variants have a statistically significant level of confidence.
Coverage?
Considerations regarding coverage involve balancing depth and breadth; deeper sequencing provides more comprehensive coverage of the transcriptome but can be computationally intensive. The optimal balance depends on the research question, with broader coverage potentially capturing rarer isoforms or variants.
PTMs?
Post-translational modifications (PTMs) significantly impact protein function and are routinely investigated through enrichment strategies during mass spectrometry analysis. Researchers employ targeted searches to identify specific PTMs, such as phosphorylation or glycosylation, which can be associated with genetic variants.
Databases?
Sample-specific databases are crucial for accurate variant identification and interpretation. These databases are built by incorporating genomic information alongside RNA-seq data, ensuring that the mass spectrometry analysis focuses on relevant peptide sequences specific to the analyzed sample.
Quant?
Quantitative proteogenomic analyses can be performed using label-free quantification methods or through the identification of tags associated with specific variants. These approaches allow for relative comparisons of protein abundance and expression levels, providing insights into disease mechanisms.
Validation?
Independent validation is essential to confirm the findings generated by proteogenomic analysis. This often involves employing orthogonal methods such as antibody-based techniques or alternative mass spectrometry approaches to corroborate the identified variant peptides.
Clinical?
Proteogenomics holds significant potential for clinical applications, particularly in biomarker discovery and development. The technology can be integrated into biomarker pipelines for disease diagnosis, prognosis, and treatment monitoring based on protein expression profiles.
Compute?
The computational aspects of proteogenomics involve complex workflows and substantial storage requirements for genomic data, mass spectrometry results, and database resources. Efficient algorithms and scalable infrastructure are necessary to manage the large datasets generated during analysis.
Reproducibility?
Ensuring reproducibility in proteogenomic studies requires adherence to established standards for sample preparation, data acquisition, and analysis. Sharing of raw data, standardized protocols, and robust quality control measures are crucial components of a reproducible workflow.
Outlook?
The future of proteogenomics lies in the routine integration of joint genomic and proteomic analyses for comprehensive disease characterization. This approach promises to revolutionize our understanding of complex biological systems and accelerate the development of targeted therapies.
Try it live
Everything above runs in your browser — open Proteogenomics Explorer: Genome-to-Spectrum Matching and change the parameters while it is running. Nothing is installed, nothing is uploaded, the whole model lives in one tab.
▶ Open Proteogenomics Explorer: Genome-to-Spectrum Matching simulation