Multi-omic molecular database powering researcher model selection
A single PDX model represents one patient's tumor. A PDX biobank represents an attempt to capture the genomic diversity of an entire cancer type — which requires accumulating hundreds to thousands of individually engrafted, characterized, and annotated models before it becomes a statistically useful research resource.
A single patient-derived xenograft model is scientifically informative about one tumor. But cancer of any given type — breast, lung, colorectal, pancreatic — is not one disease but a heterogeneous collection of molecular subtypes, each potentially responding differently to a given therapy. A biobank's value comes from aggregating enough individually engrafted models to statistically represent this diversity: enough KRAS-mutant models, enough microsatellite-instable models, enough of each relevant molecular subtype that a researcher can identify a cohort matching their specific hypothesis rather than relying on a single, potentially unrepresentative model.
Major collaborative biobanking efforts — the NCI's PDXNet consortium, the European EurOPDX consortium, and large commercial biobanks operated by contract research organizations — have each accumulated on the order of one to several thousand annotated models spanning dozens of cancer types, built up over many years of continuous patient sample intake.
Getting from a single successfully engrafted PDX model (as described in the companion PDX Tumor Engraftment simulation) to a catalog-ready biobank entry requires a substantial pipeline: expanding the model through several passages to build sufficient cryopreserved stock, running the full molecular annotation workflow described in Stages 2-3, verifying identity and absence of cross-contamination between models (a real risk when handling large numbers of similar samples in parallel), and compiling clinical metadata from the originating patient (with appropriate consent and de-identification).
Only after this full pipeline completes — typically six to twelve months after the original patient sample was received — is a model considered "catalog ready" and made available for search and distribution to the broader research community.
Because not every implanted patient sample successfully engrafts (see the take-rate statistics covered in the companion PDX Tumor Engraftment simulation), a biobank's catalog composition is inherently shaped by which tumor types and molecular subtypes engraft most readily — a selection bias every biobank curator must document and account for.
At scale, a biobank is organized hierarchically: by cancer type and subtype, by molecular biomarker status (e.g., HER2-positive, MSI-high, specific driver mutations), by treatment history of the originating patient (treatment-naive versus post-chemotherapy versus post-targeted-therapy), and by validated passage number and stromal purity (per the companion PDX Genomic Drift Monitoring simulation). This multi-axis organization is what makes the database queryable in the sophisticated ways described in Stage 4 — a flat, unorganized list of a thousand models would be far less useful than the same collection organized and cross-indexed by these clinically and molecularly meaningful axes.
The first and most foundational annotation layer applied to every biobanked model is its genomic profile — the somatic mutations, copy-number alterations, and structural variants that define the tumor's molecular identity and are most directly actionable for matching models to specific research questions.
Most large PDX biobanks perform whole-exome sequencing (WES) or a large targeted cancer gene panel (commonly 400-600 clinically and biologically relevant genes) on each model, typically at a validated early passage to minimize the confounding effects of passage-related genomic drift described in the companion drift monitoring simulation. Where a matched normal (non-tumor) sample from the same patient is available, it is sequenced alongside the tumor to improve confident identification of true somatic variants versus inherited germline variation — though matched normal tissue is not always obtainable, particularly for archival or externally sourced samples.
Copy-number analysis (from the same sequencing data or supplemented with a dedicated array-based approach) identifies larger chromosomal gains, losses, and amplifications, capturing structural genomic features that point mutation calling alone would miss.
Raw sequencing output is processed through a standardized bioinformatics pipeline that calls variants, filters known sequencing artifacts, and annotates each variant against reference databases (COSMIC, ClinVar, gnomAD) to flag known cancer driver mutations, distinguish likely pathogenic from benign variants, and calculate summary biomarkers researchers commonly search for directly — tumor mutational burden (TMB), microsatellite instability (MSI) status, and homologous recombination deficiency (HRD) scores among the most requested.
Standardizing this pipeline across every model in the biobank — rather than allowing pipeline variation between batches or submitting labs — is essential for making cross-model comparisons and searches statistically valid; a biobank with inconsistent variant-calling methodology between models cannot reliably support search queries across its full catalog.
A researcher searching for "KRAS G12C mutant, TP53 wild-type, TMB-low pancreatic models" is relying entirely on this standardized genomic annotation layer being accurate, complete, and consistently called across every model in the catalog — data quality here directly determines the biobank's scientific usefulness.
Every genomic annotation record is stamped with the specific passage number at which the sequencing was performed and links back to the model's full passage lineage record. This is what allows the database to flag, for example, that a model's genomic annotation was performed at F2 and remains valid through its documented passage-limit (per the companion drift monitoring simulation) — preventing a researcher from unknowingly using late-passage stock whose genomic profile may have drifted from the annotated baseline.
Genomic mutation data alone tells only part of a tumor's story. Layering transcriptomic expression profiles and digitized histopathology on top of the genomic annotation builds a far richer, multi-dimensional model profile that captures function and morphology alongside genotype.
A mutation catalog describes what could go wrong in a cell; RNA sequencing describes what the tumor is actually doing at the moment of sampling — which genes and pathways are actively transcribed, and at what relative levels. Bulk RNA-seq performed on each biobanked model (again typically at an early, drift-validated passage) enables expression-based analyses that pure genomic annotation cannot: calling established molecular subtypes for cancers where subtyping is expression-driven (such as the PAM50 intrinsic subtypes for breast cancer), identifying pathway activation signatures relevant to specific drug mechanisms of action, and flagging gene fusion events detectable at the transcript level that DNA sequencing alone might miss.
Because RNA-seq on a PDX sample captures transcripts from both human tumor and any residual mouse stromal cells, species-specific read alignment (as discussed in the companion drift monitoring simulation) is applied here as well, ensuring the reported expression profile reflects the human tumor compartment specifically.
Every biobanked model typically has hematoxylin and eosin (H&E) stained tissue sections scanned into high-resolution whole-slide images (WSI), reviewed by a pathologist to confirm tumor grade, histologic subtype, and percentage of viable versus necrotic tissue. Selected models also carry immunohistochemistry (IHC) staining for clinically relevant protein markers (e.g., ER/PR/HER2 for breast models, PD-L1 for immuno-oncology-relevant models), providing a protein-level readout that complements the genomic and transcriptomic layers.
Modern biobank databases increasingly make these digitized slides directly viewable and searchable online, sometimes incorporating quantitative digital pathology metrics (cellularity, mitotic index, computationally scored IHC staining intensity) as additional structured, queryable annotation fields rather than leaving histology as an unstructured image attachment.
A growing number of large biobanks are adding further omic layers — targeted or global proteomics, DNA methylation profiling — to select subsets of high-value models, recognizing that genomic and transcriptomic data alone still miss important post-translational and epigenetic dimensions of tumor biology relevant to certain drug mechanisms.
The real analytical power of a multi-omic biobank emerges only when these layers are integrated, not merely stored side by side. A unified per-model database record links a given model's specific mutation calls to its expression subtype, its histologic grade, and its clinical annotation, allowing cross-layer queries — for example, identifying whether a specific genomic alteration correlates with a particular expression signature or histologic feature across the catalog's full model population, a question no single omic layer could answer alone.
All of this annotation effort exists to serve one practical purpose: letting a researcher who needs a specific kind of tumor model find it quickly, using a query interface that spans genomic, transcriptomic, histologic, and clinical search criteria simultaneously.
A typical biobank query interface allows researchers to combine search criteria across every annotation layer described in Stages 2-3: specific driver mutations or mutation classes (e.g., "any RAS pathway mutation"), copy-number events, expression subtype, TMB/MSI/HRD status, histologic grade and subtype, originating patient clinical features (age, stage, prior treatment lines, response to specific therapies where documented), and model-level metadata such as validated passage limit and available quantity of cryopreserved stock.
Queries are typically combinable with Boolean logic — for example, "KRAS G12D mutant AND TP53 wild-type AND treatment-naive AND passage-validated through F5" — letting a researcher narrow a catalog of thousands of models down to a precise, scientifically relevant shortlist in minutes rather than manually reviewing published literature or contacting multiple biobanks individually.
Beyond the interactive web search interface, most major public biobanks (and many commercial ones) also expose their annotation database through an API, allowing computational researchers to query and retrieve model annotation data programmatically — essential for bioinformatics workflows that need to cross-reference biobank model characteristics against external datasets (patient cohort mutation frequencies, drug-target databases, or a researcher's own prior sequencing data) at scale, rather than one model at a time through a web form.
This programmatic access has become increasingly important as PDX biobanks are used not just to select individual models for wet-lab experiments, but as reference datasets for computational and machine-learning approaches to predicting drug response from molecular features.
Search match quality depends entirely on annotation completeness and consistency — a biobank with gaps in its RNA-seq coverage or inconsistent variant-calling pipelines between model batches will silently under-return relevant models for a given query, which is why the standardization emphasized in Stage 2 is not a bureaucratic nicety but a direct determinant of the database's practical usefulness.
Well-designed biobank query interfaces do not simply return a list of matching model IDs — they surface enough context for a researcher to judge match quality and model suitability directly from the search results: how the mutation was called (which pipeline, what confidence), how many passages the model has undergone and whether it remains within its validated limit, what quantity of cryopreserved stock is currently available for distribution, and links through to the underlying histology images and full annotation record for deeper review before committing to request a specific model.
The final step closes the loop between biobank and bench: a researcher's shortlist of best-matching models is converted into an actual request, and cryopreserved tumor stock — accompanied by its full annotation package — is shipped and re-established in the requesting lab.
Once a researcher has identified a shortlist of candidate models through the query interface described in Stage 4, formal distribution typically requires submitting a model request through the biobank's administrative process, which verifies the requesting institution, confirms appropriate research use, and executes a material transfer agreement (MTA) — a legal document governing how the model may be used, whether results must be shared back with the biobank, and any restrictions on commercial use, particularly important given that most models originate from patient tissue donated under specific informed-consent terms that constrain downstream use.
Once the MTA is executed and stock availability confirmed, the biobank retrieves the appropriate cryopreserved vials — ideally from the earliest validated passage that still meets the researcher's quantity needs, to give the receiving lab maximum room to passage the model further before approaching its documented drift limit.
Cryopreserved tumor fragments are shipped in liquid nitrogen dry shippers (maintaining ultra-low temperature without requiring external power during transit) directly to the receiving laboratory's vivarium. Upon arrival, the receiving lab thaws and re-implants the fragments following the same subcutaneous or orthotopic implantation protocols described in the companion PDX Tumor Engraftment and Orthotopic vs Subcutaneous Implantation simulations, re-establishing the living tumor line in their own animal colony.
Most biobanks recommend — and some require — that the receiving lab perform a basic identity confirmation (e.g., short tandem repeat/STR profiling, or a targeted mutation check against a few of the model's known driver mutations) after re-establishment, both to confirm no sample mix-up occurred during banking or shipping and to provide the receiving lab with fresh baseline data for their own local drift monitoring going forward.
Distribution closes a full-circle system: the annotation database that helped the researcher choose the right model also travels with the model itself, giving the receiving lab a validated genomic, transcriptomic, and histologic starting reference for every experiment they run on that line.
Many biobank programs request that researchers using distributed models report back key findings — drug response data, additional characterization performed independently, or any observed discrepancies from the provided annotation — creating a feedback loop that improves the shared database's accuracy and value over time. This community-contributed data layer, when systematically captured and integrated back into the searchable database, is increasingly recognized as one of the most valuable long-term assets a large, actively used PDX biobank can build: not just a static catalog of models, but a continuously improving, crowd-validated resource for the entire cancer research community.