Sequence Analysis Foundations
Bioinformatics: the application of computational tools to understand biological data — primarily DNA, RNA, and protein sequences. DNA sequencing revolution: Sanger sequencing (1977) → Human Genome Project ($3B, 13 years) → Illumina NGS ($600, 2 days) → Nanopore sequencing (portable, real-time, long reads >4 Mbp). Sequence alignment: comparing sequences to find similarity — evidence of evolutionary, structural, or functional relationships. BLAST (Basic Local Alignment Search Tool): the most-used bioinformatics tool — compares a query sequence against databases in seconds. Smith-Waterman: optimal local alignment (dynamic programming, O(mn) complexity). Multiple Sequence Alignment (MSA): ClustalW, MUSCLE, MAFFT, T-Coffee — aligning 3+ sequences reveals conserved regions. Scoring matrices: BLOSUM62, PAM250 — quantify amino acid substitution probabilities. Hidden Markov Models (HMMs): HMMER, Pfam database — profile-based search for protein families and domains.
Genomics and Databases
GenBank (NCBI): >2.5 trillion bases in >250 million sequences — the primary nucleotide database. UniProt: comprehensive protein database — reviewed (Swiss-Prot, 570K entries) and unreviewed (TrEMBL, 250M+ entries). PDB (Protein Data Bank): >200,000 experimentally determined 3D structures. Ensembl / UCSC Genome Browser: genome annotation and visualization platforms. Gene Ontology (GO): standardized vocabulary for gene/protein function across species — molecular function, biological process, cellular component. Metagenomics: sequencing environmental samples without isolating organisms — studying entire microbial communities. Single-cell genomics: sequencing individual cells reveals cellular heterogeneity — 10x Genomics Chromium, Drop-seq. Spatial transcriptomics: gene expression mapped to tissue location — Visium (10x), MERFISH, Slide-seq. Multi-omics integration: combining genomics, transcriptomics, proteomics, metabolomics for systems-level understanding.
Phylogenetics and Evolution
Phylogenetics: reconstructing evolutionary relationships from sequence data — building "tree of life." Methods: Maximum Parsimony (fewest evolutionary changes), Maximum Likelihood (most probable tree given model), Bayesian Inference (posterior probability of trees). Bootstrapping: statistical support for tree branches — 1000+ replicates standard practice. Molecular clock: mutations accumulate at roughly constant rate → calibrate divergence times. Horizontal Gene Transfer (HGT): gene sharing between unrelated species — complicates tree-of-life for prokaryotes. Phylogenomics: using whole genomes for phylogenetic analysis — resolves difficult relationships. SARS-CoV-2 genomic epidemiology: real-time phylogenetics tracked viral evolution, variant emergence, transmission chains. Nextstrain: open-source platform for pathogen genomic surveillance — used worldwide during COVID-19. Ancient DNA: sequencing Neanderthal, Denisovan genomes revealed interbreeding with modern humans (Svante Pääbo, Nobel 2022).
Structural and AI Bioinformatics
Structural bioinformatics: predicting and analyzing 3D structures of biomolecules. Homology modeling: building protein structure from known template — SWISS-MODEL, Modeller. AlphaFold 2 (DeepMind): revolutionized structure prediction — atomic-level accuracy from sequence alone. AlphaFold DB: 200M+ predicted structures freely available. Molecular docking: predicting how small molecules bind to protein targets — AutoDock, Glide, GOLD. Molecular dynamics (MD): simulating atomic movements over time — GROMACS, AMBER, NAMD. Drug discovery pipeline: target identification → virtual screening → lead optimization → ADMET prediction → clinical candidate. AI in bioinformatics: deep learning for variant effect prediction (DeepVariant, EVE), gene expression imputation, drug-target interaction prediction. Large language models for biology: ESM (Meta), ProGen, ProtTrans — protein language models trained on billions of sequences. Single-cell AI: scVI, scBERT, CellTypist — automated cell type annotation and trajectory inference. Future: digital twins of cells and organisms — complete computational models for personalized medicine.
❓ Frequently Asked Questions
Bioinformatics: the application of computational tools to understand biological data — primarily DNA, RNA, and protein sequences. DNA sequencing revolution: Sanger sequencing (1977) → Human Genome Pro...
GenBank (NCBI): >2.5 trillion bases in >250 million sequences — the primary nucleotide database. UniProt: comprehensive protein database — reviewed (Swiss-Prot, 570K entries) and unreviewed (TrEMBL, 2...
Phylogenetics: reconstructing evolutionary relationships from sequence data — building "tree of life." Methods: Maximum Parsimony (fewest evolutionary changes), Maximum Likelihood (most probable tree ...
Structural bioinformatics: predicting and analyzing 3D structures of biomolecules. Homology modeling: building protein structure from known template — SWISS-MODEL, Modeller. AlphaFold 2 (DeepMind): re...
Try it live
Everything above runs in your browser — open Sequence Alignment: DNA Matching in 3D and change the parameters while it is running. Nothing is installed, nothing is uploaded, the whole model lives in one tab.
▶ Open Sequence Alignment: DNA Matching in 3D simulation