🧫 Autoantibody Signature Autoimmune Diagnosis
This simulation provides a panel of autoantibodies for early diagnosis of autoimmune diseases.
Building the Discovery Cohort — Serum Biobanking for Autoantibody Profiling
Every autoantibody signature study begins not with an array but with a carefully phenotyped cohort. Case definition, preanalytical handling, and control matching determine whether downstream reactivity differences reflect true disease biology or confounding artifacts of collection, storage, and demography. For undifferentiated autoimmune syndromes — where patients present with overlapping symptoms years before a specific diagnosis (SLE, Sjögren's, RA, mixed connective tissue disease) crystallizes clinically — serum biobanked at first presentation is the only material that can support genuinely early diagnostic discovery.
- 412: Enrolled participants (206 cases / 206 controls)
- 3: Recruiting centers (tertiary rheumatology clinics)
- 7.4 mo: Median symptom duration (at serum draw, cases)
- <2 h: Serum processing window (venipuncture to −80°C)
Case definition, matching, and preanalytical standardization
Cohort design for autoantibody discovery must balance clinical realism against statistical power:
Case ascertainment: • Inclusion: new-onset polyarthralgia, photosensitive rash, sicca symptoms, or Raynaud's phenomenon without a confirmed connective-tissue-disease diagnosis at enrollment • Adjudication: 2 independent rheumatologists blind to array data assign a reference diagnosis 24 months later (ACR/EULAR classification criteria) • This "future diagnosis" reference standard lets the array signature be tested as a true early-detection tool rather than a diagnosis-confirmation tool
Control matching: • 1:1 age (±3 yr) and sex matching; no personal or first-degree family history of autoimmune disease • Excludes recent infection (<4 weeks), pregnancy, and active malignancy — all known to perturb systemic IgG reactivity
Preanalytical variables (the largest source of irreproducibility in serum proteomics): • Tube type: serum separator tubes (BD SST) processed identically across all 3 sites • Time-to-spin: >4h at room temperature measurably degrades complement-sensitive epitopes and elevates background reactivity — protocol mandates centrifugation within 2h • Freeze–thaw: each cycle measurably increases nonspecific IgG aggregation signal on protein arrays; aliquots are single-use, capped at 3 cycles, and freeze-thaw count is logged per vial • Hemolysis index and lipemia are scored visually and by spectrophotometry (414nm); samples with free hemoglobin >0.5 g/L are excluded from the discovery set
Biobank logistics: • 200μL serum aliquots in barcoded cryovials, 2D-barcode racking for automated retrieval • Storage at −80°C in redundant, alarm-monitored freezers; chain-of-custody logged at every thaw event
HuProt Proteome Microarrays — Screening Nearly the Entire Human Proteome for Autoreactivity
The core experimental innovation enabling autoantibody signature discovery is the full-proteome protein microarray: thousands of individually purified, full-length recombinant human proteins printed at defined coordinates on a single slide, allowing one drop of patient serum to be screened against nearly the entire human proteome in a single assay. Where classical autoimmune serology tests one autoantigen at a time (anti-dsDNA ELISA, anti-CCP, ANA by IFA), the array format converts autoantibody discovery from a hypothesis-driven, one-at-a-time search into an unbiased, proteome-wide scan.
- 9,483: Proteins on HuProt v4.0 (~81% of the coding proteome)
- 2: Print replicate spots (per protein, per subarray)
- 1:200: Serum dilution (90 min RT incubation)
- 5 μm: Scanner resolution (GenePix 4400A, dual PMT gain)
Array manufacture and the serum incubation workflow
Full-length human protein microarrays are manufactured by expressing thousands of ORFs individually and purifying each protein before printing:
Protein production: • ORF clones (human ORFeome collaboration / CCSB collection) expressed as GST- or His-tagged fusions in yeast (S. cerevisiae) using a high-throughput induction pipeline • Each protein affinity-purified individually (glutathione or Ni-NTA resin), quality-checked by anti-tag Western blot for full-length expression • ~9,483 proteins printed in duplicate spots on nitrocellulose-coated glass slides using a contact microarray printer (~180μm spot diameter, 400μm pitch)
Hybridization protocol: 1. Slide blocking: 5% BSA/PBS-T, 1h, RT — saturates nonspecific protein-binding sites on nitrocellulose 2. Serum probing: patient serum diluted 1:200 in blocking buffer, applied to subarray chamber, 90 min RT with gentle agitation 3. Washing: 3× PBS-T, 5 min each — removes unbound serum proteins 4. Detection: Alexa Fluor 647-conjugated goat anti-human IgG (Fc-specific), 1:1000, 45 min in the dark 5. Final wash and spin-dry; slides scanned within 2h to minimize fluorophore photobleaching
Scanning and image capture: • GenePix 4400A confocal laser scanner, 635nm excitation channel for AF647 • Dual PMT gain acquisition (500 and 600) captures both weak and saturating spots within one dynamic range • 5μm pixel resolution resolves individual 180μm spots as ~35×35 pixel regions for robust median-intensity extraction
Alternative platforms in clinical translation: • Peptide microarrays (JPT PepStar, NimbleGen): overlapping 15–20mer peptides map linear epitopes at single-residue resolution — critical for conformational vs. linear epitope discrimination • Bead-based suspension arrays (Luminex xMAP): lower-plex (up to ~500) but higher throughput and CLIA-friendly, used for panel-locked clinical assays after discovery (see Stage 5)
From Raw Fluorescence to Reactivity Calls — Normalization and Statistical Filtering
A raw protein-array scan is dominated by technical noise: uneven hybridization, spatial gradients across the slide, batch-to-batch scanner drift, and background autofluorescence of the nitrocellulose substrate. Converting 9,483 raw spot intensities per sample into a small number of statistically defensible "reactive" calls requires a rigorous normalization and multiple-testing pipeline — the same statistical discipline used in microarray gene-expression analysis, adapted to the antibody-reactivity setting.
- 9,483: Antigens tested per sample (duplicate spots averaged)
- 6: Hybridization batches (cyclic-loess corrected)
- Z > 3: Reactivity Z threshold (vs. control distribution)
- 612: FDR-significant hits (Benjamini–Hochberg q<0.05)
Quantile normalization, batch correction, and the reactivity call
Signal extraction and normalization proceed in defined stages:
1. Local background subtraction: • GenePix Pro computes foreground median and a local background median (annular region around each spot) per feature • Net signal = FG_median − BG_median; negative values floored at a small positive constant to allow log-transform • Duplicate spots averaged (CV<15% required, else feature flagged as unreliable)
2. Between-array (between-sample) normalization: • Quantile normalization (limma normalizeBetweenArrays): forces the full intensity distribution of every sample to match a common reference distribution, removing sample-to-sample scaling differences unrelated to true biology • Cyclic loess applied afterward to remove residual nonlinear intensity-dependent bias between pairs of arrays
3. Batch effect correction: • 412 samples processed across 6 hybridization batches (different days, secondary-antibody lots) • ComBat empirical Bayes batch correction removes batch mean/variance shifts while preserving case/control biological variance • Batch as a factor is confirmed non-significant in a post-correction PCA/PERMANOVA check
4. Reactivity Z-score and thresholding: • Per-antigen robust Z-score = (sample signal − median of control-serum distribution) / MAD(control distribution) × 1.4826 • Z>3 (≈99.7th percentile of a normal reference) called reactive for that sample-antigen pair • Case-vs-control reactivity frequency compared via Fisher's exact test per antigen across the full 9,483-antigen panel
5. Multiple-testing correction: • Benjamini–Hochberg FDR applied across all 9,483 simultaneous tests; q<0.05 retained • 612 antigens survive FDR filtering — a >15-fold reduction from the full array, discarding antigens whose apparent case-enrichment is consistent with chance given 9,483 comparisons • These 612 form the candidate pool entering feature selection (Stage 4)
From 612 Candidates to a 24-Marker Panel — Regularized Feature Selection
A diagnostic test cannot use 612 antigens — clinical assays need a small, fixed, manufacturable panel. The feature-selection stage compresses statistically significant but highly correlated reactivity data (many autoantibodies co-occur because they target the same apoptotic-cell-derived antigen complexes, or arise from linked B-cell clonal expansions) into a minimal marker set that preserves classification performance. This is the step that converts a discovery-stage observation into a candidate clinical assay.
- 612: Candidate antigens entering (FDR<5% from Stage 3)
- 24: Final panel size (autoantibodies)
- 0.7: Elastic-net mixing (α) (glmnet, 10-fold CV)
- 0.91: Discovery-cohort AUC (95% CI 0.87–0.94, nested 10×10 CV)
Regularized regression, ensemble feature selection, and cross-validation
Panel construction combines complementary feature-selection methods and validates honestly with nested cross-validation:
1. LASSO / elastic-net logistic regression (glmnet): • Objective: minimize −log-likelihood(case vs. control) + λ[(1−α)‖β‖₂²/2 + α‖β‖₁], α=0.7 • The L1 component drives most of the 612 antigen coefficients exactly to zero; the retained nonzero-coefficient antigens form the LASSO-selected set • λ chosen by 10-fold cross-validation at the "λ.1se" rule (largest λ within 1 SE of minimum CV error) — favors a smaller, more stable panel over the absolute-minimum-error model
2. Random forest variable importance (ntree=2000): • Mean decrease in Gini impurity ranks each antigen's contribution to case/control separation across an ensemble of 2000 bootstrap-resampled decision trees • Captures nonlinear and interaction effects that a linear LASSO model can miss (e.g., an antigen only informative in combination with a second)
3. SVM recursive feature elimination (SVM-RFE): • Linear-kernel SVM trained on all 612 features; features ranked by |weight|, lowest-ranked 10% removed, retrained iteratively • Converges on a minimal feature subset that preserves margin-based separability
4. Consensus panel selection: • An antigen is retained in the final panel only if selected by ≥2 of the 3 methods — reduces overfitting to any single algorithm's idiosyncrasies • Intersection yields 24 autoantibodies spanning known lupus/Sjögren's-associated antigens (Ro52/TRIM21, Ro60, La/SSB, Sm-D1, U1-snRNP-70K) alongside novel candidates (centromere-associated protein E fragments, histone H2B variants, RNA-binding motif proteins) not previously reported as diagnostic markers
5. Performance estimation without optimism bias: • Nested 10×10 cross-validation: an outer loop holds out test folds never seen during any part of feature selection or hyperparameter tuning in the corresponding inner loop • This nested design is essential — selecting features on the full dataset before a single non-nested CV step inflates AUC estimates, a well-documented pitfall in biomarker panel papers • Honest nested-CV estimate: AUC 0.91 (95% CI 0.87–0.94)
Locking the Panel — Independent Validation and Serological Subtyping
A biomarker panel is only a candidate until it survives contact with data it never influenced. The final stage freezes the 24-antibody model exactly as trained, transfers it to a clinically deployable multiplex format, and tests it against an independent cohort recruited after the model was locked. This stage also asks a second question beyond binary diagnosis: does the reactivity pattern itself stratify patients into biologically and clinically distinct subtypes?
- 167: Independent validation cohort (2 external sites)
- 0.94: Validation AUC (sensitivity 89% / specificity 92%)
- Luminex xMAP: Assay format for deployment (24-plex, CLIA-compatible)
- 3: Serological subtypes resolved (consensus clustering, k=3)
Assay transfer, lock-down validation, and unsupervised subtype discovery
Translating a discovery-array signature into a clinical test requires both a platform change and a strict validation discipline:
Platform transfer: • The 24 selected antigens are individually conjugated to spectrally distinct magnetic Luminex xMAP microbeads, replacing the planar protein array • Bead-based multiplexing is CLIA-compatible, requires only 12.5μL serum, and returns quantitative median fluorescence intensity (MFI) per antigen in a single well — a format compatible with clinical laboratory throughput (hundreds of samples/day) • Cross-platform concordance (array MFI vs. bead MFI) is confirmed by Deming regression on a 40-sample bridging set (r>0.85 required per antigen before proceeding)
Lock-down validation protocol: • The logistic model — 24 coefficients plus intercept, fixed from Stage 4 — is frozen before any validation-cohort data is examined (a pre-registered analysis plan) • 167 new patients from 2 sites not involved in discovery are tested; laboratory personnel running the Luminex assay are blinded to clinical diagnosis • Locked model performance: AUC 0.94, sensitivity 89%, specificity 92% at the pre-specified probability cutoff (0.42) chosen in the discovery cohort — not re-optimized on validation data
Unsupervised serological subtyping: • Consensus clustering (ConsensusClusterPlus, k=2–6 tested, k=3 selected by cluster-consensus CDF stability) applied to the 24-marker reactivity matrix across all validated cases • Subtype A: dominant anti-Ro52/Ro60/La reactivity, enriched for HLA-DRB1*03:01, associated with sicca symptoms and photosensitivity • Subtype B: dominant anti-Sm/RNP reactivity, enriched for HLA-DRB1*15:01, associated with higher rates of nephritis at 24-month follow-up • Subtype C: low-titer polyreactive pattern across many panel members, intermediate HLA association, slower progression to classifiable disease • These subtypes were not defined by clinical criteria — they emerge purely from serological reactivity patterns, yet they stratify 24-month treatment response and organ-involvement risk, suggesting the autoantibody signature captures pathophysiologically distinct disease programs rather than a single continuous severity axis.
Because the 24-marker signature was frozen before validation-cohort data were examined, and the validation cohort was recruited at sites independent of discovery, AUC=0.94 is a genuinely unbiased estimate of real-world diagnostic performance — not an artifact of feature selection leaking into performance estimation, which is the single most common flaw in published biomarker-panel studies. Combined with the emergent 3-subtype structure, this positions the panel as both an early-diagnosis tool and a stratification tool for subtype-tailored treatment trials.
This simulation provides a panel of autoantibodies for early diagnosis of autoimmune diseases.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install