Scoring and allocating in-person, home-health, and telehealth capacity across a decentralized clinical trial's investigator network, site by site
Before a single mile of catchment is modeled, hybrid site selection starts as a data-mining exercise. Sponsors and CROs pull historical site-performance records from Clinical Trial Management Systems (Veeva SiteVault, Medidata Rave CTMS), overlay them against real-world claims and EHR data from aggregators like IQVIA, Komodo Health, and TriNetX, and rank thousands of candidate investigators by therapeutic-area patient volume, past enrollment velocity, and protocol-deviation history — long before anyone asks whether a given site can support a decentralized visit model.
Feasibility mining pulls from three converging data streams:
1. CTMS historical performance: • Enrollment velocity: patients randomized per site per month, normalized against protocol complexity tier • Screen-fail ratio: proportion of screened patients who fail eligibility — high rates indicate poor pre-screening infrastructure • Protocol deviation rate: per FDA Form 1572 audit history and sponsor monitoring visit (SMV) reports • Query rate: EDC query volume per case report form (CRF) page — proxy for site data-entry discipline
2. Claims and EHR aggregators: • IQVIA National Prescription Audit and Longitudinal Access and Adjudication Data (LAAD): therapeutic-area patient counts by ZIP-3 • TriNetX federated EHR network: real-time query of ICD-10/SNOMED-coded patient cohorts across 120+ health systems without moving PHI • Komodo Health Healthcare Map: 330M+ patient claims records for prevalence-weighted site targeting • Output: an estimated addressable patient population within each candidate PI's practice or affiliated health system
3. Site technology and DCT-readiness pre-screen: • eConsent platform compatibility (DocuSign CTMS, Medidata Rave RTSM, Signant Health) • Existing telehealth infrastructure (Epic MyChart video visits, dedicated eCOA tablets) • Home-health vendor relationships already on file (Illingworth Research, Paradigm, Emerald Health) • Sites lacking any DCT infrastructure are not excluded here — they are flagged for later remote-support decisions
Scoring output: each of the 1,240 candidates receives a preliminary feasibility index (0–100) combining historical performance (50% weight) and addressable population (50% weight). Only sites clearing a 40-point floor advance to geospatial modeling — typically retaining roughly a third of the initial universe.
Geospatial modeling answers the question a spreadsheet cannot: which patients can physically reach a site, and which need the trial to come to them instead? By layering Census block-group density, CDC PLACES chronic-disease prevalence estimates, and home-health/mobile-phlebotomy vendor coverage footprints onto drive-time isochrones around each candidate site, the optimizer identifies exactly where a hybrid decentralized layer closes an access gap that a traditional brick-and-mortar-only design would leave unenrolled.
Catchment modeling proceeds in three geospatial passes:
1. Drive-time isochrone generation: • Road-network routing engine (OSRM or HERE Technologies) generates 30/60/90-minute drive-time polygons around each candidate site, not simple radius circles • Isochrones account for real road topology — a rural site's 60-minute isochrone may cover 3x the area of an urban site's due to highway access • Traffic-adjusted variants generated for peak commute windows, since many protocol visits require fasting labs drawn before 9am
2. Population and prevalence overlay: • US Census Bureau block-group population density joined spatially to each isochrone • CDC PLACES tract-level prevalence estimates for the target condition (e.g., type 2 diabetes, heart failure, moderate-to-severe plaque psoriasis) weight the raw population count toward an addressable-patient estimate • SEER cancer registry data substituted for oncology protocols; county-level substituted where tract data is suppressed for small-cell privacy rules
3. Decentralization gap identification: • Home-health nurse vendor coverage polygons (contracted radius from vendor's staffing hubs) overlaid against the site isochrone • Mobile phlebotomy and imaging-unit vendor footprints layered in for procedure-heavy protocols • Any populated area with high prevalence but outside both the site isochrone AND home-health coverage is flagged as an "access desert" — a candidate for telehealth-only enrollment or vendor network expansion • Sites whose isochrone + decentralized layer together cover under a 55% addressable-population threshold are deprioritized
Result: candidate pool narrows from 1,240 to 410 sites, each now carrying a catchment-adjusted addressable population and a median drive-time-to-site metric that becomes the baseline the modality-allocation stage tries to shrink.
A hybrid trial is not "some sites virtual, some in person" — it is a per-visit decision made against the Schedule of Assessments (SoA) itself. Each protocol visit is tagged in-person, remote-capable, or hybrid based on a visit-complexity rubric, and the allocation engine then assigns the actual delivery channel (site visit, home-health nurse, telehealth video) per site cluster to minimize patient travel burden without compromising GCP source-data verification or investigator oversight obligations under ICH E6(R2)/(R3).
Modality tagging runs against a standardized visit-complexity rubric before any optimization occurs:
Always in-person (site or clinical unit): • IP infusion/injection administration requiring on-site pharmacy dispensing under 21 CFR Part 11 chain-of-custody • 12-lead ECG with central cardiac safety review, advanced imaging (MRI, DEXA), biopsy procedures • First-dose safety visits where investigator must be physically present per protocol-specified REMS or risk mitigation plan
Hybrid-eligible (home-health or mobile unit): • Routine phlebotomy, vital signs, physical exam components deliverable by a trained research nurse • Investigational product administration for self-injectables after in-clinic training visit • Local labs processed via mobile courier to central lab (Q2 Solutions, Labcorp Drug Development) within stability windows
Remote-capable (telehealth/digital): • PI safety review visits deliverable via synchronous video (platform: Science 37, Medable, or sponsor-hosted telehealth) • ePRO/eCOA questionnaire capture, medication diary review, adverse event check-ins • Concomitant medication reconciliation, general well-being assessments
Constrained assignment engine: • Objective function: minimize Σ(patient travel time × visit frequency) subject to protocol-mandated in-person minimums • Constraint: investigator must retain documented oversight per ICH E6(R3) even for delegated remote visits — sub-investigator or qualified home-health RN must be listed on the delegation log • Constraint: informed consent and any dose-escalation safety visits remain non-negotiable in-person unless protocol explicitly designed as fully remote (rare, typically post-marketing/registry studies) • Output: a per-site, per-visit modality map feeding directly into the site's activation package and the IRT/RTSM visit scheduling logic
With catchment and modality data in hand, every surviving candidate site is scored on a weighted multi-criteria decision analysis (MCDA) spanning nine operational dimensions — from enrollment velocity to EHR interoperability to projected patient diversity reach. The resulting single composite score (0–100) is what actually drives go/no-go site activation decisions, replacing the historically ad hoc "the PI is a friend of the sponsor" method with a defensible, auditable ranking.
Composite score = Σ(dimension score × weight), nine dimensions:
1. Enrollment velocity (weight 18%) — patients/month normalized to protocol complexity, from Stage 1 CTMS mining 2. Diversity reach (weight 15%) — projected representativeness of catchment population against FDA 2022 Diversity Action Plan guidance benchmarks for the target condition 3. Staff DCT-technology readiness (weight 12%) — trained-staff count on eConsent, eCOA, and telehealth platforms; verified via site initiation visit (SIV) checklist 4. IRB/EC cycle time (weight 10%) — median days from submission to approval; central IRB (WCG, Advarra) sites typically 3–4 weeks faster than local IRB 5. EHR interoperability (weight 10%) — HL7 FHIR R4 API availability for direct EDC-to-EHR data capture, reducing transcription query rate 6. Device/kit logistics capacity (weight 10%) — cold-chain and direct-to-patient (DTP) shipment infrastructure, courier SLA compliance history 7. PI DCT experience (weight 10%) — number of prior hybrid/fully-remote protocols the investigator has run to completion 8. Budget efficiency (weight 8%) — projected cost-per-randomized-patient versus network median 9. Dropout/retention risk (weight 7%) — historical discontinuation rate, weighted by visit burden of the current protocol
Threshold and rejection logic: • Composite score <55: auto-flagged for rejection or re-scoping as a remote-only satellite feeding a nearby hub • Composite score 55–70: conditional approval pending a site-specific mitigation plan (usually additional DCT staff training or a home-health vendor contract) • Composite score >70: fast-tracked to contract and IRB submission • Score distribution is reviewed against FDA's decentralized clinical trials guidance (final guidance issued 2024) to ensure no single dimension's weighting systematically excludes historically underrepresented catchments — a bias audit run before finalizing the ranked list
A 2023 Tufts Center for the Study of Drug Development analysis found that MCDA-driven hybrid site selection reduced median site activation time by 5.4 weeks and improved 12-month retention by 9 percentage points versus geography-only site selection, largely by front-loading the DCT-readiness and IRB-cycle-time dimensions that traditional feasibility questionnaires had ignored.
Top-ranked sites are not activated in isolation — they are reconfigured into a hub-and-satellite network where a small number of high-throughput academic or health-system hubs anchor complex, equipment-heavy procedures while community satellite sites and contracted home-health vendors absorb the routine visit volume within their local catchment. A mixed-integer linear programming (MILP) solver balances projected enrollment yield against per-site activation cost and aggregate patient travel burden across the full target cohort.
The network optimization problem is formulated as a facility-location variant:
Decision variables: • x_i ∈ {0,1}: whether candidate site i is activated as a hub • y_ij ∈ {0,1}: whether satellite/patient-cluster j is assigned to hub i • z_j ∈ {0,1}: whether home-health-only coverage (no physical satellite) serves cluster j
Objective function (minimize): • w1 × Σ(activation_cost_i × x_i) + w2 × Σ(travel_time_ij × y_ij) + w3 × Σ(unmet_demand_j) • Typical weighting: 35% cost, 45% travel burden, 20% unmet-demand penalty — sponsor-adjustable per protocol priority
Constraints: • Capacity: each hub i can support at most cap_i concurrent randomized patients per month (derived from Stage 1 historical enrollment velocity) • Coverage: every patient cluster j must be assigned to exactly one hub or home-health-only pathway • Minimum network size: at least N_min hubs required for statistical/operational redundancy against single-site dropout • Regulatory: hubs must independently satisfy IRB and 1572 requirements; satellites operating under a hub's IRB via reliance agreement (per the NIH/FDA single-IRB mandate for multi-site trials)
Solver and runtime: • Solved via commercial MILP solver (Gurobi or CPLEX) or open-source CBC for smaller networks • Typical solve time: 40 sites × 400 clusters converges in under 3 minutes on standard cloud compute • Sensitivity analysis re-runs the solve across ±20% activation-cost scenarios to test network robustness before final sign-off
Output interpretation: • Academic medical center hubs (typically 8–12 of the 52 final sites) anchor imaging, IP infusion, and safety-escalation visits • Community satellite sites absorb routine follow-up and screening visits within a tighter 15–20 mile radius • Home-health-only clusters (no physical satellite) served entirely by mobile nursing and courier-based lab logistics, reserved for the lowest-density, highest-prevalence rural pockets identified in Stage 2
Site selection does not end at activation. Once first-patient-in occurs, real-time enrollment, retention, and data-query telemetry stream back from the CTMS and eCOA layer, and the same scoring engine that selected the network continuously re-ranks sites in production — deprioritizing underperformers for future patient allocation while expanding the remote-visit budget of hybrid satellites that are outperforming projections. This adaptive loop is what separates a modern optimizer from a one-time feasibility spreadsheet.
Adaptive rebalancing runs on a standing weekly data pipeline:
1. Telemetry ingestion: • CTMS enrollment counts, screen-fail reasons, and protocol deviation logs pulled weekly • eCOA/ePRO compliance rates (percentage of expected diary entries actually captured) per site and per modality channel • RTSM/IRT randomization pace compared against the original MILP-projected capacity per hub • Central lab turnaround-time and query-resolution-time feeds from the data management vendor
2. Rebalancing triggers: • Site enrollment pace <60% of projected cap_i for two consecutive months → flagged for support intervention or deprioritization • Site query rate >2 standard deviations above network mean → DCT-technology retraining triggered before deprioritization • Remote-visit compliance >95% sustained for 3 months → satellite's remote-visit budget expanded, shifting additional visit types from in-person to telehealth/home-health for that cluster • Home-health vendor no-show rate >8% → vendor reallocated or backup courier/nursing contract activated
3. Reallocation mechanics: • Freed enrollment capacity from deprioritized sites reallocated to top-quartile performers via an updated RTSM randomization ratio • New satellite activation triggered if an access-desert cluster (identified in Stage 2 but initially deferred) shows unexpectedly strong referral volume • All rebalancing actions logged and reportable to the DSMB/IRB as protocol-consistent operational adjustments, not protocol amendments, provided the SoA and eligibility criteria remain unchanged
4. Closing metrics: • Networks using adaptive rebalancing report a 10–15% higher final enrollment yield within the same trial timeline versus static site allocation, primarily by redirecting capacity before a lagging site triggers a formal corrective action plan
In a 2024 hybrid Phase III trial across a 52-site cardiometabolic network, adaptive rebalancing reallocated capacity away from 9 underperforming community sites by month 5 and expanded the remote-visit budget at 14 high-compliance satellites — closing enrollment 6 weeks ahead of the original static-allocation timeline and reducing total protocol deviations by 22% relative to the sponsor's prior non-adaptive DCT trial.