💊 DTx Engagement Retention Curve Simulator
Simulator of user retention curve for digital therapeutic applications over time.
The Week-1 Novelty Cliff in Digital Therapeutics
Digital therapeutics (DTx) — software-delivered, clinically validated interventions such as reSET for substance use disorder, EndeavorRx for pediatric ADHD, or Somryst for chronic insomnia — face an engagement problem unlike any pharmaceutical: the product must be used repeatedly to work, and most users stop using it almost immediately. Real-world engagement telemetry across dozens of published DTx deployments shows the same universal pattern: a precipitous week-1 falloff, followed by a decelerating long tail.
- 55–70%: Typical D1 retention (consumer health apps)
- 20–35%: Typical D7 retention (steepest drop window)
- 8–20%: Typical D30 retention (general mHealth median)
- +2–3×: DTx vs. consumer apps (higher retention when prescribed)
Why the first week is different from every week after it
The week-1 cliff is driven by a specific, well-documented mix of factors that behavioral analytics teams at digital health companies track obsessively:
• Onboarding friction: account creation, consent forms, baseline symptom questionnaires (PHQ-9, GAD-7), and permission requests (notifications, camera, health data) create multiple exit points before any therapeutic content is delivered • Novelty decay: initial curiosity and installation intent (often driven by a clinician referral or insurance nudge) is not the same as intrinsic motivation to sustain use • Expectation mismatch: many users expect immediate symptom relief; DTx interventions modeled on CBT, exposure therapy, or behavioral activation require weeks of consistent practice before effects appear • Cognitive/symptom burden: the target population (depression, anxiety, insomnia, substance use) frequently has the exact symptoms — low motivation, executive dysfunction, avoidance — that make sustained app engagement hardest • No clinician accountability yet: unlike week 4+ where a check-in cadence has been established, week 1 has no external accountability loop
Eli Lilly's meta-analysis of published mHealth RCTs (Torous et al., 2020, World Psychiatry) found a median 30-day retention of just 3.9% for unguided consumer mental health apps, versus 40–60% for the same intervention delivered inside a clinical trial with study-coordinator contact — showing that retention is not a fixed property of the software but of the surrounding care model.
Baumel et al. (2019, JMIR) analyzed engagement data from 93 mental health apps and found a strikingly consistent decay curve: a median of only 3.3% of users were still active 15 days after download, reinforcing that the "novelty cliff" is the modal DTx engagement pattern, not an exception.
Push Notifications and Onboarding Checklists as Early Levers
Between weeks 2 and 4, the surviving user segment is actively forming — or failing to form — a usage habit. This is the window where lightweight behavioral-design levers have outsized effect, because the user has not yet accumulated enough sessions for clinical benefit to become self-reinforcing.
- +40–90%: Push notif. open-rate lift (session return within 24h)
- 3–5/wk: Optimal push frequency (beyond this, opt-out risk rises)
- +15–25%: Onboarding checklist effect (D7 retention (Nir Eyal Hook data))
- ~30%: Notification opt-out rate (within first month if over-sent)
Behavioral design levers and their dose–response limits
Push notifications function through a simple mechanism: an external, low-cost cue that interrupts the user's attention and re-triggers app-opening behavior before the habit is self-sustaining. But the relationship between notification frequency and retention is not monotonic — it is an inverted U.
• Under-notifying: fewer than 1–2 relevant prompts per week fails to counteract natural forgetting curves; DTx apps with silent weeks show the same decay as apps with zero notifications • Optimal zone: 3–5 contextually relevant, personalized prompts per week (e.g., "your evening wind-down reminder" rather than generic "come back!" messages) lift 24-hour return rates by 40–90% in published mobile analytics benchmarks • Over-notifying: beyond ~7 prompts/week, opt-out and app-uninstall rates climb sharply — users perceive the app as nagging, which is especially counterproductive in anxiety and depression populations where notification anxiety is itself a symptom trigger
Just-in-time adaptive interventions (JITAIs) — a research framework formalized by Nahum-Shani et al. (2018, Annals of Behavioral Medicine) — improve on blanket scheduling by using contextual signals (time of day, recent app inactivity, self-reported mood) to decide not just whether to send a prompt but when it will have the highest probability of producing a beneficial response, materially outperforming fixed-schedule notifications in several DTx trials.
Power-Law Decay and the Sustained-Use Segment
By month two or three, the raw attrition rate has slowed dramatically. What remains is not a random sample of the original cohort — it is a self-selected group whose usage pattern typically follows a power-law (heavy-tailed) distribution rather than exponential decay, meaning a small fraction of highly engaged "super-users" accounts for a disproportionate share of total sessions and, in most published trials, most of the measured clinical benefit.
- ~10–15%: Sustained-use segment (of original cohort by month 3)
- ~60–70%: Share of total sessions (generated by top quartile)
- Power-law: Curve shape (not exponential — heavy tail)
- 8–12 wks: Habit-formation threshold (median for routine to stabilize)
Why DTx decay is heavy-tailed, not exponential
Naively, one might model engagement decay as exponential (constant per-capita dropout probability). Real DTx telemetry rarely fits this. Instead, dropout probability itself declines over time — the longer someone has already stuck with the program, the less likely they are to quit in the next period. This produces a power-law or stretched-exponential survival curve, a pattern also seen in habit-formation research (Lally et al., 2010, European Journal of Social Psychology, which found habit automaticity plateaus after a median of 66 days, with a range of 18–254 days depending on behavior complexity).
Practically, this means:
• Early weeks disproportionately determine long-run cohort composition — losing a user in week 2 is far more common, and far more "recoverable" with intervention, than losing one in month 4 • The surviving long-tail segment increasingly resembles a distinct sub-population with higher baseline motivation, milder symptom severity, or stronger social/clinical support — a selection effect that complicates naive interpretation of "the app works" from retention data alone • Clinical trial designs for DTx increasingly report both intention-to-treat (ITT) and per-protocol (completers-only) outcomes precisely because the sustained-use segment shows substantially larger effect sizes than the full randomized cohort
Gamification vs. Clinician Check-ins vs. Adaptive Notifications
Three retention lever families dominate DTx product design, each with different mechanisms, effect sizes, and risk profiles. No single lever is sufficient alone — the strongest-performing published DTx products layer all three with careful ethical guardrails against manipulative patterns.
- +10–20%: Gamification D30 lift (streaks, badges, progress bars)
- +25–45%: Clinician check-in D30 lift (largest single-lever effect)
- +15–30%: Adaptive (JITAI) push lift (vs. fixed-schedule push)
- +50–70%: Combined multi-lever lift (vs. no-lever baseline)
Comparative mechanisms and evidence strength
Gamification (streak counters, achievement badges, progress visualizations) exploits variable and fixed reward psychology to sustain short-term engagement, but effect sizes are typically modest (10–20% D30 lift) and can produce "streak anxiety" or a what-the-hell dropout effect once a streak breaks — an important caveat covered in depth in dedicated gamification-decay research.
Clinician check-ins — even brief asynchronous messages or a weekly 10-minute telehealth touch-point — consistently show the largest single-lever effect in published DTx literature, often lifting 30-day retention by 25–45 percentage points relative to unguided self-use. The mechanism is accountability: a human expecting to review your data changes the psychology of a missed session from "no consequence" to "something to explain." This is the core rationale behind blended-care and coach-supported DTx models (e.g., Lantern, Meru Health), which report substantially higher completion rates than app-only comparators in head-to-head studies.
Adaptive, context-aware notifications (JITAIs) outperform fixed-schedule push by using real-time signals — geolocation, time-of-day usage history, wearable-derived stress markers — to time prompts for moments of highest receptivity, avoiding both under- and over-notification failure modes described in Stage 2.
The strongest-performing published DTx combine all three, but layering must be done carefully: stacking gamification, notifications, and human check-ins simultaneously without user control over intensity risks recreating the dark-pattern dynamics seen in addictive consumer apps — a tension DTx regulators (FDA Digital Health Center of Excellence) increasingly scrutinize.
The Retention–Outcome Correlation in Published DTx Trials
The central premise justifying DTx retention engineering is clinical, not commercial: unlike a pill that works whether or not the patient thinks about it between doses, a DTx intervention's active ingredient is repeated engagement itself. This makes cumulative engaged-session count a legitimate dose metric, and several pivotal trials have now formally quantified the dose–response relationship.
- r≈0.4–0.5: Pear Therapeutics reSET-O (sessions vs. abstinence days)
- 6 sessions: Somryst (insomnia CBT-I) (minimum for durable ISI improvement)
- 25 days: EndeavorRx (pediatric ADHD) (required per FDA label, 5×/wk)
- ~0.6: Meta-analytic dose–response (effect-size correlation, engagement tertiles)
Engagement as the active ingredient — evidence and regulatory implications
FDA-cleared DTx products increasingly specify a minimum effective dose in their labeling, directly analogous to a drug's dosing schedule. EndeavorRx (akili Interactive), the first FDA-cleared video-game DTx for pediatric ADHD, was cleared based on a 25-consecutive-weekday, 25-minutes-per-day protocol — engagement below this threshold was not evaluated for efficacy and is not covered by the label claim. Somryst (Pear Therapeutics/Nox Health), a CBT-I DTx for chronic insomnia, requires completion of at least 6 of 9 core sessions for its labeled Insomnia Severity Index improvement to apply.
Across published DTx RCTs, secondary dose–response analyses (comparing outcome improvement across tertiles or quartiles of realized engagement) consistently find a positive, often near-linear relationship between cumulative sessions completed and symptom improvement — though this correlational evidence is confounded by the same selection effect noted in Stage 3 (more motivated or less-severe patients both engage more AND improve more, independent of any causal engagement effect).
This dual nature — engagement as both a plausible mechanism and a confounded proxy — is why leading DTx developers now run engagement-optimization A/B tests as pre-registered secondary trial endpoints, and why payers and the FDA increasingly request real-world engagement telemetry as part of post-market surveillance, treating retention curves as an early-warning signal for real-world effectiveness that may diverge from tightly-monitored trial conditions.
A 2021 systematic review of 34 FDA-cleared and CE-marked DTx products (Wang et al., npj Digital Medicine) found that only 40% reported real-world engagement data post-launch, and of those, realized engagement was on average 30–50% lower than trial-phase engagement — underscoring that retention engineering is not a growth-hacking afterthought but a clinical necessity for DTx to deliver its trial-demonstrated benefit in the real world.
Simulator of user retention curve for digital therapeutic applications over time.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install