🗣 Behavioral Nudge Vaccination Uptake Experiment
This simulation tests the effectiveness of different behavioral nudges in increasing vaccination uptake. Users can design and compare various interventions to see their impact on public behavior towards getting vaccinated.
Study Design and Randomization — The Foundation of Causal Nudge Evidence
To know whether a text message actually causes more people to get vaccinated — rather than merely correlating with people who were already inclined to — researchers rely on randomized controlled trials (RCTs). The largest of these, the "Behavior Change for Good" megastudy led by Katy Milkman, Angela Duckworth, and collaborators, tested dozens of behavioral nudges simultaneously against a shared control group, published as "A Megastudy of Text-Based Nudges Encouraging Patients to Get Vaccinated" (PNAS, 2021).
- RCT: Megastudy design (randomized, pre-registered)
- ~47,306: Penn Medicine trial size (patients, 19 message arms)
- ~689,693: Walmart Pharmacy trial size (patients, 22 message arms)
- Milkman, Duckworth: Lead investigators (Behavior Change for Good Initiative)
Why randomization, and why a "megastudy"
A randomized controlled trial solves the fundamental identification problem in behavioral science: people who opt into a wellness text-message program probably already care more about their health, so simply comparing "signed up" vs. "did not sign up" people would hopelessly confound the message's effect with pre-existing motivation. Randomization severs this link — by flipping a coin (or, in practice, a computer-generated random assignment) for every eligible patient, the control and treatment arms are, in expectation, identical on every characteristic that might affect vaccination, observed or unobserved.
The "megastudy" innovation, pioneered by the Behavior Change for Good Initiative (a partnership between the University of Pennsylvania and Behavioral scientists including Katy Milkman, Angela Duckworth, and Katherine Milkman's collaborators), goes further: instead of one team testing one intervention against one control group, dozens of independently-designed nudges — submitted by different behavioral science teams — are tested simultaneously against a single shared control group in the same population and time window. This design lets researchers directly rank interventions by effect size using identical statistical footing, something no single-intervention RCT can do.
The two field experiments
The published megastudy combined two parallel field experiments run during the 2020 flu season:
1. Penn Medicine primary care trial: approximately 47,306 patients across primary care practices, randomized into a control group and 18 distinct nudge message arms (differing in wording, framing, and timing), with vaccination status tracked via electronic health records
2. Walmart Pharmacy trial: approximately 689,693 pharmacy customers, randomized into a control group and 22 distinct SMS nudge arms, with the outcome measured as an actual flu shot administered at a Walmart pharmacy within a fixed follow-up window
Combining a health-system population (patients with an existing care relationship) and a retail-pharmacy population (a much larger, more transactional customer base) let the researchers test whether nudge effectiveness generalized across very different contexts and baseline uptake rates — a critical external-validity check that single-site studies cannot provide.
Control Arm — The Organic Baseline Uptake Funnel
Before any nudge can be evaluated, the control arm establishes what happens without intervention: standard-of-care communication only. In the funnel visualization, control-arm particles flow from "eligible" through "reminded," "scheduled," and finally "vaccinated" at whatever rate ordinary friction and motivation produce.
- ~30–35%: Typical baseline uptake (flu season, primary care setting)
- largest: Eligible→reminded drop-off (many patients never see standard reminder)
- moderate: Reminded→scheduled drop-off (intention-action gap)
- smallest: Scheduled→vaccinated drop-off (once scheduled, most follow through)
The funnel structure and where people fall out
Vaccination uptake is not a single decision but a chain of smaller steps, each with its own drop-off, a structure directly analogous to a marketing or product conversion funnel:
1. Eligible: a patient is due for a flu shot per clinical guidelines 2. Reminded: the patient actually receives and notices a reminder (many standard mailed or portal reminders go unseen) 3. Scheduled: the patient takes the concrete step of booking or deciding on a specific appointment/visit 4. Vaccinated: the patient completes the shot
In the control arm, receiving only routine, low-salience reminders (a mailed postcard, a buried patient-portal message), the largest drop-off typically occurs between "eligible" and "reminded" — most patients simply never engage with routine outreach — and a second substantial drop-off occurs between "reminded" and "scheduled," reflecting the classic behavioral-economics "intention-action gap": people intend to get vaccinated but never convert that intention into a concrete calendar commitment.
Why a shared, well-powered control group matters
In the actual megastudy, the shared control group was deliberately made large relative to any single nudge arm specifically to give the baseline uptake estimate itself a tight confidence interval — since every treatment-arm comparison inherits the control estimate's uncertainty. A noisy control estimate would make it impossible to distinguish a genuinely effective nudge from statistical noise, regardless of how large the treatment effect might be.
Across both the Penn Medicine and Walmart samples, baseline (control) flu vaccination uptake fell in a broadly similar range — illustrating that even without any experimental message, a substantial share of eligible patients get vaccinated through routine channels, existing intention, or standing appointments; the nudges compete against this already-nontrivial baseline, which is part of why realistic nudge effect sizes are usually measured in single-digit percentage points rather than a doubling or tripling of uptake.
Applying the Nudge — Four Behavioral Mechanisms Tested in the Funnel
Once the nudge is applied, the funnel visibly widens: more particles survive each gate. The specific mechanism matters — different nudges intervene at different points in the funnel, and the megastudy literature shows some mechanisms consistently outperform others.
- ~4+ pts: Default appointment lift (Chapman et al. 2010; opt-out framing)
- ~4.2 pts: Commitment device lift (Milkman et al. 2011, PNAS)
- ~1–3 pts: SMS reminder lift (megastudy text-message arms)
- modest: Social norm messaging lift (weaker than active-choice nudges)
SMS reminder and social-norm messaging
SMS reminders work primarily by widening the "eligible → reminded" gate: a text message has far higher open and attention rates than mail or portal notifications, so simply making sure the reminder is actually seen recovers some of the funnel's largest natural drop-off. In the megastudy's Walmart and Penn Medicine arms, plain reminder texts produced modest but reliably positive lifts, and — notably — sending a second reminder outperformed a single reminder, suggesting attention/salience decays and needs reinforcement.
Social-norm messaging draws on Robert Cialdini's foundational work on descriptive norms (Cialdini, "Crafting Normative Messages to Protect the Environment," 2003, and related public-health applications): telling people "most people like you already got their flu shot" leverages the human tendency to conform to perceived typical behavior. In the megastudy, purely norm-based messages tended to produce smaller effects than messages that additionally prompted a concrete action step — suggesting that shifting belief about what others do is necessary but not sufficient; the message also needs to close the intention-action gap.
Default appointments and commitment devices — the two strongest mechanisms
The two most consistently powerful nudge mechanisms operate on the "scheduled" gate directly, rather than on attention or belief:
Default/opt-out scheduling: rather than asking patients to actively book an appointment, the system pre-assigns (defaults) a specific date and time, and the patient must actively act to opt out or reschedule rather than to opt in. This flips the default from "nothing happens unless you act" to "vaccination happens unless you act" — exploiting status-quo bias in the pro-social direction. Chapman et al. (2010, "Opt-Out Influenza Vaccination Reminders," Health Affairs / related default-scheduling studies) documented substantial uptake gains from this approach; the "active choice" variant (Keller, Harlam, Loewenstein & Volpp, "Enhanced Active Choice," Health Affairs 2011) — forcing an explicit yes/no decision rather than a silent default — also reliably outperforms passive reminders, though typically by a smaller margin than a true default appointment.
Commitment devices: Milkman, Beshears, Choi, Laibson & Madrian ("Using Implementation Intentions Prompts to Enhance Influenza Vaccination Rates," PNAS, 2011) found that simply asking employees to write down the specific date and time they planned to get vaccinated increased vaccination rates by roughly 4.2 percentage points relative to a plain reminder — one of the largest documented low-cost nudge effects in the vaccination literature. The mechanism is "implementation intentions" (Gollwitzer, 1999): converting a vague goal ("I should get vaccinated sometime") into a specific if-then plan closes the intention-action gap directly.
Effect Size Computation — Absolute and Relative Lift
With both funnels complete, the core causal quantity of interest is the difference in final uptake between arms: the absolute lift (in percentage points) and the relative lift (as a percentage increase over baseline) — the two standard ways an RCT result is reported and compared across studies.
- p_nudge − p_ctrl: Absolute lift formula (in percentage points)
- (p_n−p_c)/p_c: Relative lift formula (as a percentage)
- ~4+ pts abs.: Best megastudy arms (default/active-choice & commitment)
- small, positive: Median megastudy arm (most of 22 arms beat control)
Reading absolute vs. relative lift correctly
Absolute lift — the simple difference between nudge-arm and control-arm uptake rates — is the number that matters for public-health impact: a 4-percentage-point absolute lift on a population of one million eligible patients means roughly 40,000 additional vaccinations, regardless of what the baseline rate happened to be.
Relative lift — the absolute lift divided by the control rate — is useful for comparing an intervention's "multiplicative" effect across settings with very different baselines, but can be misleading in isolation: a nudge that lifts uptake from 2% to 4% has a 100% relative lift but only a 2-point absolute lift, while a nudge lifting uptake from 32% to 36% has a "only" 12.5% relative lift but a much larger 4-point absolute lift and far greater real-world impact at scale. The megastudy analysis reports both, but public-health decision-makers should weight absolute lift most heavily when the goal is total additional people vaccinated.
What the megastudy found across 19–22 competing arms
Across both the Penn Medicine and Walmart trials, the headline finding was that the large majority of tested nudges produced a positive — if often small — effect on vaccination, but effect sizes varied severalfold across mechanisms. The best-performing message variants used "reserved for you" ownership/endowment-style framing (e.g., language implying a dose had specifically been set aside for the recipient), echoing the psychological finding that people work harder to avoid losing something already framed as theirs than to gain an equivalent new benefit (loss aversion, Kahneman & Tversky).
A second consistent finding: sending patients two reminder texts rather than one meaningfully outperformed a single message — suggesting that for text-based nudges specifically, repetition/reinforcement of a low-cost message adds real incremental value, a design detail cheap enough for any health system to adopt regardless of which specific wording performs best.
Statistical Significance — Confidence Intervals, p-Values, and Sample Size
A positive lift in a sample is not proof of a real effect until its statistical significance is established. The two-proportion z-test formalizes how confident we can be that the observed gap between arms reflects a true underlying difference rather than random sampling noise — and sample size is the single biggest lever over that confidence.
- 2-proportion z-test: Test used (or chi-square equivalent)
- p < 0.05: Conventional threshold (95% confidence standard)
- ∝ 1/√n: CI width vs. sample size (diminishing returns per added n)
- shared control: Megastudy advantage (boosts power for every arm)
The two-proportion z-test and why sample size is decisive
To test whether the nudge-arm uptake rate p_n differs from the control-arm rate p_c, the standard approach is a two-proportion z-test:
z = (p_n − p_c) / sqrt( p̂(1−p̂)(1/n_n + 1/n_c) )
where p̂ is the pooled uptake rate across both arms and n_n, n_c are the two arm sample sizes. The resulting z-score maps to a p-value: the probability of observing a lift at least this large purely by chance if the true effect were zero.
The denominator makes explicit why sample size matters so much: standard error shrinks proportionally to 1/√n, so quadrupling the sample size only halves the confidence interval width — a diminishing-returns relationship that is why megastudies deliberately recruit hundreds of thousands of participants rather than a few hundred: with a true effect size of only a few percentage points (typical for realistic nudges), a small study is simply underpowered to distinguish that real effect from noise, and would report a misleadingly wide confidence interval or a false-negative "no significant effect" result.
Why the shared-control megastudy design is statistically efficient
Because every nudge arm in a megastudy is compared against the same large, shared control group rather than each having its own dedicated, smaller control, statistical power for every single comparison is higher than an equivalent set of 19–22 fully independent two-arm trials would achieve with the same total sample size. This is precisely why the Behavior Change for Good megastudy design could confidently rank interventions that differed by only 1–3 percentage points in absolute lift — differences that would likely be statistically indistinguishable from zero in a conventional, single-intervention trial with a few thousand participants per arm.
In the interactive simulator here, moving the "Sample Size" slider visibly demonstrates this relationship: at small sample sizes, the same true underlying lift produces a wide, overlapping pair of outcome distributions and a large p-value (not significant); at the megastudy's actual scale (tens to hundreds of thousands per arm), even a modest true lift produces two clearly separated distributions and a vanishingly small p-value.
Scale-Up Projection — From a Trial Effect to a National Impact Estimate
The final step of a nudge RCT is translating a validated percentage-point lift, measured in a trial of tens of thousands of people, into a projection of impact if the winning intervention were deployed across an entire health system or country — the step that turns an academic finding into a public-health policy decision.
- ~250M+: US annual flu-eligible pop. (CDC-recommended annual flu shot)
- millions: Projected extra vaccinations (at scale, best-performing nudge)
- ~cents: Marginal cost per SMS nudge (per patient reached)
- context drift: Deployment risk factor (effect may shrink outside trial setting)
The arithmetic of scale-up, and its caveats
Scale-up projection is conceptually simple arithmetic: multiply the trial's estimated absolute lift by the size of the target deployment population. A nudge shown to lift uptake by roughly 2 percentage points, deployed across a health system serving 10 million eligible patients, projects to roughly 200,000 additional vaccinations per season if the effect holds at scale — an impact far beyond what any single RCT population could demonstrate directly, and the central practical argument for running large, cheap, easily-scalable nudges (a text message costs a small fraction of a cent) rather than expensive, high-touch interventions with similar effect sizes.
But naive scale-up projection carries real risks the megastudy authors explicitly caution against:
• Effect decay outside the trial context: a novel message loses its salience once patients (or an entire population) have seen similar messages repeatedly across seasons — the "reserved for you" framing that worked as a novel treatment in a one-time trial may lose potency if reused as standard annual messaging • Population non-representativeness: the trial population (Penn Medicine patients, Walmart pharmacy customers) may differ systematically from the full target deployment population in ways that change the true effect size • Implementation fidelity: real-world deployment (message timing, technical delivery reliability, opt-out rates) rarely matches a tightly-controlled trial's execution exactly
Why the megastudy approach remains a policy-relevant benchmark
Despite these caveats, the Behavior Change for Good megastudy design was explicitly built to generate policy-actionable evidence: by testing dozens of realistic, low-cost, easily-scalable message variants within actual health-system and retail-pharmacy delivery infrastructure (rather than a lab setting), the results transferred directly into real deployment decisions. Following the study, several of the tested message strategies — particularly two-reminder sequences and ownership-framed language — were adopted into routine patient outreach by participating health systems and pharmacy chains for subsequent flu seasons.
The broader lesson for public-health policy is that a rigorously measured few-percentage-point lift, deployed essentially for free across a population of hundreds of millions, can outperform far more expensive interventions on a cost-per-additional-vaccination basis — which is precisely why behavioral-economics nudges have become a standard tool alongside (not instead of) traditional public-health communication and access interventions.
This simulation tests the effectiveness of different behavioral nudges in increasing vaccination uptake. Users can design and compare various interventions to see their impact on public behavior towards getting vaccinated.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install