Blog · Gamification Analysis Contact Me
Test, Learn, Adapt: S-Tier Behavioral Designer’s Guide
Gamification Analysis

Test, Learn, Adapt: S-Tier Behavioral Designer’s Guide

Every public-sector team I have ever advised believes it tests its policies. Almost none of them do. They pilot. They consult. They evaluate. They run before-and-after comparisons that compare apples in March to oranges in November and call the resulting tea leaves “evidence.”

The UK Cabinet Office Behavioural Insights Team called this bluff in 2012, in a 32-page paper that sits behind every credible behavioural-insights deployment since. Laura Haynes, Owain Service, Ben Goldacre, and David Torgerson titled it Test, Learn, Adapt: Developing Public Policy with Randomised Controlled Trials, with a foreword by David Halpern. The argument was unfussy. If you are going to call a policy “evidence-based,” the evidence has to actually exist, and almost the only way to produce it inside a real government is the randomised controlled trial.

That sounded radical to civil servants in 2012. A decade later, it is the spine of behavioural-insights practice on three continents. The nine steps reorganise how a policy team thinks about its own ignorance. The Test, Learn, Adapt loop refuses the conceit that intuition plus committee plus communication strategy equals knowledge.

The post below is the S-Tier Behavioral Designer’s Guide to that protocol — what Haynes and her co-authors actually said, where it sharpens every other behavior model you already know, where it falls apart, and how to wire it into an Octalysis-grade build so the trial earns its political budget instead of burning it.

Speed Run Notes

  • Most “evidence-based policy” is not. Haynes, Service, Goldacre & Torgerson’s 2012 BIT paper named the bluff and replaced it with a nine-step protocol built around the randomised controlled trial.
  • Nine steps, three verbs. Test (steps 1-7) generates causal evidence. Learn (step 8) interrogates what it means. Adapt (step 9) routes the result back into policy and the next test.
  • Randomisation does political work, not just statistical work. The unit you randomise on decides whose lobby gets to claim credit and whose excuse survives. Pick the unit honestly or the trial reads as theatre.
  • Powering matters more than novelty. Underpowered trials hand undeserved confidence to whichever side was loudest, then the policy scales on noise. Pre-register, or do not run.
  • Publication of null results is a feature, not a courtesy. A protocol that only reports wins is a marketing pipeline. The BIT rule is publish what you measured, or do not measure it.
  • Octalysis Crosswalk: each Test/Learn/Adapt step has a Core Drive that applied teams reliably skip. Honest hypothesis = CD3. Honest power = CD8. Honest reporting = CD1. Design at the Core Drive level.

Author Credibility: Yu-kai Chou

Yu-kai Chou — creator of the Octalysis Framework

Yu-kai Chou created the Octalysis Framework after studying gamification since 2003 — years before the term entered mainstream vocabulary. As a Human-Systems Architect & Behavioral Designer, his framework has been applied by LEGO, Microsoft, Porsche, Coca-Cola, Salesforce, and MrBeast, impacting over 1.5 Billion Users.

Chou has taught the Octalysis methodology at Harvard, Stanford, Yale, Tesla, Google, BCG, and IDEO.

His work has been cited by Harvard, Stanford, MIT, Forbes, Wall Street Journal, Wired, US Department of Energy, NIST, NSF, NCBI, US Department of Education, ClinicalTrials.gov, and Google Scholar — with 3,700+ more academic publications. Explore his books here.

What Is Test, Learn, Adapt?

Test, Learn, Adapt is the 2012 paper that made randomised controlled trials a routine instrument of UK public policy. Laura Haynes was a senior advisor at the Behavioural Insights Team. Owain Service was the team’s first managing director. Ben Goldacre brought the medical-RCT tradition and the rhetorical force of Bad Science. David Torgerson directed the York Trials Unit. The four wrote one document together and shipped it under the Cabinet Office imprint with David Halpern’s foreword. The argument was direct enough to fit on a postcard, and the implementation tight enough to fit inside a Whitehall delivery cycle.

The argument has two halves. First, almost every claim that a public-sector intervention “works” rests on weaker evidence than the public assumes. Pre-post comparisons mistake regression to the mean for effect. Pilots without comparison groups confuse selection effects with policy. Focus groups confuse articulation with behavior. None of these methods can answer the only question that matters: did the intervention cause the outcome, or would the outcome have happened anyway?

Second, the only method that reliably answers that question inside the constraints of government is the randomised controlled trial. The paper walks through the nine practical steps of mounting one. It explains how to choose between individual, cluster, and stepped-wedge randomisation. It addresses the political objections (it is unfair to randomise) and the practical ones (the agency cannot afford to wait for results). It points at twelve early BIT trials as proof of concept, on subjects ranging from court-fine collection to organ donation.

The framework’s design genius is the verb sequence in the title. Test is operational: run the trial. Learn is epistemological: read the result against the prior. Adapt is political: route the result back into policy. Most policy shops do one or two of these. The BIT pattern is to do all three, in order, as a routine rather than a project, until the loop is faster than the policy cycle that depends on it.

The Nine Steps of Test, Learn, Adapt

The paper organises practice into nine numbered steps. The first seven sit inside Test. Step 8 is Learn. Step 9 is Adapt. The numbering looks bureaucratic and it is. The bureaucracy is load-bearing because each step is also a place where a normal policy process quietly skips a decision and then pretends it never had to be made. Reading the steps as a checklist of skipped decisions is the right way in.

Step 1: Identify Two or More Policy Interventions to Compare

The most-skipped step in the entire protocol. Departments routinely set up trials of “the intervention” against “the status quo,” then discover that “the status quo” is a moving target because three other reforms went live in the same quarter. The BIT discipline is to identify, before any randomisation, the exact contrast you are testing. Treatment A vs Treatment B. Treatment vs business-as-usual specified at the level of a letter template, a script, a default setting. If you cannot describe the contrast in one sentence, the trial will not be interpretable when it lands.

A second, harder version of Step 1 is to identify a contrast that the policy maker is willing to act on. A trial that compares two interventions neither of which the minister can endorse if they win is a trial nobody learns from. Test, Learn, Adapt insists that the test only counts if Adapt can run on its results.

Step 2: Determine the Outcome the Policy Is Intended to Influence

Specify the outcome operationally. “Tax compliance” is not an outcome; “the proportion of self-assessment taxpayers who pay the full amount owed within 30 days of receiving the reminder letter” is. The BIT paper is unsparing on this. The most common failure is measuring what is easy (calls received, leaflets distributed) rather than what matters (behavior change downstream of the leaflet). Pick the downstream outcome before the trial design, not after the data lands.

A pre-registered primary outcome also keeps everyone honest at Step 8. With one declared primary outcome and a small set of pre-specified secondaries, the result is what it is. Without them, an honest team can still talk itself into a finding by running enough sub-group analyses, and a dishonest one will.

Step 3: Decide on the Randomisation Unit

Individual, cluster, or stepped wedge. Each carries a different price. Individual randomisation gives the smallest sample size for a given power but only works when treatments do not bleed across people in the same setting. Cluster randomisation (whole schools, whole job centres, whole hospitals) handles bleed-over but costs in sample size and demands cluster-level analysis. Stepped wedge rolls the intervention out across units in randomised order, so every unit eventually gets the intervention — the politically defensible option when withholding feels unfair.

The choice is not technical, it is political. A cluster trial says “the intervention reaches everyone in this office and nobody in that one.” A stepped wedge says “everyone gets it, in a randomised sequence.” A team that does not declare which it is doing in Step 3 will get caught between the two when the policy lead asks why some constituents got the new letter and others did not.

Step 4: Determine How Many People or Units Are Required

Sample-size calculation against the minimum detectable effect that would justify rolling the intervention out. This is where most public-sector trials secretly die. A team picks a number that fits the budget, runs the trial, and either (a) gets a result and over-interprets it, or (b) gets a null and cannot tell whether the intervention does nothing or whether the trial was too small to see what it does. Either failure mode discredits the protocol.

The BIT discipline is to power the trial for the effect size that would change policy if true. If you would not act on a 1% lift, do not run a trial that can only see a 1% lift. The honest version of this conversation usually shrinks the menu of trial-able interventions and that is healthy.

Step 5: Randomly Assign People or Units to Treatment and Control

Operational step, but the protocol matters. Generate the allocation off-site, in a way the operational staff cannot see or game. Document the seed, the time, and the random-number routine. The single most damaging failure of a behavioural-insights trial is staff “helping” by re-allocating subjects they think will respond better — not from malice, from belief that they know what works. Allocation concealment, borrowed from the medical-RCT tradition, prevents this with procedure rather than vigilance.

Step 6: Introduce the Interventions to the Two Groups

Fidelity. The intervention has to be delivered as specified, on schedule, to the right people. Fidelity monitoring is the unsexy half of trial delivery and the reason most academic-public-sector partnerships break. Build a fidelity check into the protocol: a monthly sample of delivered letters, a random audit of calls, a comparison of intended and delivered intervention dose. When fidelity slips, document it. A trial of “the intervention as we intended to deliver it” is a different question from a trial of “the intervention as it was actually delivered.” Both are legitimate. Conflating them is not.

Step 7: Measure the Results and Determine the Impact of the Interventions

Pre-registered analysis plan, executed. The analysis plan was written before the data came in. The plan names the test, the population, the handling of missing data, the multiple-comparison correction, the sub-groups (if any) that are pre-specified, and the headline metric that goes to the policy lead. Execute the plan as written. Note any deviations. Report both.

The BIT discipline at Step 7 is to write the press release for the null result before opening the data. If you cannot live with the null version, your prior was strong enough that you should not have run the trial.

Step 8: Adapt Your Intervention to Reflect Your Findings

Step 8 is the Learn step in the verb sequence, despite its placement in the numbering. The team converts the trial result into an updated belief about what the intervention does, for whom, under what conditions. The honest version of this includes uncertainty intervals, an explicit comparison to the prior expectation, and a list of contexts in which the result would not generalise. The dishonest version reports the headline effect as a fixed truth and quietly drops the confidence intervals.

Updating is asymmetric. A positive result with a tight interval updates belief sharply. A null result with a wide interval updates belief weakly. A positive result with a wide interval (under-powered trial, lucky draw) should update belief weakly too, and almost never does in practice. Adapt requires that the team treat the result the way the data warrant, not the way the politics demand.

Step 9: Return to Step 1

The loop. The fundamental insight of Test, Learn, Adapt is that public policy is not a project, it is a process, and the unit of value is the rate at which the policy machine can compound learning. A single trial is interesting. Twenty trials in series, each refining the next, is governance.

The hardest political work inside Step 9 is to protect the loop from a single-trial mindset. The minister wants the answer. The press wants the result. The protocol wants the next test. The BIT pattern, copied successfully by units in Australia, the US, Singapore, and the Netherlands, is to embed the loop inside a delivery unit that owns its own backlog and reports against the rate of trials completed rather than the absolute size of any individual effect.

What Haynes, Service, Goldacre & Torgerson Got Right

Three observations make the paper survive a decade of follow-on critique.

The institutional fix is the protocol, not the trial. A one-off RCT inside a department is easy to dismiss as the special project of one team. A documented protocol that any team can run, on any policy lever, on a routine schedule, is harder to dismiss because the asymmetry has flipped. The question is no longer “why did you run a trial here” but “why did you not run one.” The paper put that asymmetry in writing and gave it a name.

The political objections were addressed up front. A long section of the paper walks through every standard objection (it is unfair, it is expensive, it is slow, we already know it works) and answers each one with the specific design move that addresses it. Stepped wedge for fairness. Pre-existing administrative data for cost. Iterative rolling for speed. A small pilot trial for prior calibration when intuition is strong. The paper anticipates the conversation a minister will have with a permanent secretary the day after reading the executive summary.

The exemplar list earned the rest of the argument. Twelve early BIT trials, with effect sizes and pound figures attached, gave the protocol something concrete to point at. The HMRC tax-letter trial alone recovered an order of magnitude more revenue than the entire BIT operating budget for the year. The court-fine SMS trial saved tens of millions in bailiff costs. The protocol earned its credibility by going first, in public, with the numbers.

Where Test, Learn, Adapt Falls Apart

The protocol is durable, not flawless. Three failure modes show up reliably in real deployments.

Critique 1: Internal Validity Is Not the Same Thing as Useful Knowledge

Nancy Cartwright and Angus Deaton have argued, separately and together, that RCTs answer a narrow question well and a wider question badly. The narrow question is whether the intervention worked in this trial, on this population, in this context. The wider question, which policy makers actually want answered, is whether the intervention will work in the next context. RCTs do not by themselves answer the wider question. They produce a treatment effect averaged across a specific population, in a specific moment, with a specific delivery mechanism. Scaling that effect to a different population, moment, or mechanism is an act of theory, not of replication.

Test, Learn, Adapt under-weights this gap. The Adapt step assumes generalisation is a deductive step from the trial result; in practice it is an inductive jump that requires explicit theory about why the effect occurred in the first place. The fix is to pair every trial with a mechanism statement — “this letter worked because it activated descriptive norms” — that constrains how the result generalises. A protocol that names only effects, never mechanisms, will mis-scale most of what it learns.

Critique 2: The Set of Trial-able Interventions Is Narrower Than Policy Itself

RCTs require a manipulable treatment, a measurable outcome, and a population that can be randomly assigned. That excludes most of macro policy. Monetary policy, large structural reforms, constitutional change, geopolitical posture — none of these admit an RCT design. The BIT-style behavioural intervention sits inside a narrow band of policy: small, well-specified treatments, on individual or local-cluster populations, with administrative-data outcomes. Inside that band the protocol shines. Outside it, the protocol is silent, and the temptation is to scale “what we tested” past “what testing can decide.”

This produces a structural bias in evidence-based policy: the trial-able interventions accumulate evidence faster than the non-trial-able ones, even when the non-trial-able ones matter more. A decade of BIT-style trials has produced richer evidence on letter wording than on workforce reform. That is the discipline doing its job, not failing — but the policy mix must not confuse the available evidence base for the actual choice set.

Critique 3: The Protocol Tolerates Statistically Significant Tiny Effects as if They Were Policy-Worthy

A well-powered trial detects a 1% lift on a large administrative sample. Statistical significance is achieved. The headline reports “evidence-based intervention raises X by 1%.” The accompanying article does not ask the question that matters: is the 1% lift worth the cost of the intervention, the opportunity cost of running the trial, and the political capital of declaring victory on a small effect?

The original paper does not solve this. Steps 4 and 8 hint at it (power for the effect size that would justify rollout; report uncertainty intervals) but the field has drifted toward reporting any positive p-value as a win. The corrective is to declare, in Step 1, the minimum effect size that would change policy if true and to refuse to publish a victory headline on an effect below that floor. The honest BIT-aligned units do this. Most behavioural-insights press releases do not.

What’s Really Happening Inside the Org

Test, Learn, Adapt is a protocol but its hardest work is psychological. Inside any real public-sector team, the protocol asks people to do three things their normal incentives push them away from.

It asks them to be uncertain in public. A senior official who pre-registers a hypothesis is staking professional credibility on a result they cannot control. The same official, asked to write a policy memo, can hedge and equivocate and survive any outcome. Pre-registration removes the hedge, which is its point and its institutional cost.

It asks them to share credit with chance. A policy that worked because of the intervention is a victory. A policy that worked because the comparison group also improved (regression, seasonality, policy contamination) is a non-finding. The trial protocol forces the team to give up the part of the credit that random variation would otherwise quietly award them.

It asks them to publish what they would rather forget. The Test, Learn, Adapt commitment to publish all results, positive, null, and negative, is the protocol’s most subversive feature. Most policy machinery prefers to publish wins, suppress losses, and re-frame everything else. A protocol that requires the team to publish the loss is a protocol that punishes the team for running the trial in the first place — unless the institution actively rewards null publications. The BIT solved this with a culture norm and a public-facing trial registry. Units that have tried to import the protocol without the registry have routinely watched it decay back into selective publication within two years.

The cognitive load of the protocol is therefore not the statistics. The statistics are off-the-shelf. The cognitive load is the discipline of accepting public uncertainty, sharing credit with chance, and publishing losses. A team that solves these three loops once will run the protocol for a decade. A team that has not solved them will run one trial and quietly stop.

Test, Learn, Adapt vs Other Frameworks

Test, Learn, Adapt vs MINDSPACE

The two BIT canonical documents do different jobs. MINDSPACE (Dolan et al, 2010) is a catalogue of nine intervention levers (Messenger, Incentives, Norms, Defaults, Salience, Priming, Affect, Commitments, Ego). Test, Learn, Adapt is the protocol for deciding which of those levers actually moves the dial in your context. MINDSPACE without Test, Learn, Adapt is a menu without a test kitchen. Test, Learn, Adapt without MINDSPACE is a test kitchen with no chef. The two are designed to chain: MINDSPACE generates hypotheses, Test, Learn, Adapt evaluates them.

Test, Learn, Adapt vs EAST

EAST (Easy, Attractive, Social, Timely) is the post-2014 BIT operational simplification of MINDSPACE for front-line teams. It compresses the nine levers into a four-letter heuristic for non-specialists. Test, Learn, Adapt sits above EAST in the same way it sits above MINDSPACE. EAST helps you draft the intervention; Test, Learn, Adapt decides whether the intervention worked.

Test, Learn, Adapt vs COM-B

COM-B (Michie, van Stralen, West 2011) is a diagnostic framework: it asks which of Capability, Opportunity, or Motivation is the binding constraint on the behavior. Test, Learn, Adapt is a verification framework: it tests whether the chosen intervention moved the outcome. COM-B feeds Test, Learn, Adapt with a diagnosis-grounded hypothesis. The chain is COM-B (diagnose) → MINDSPACE / EAST (design) → Test, Learn, Adapt (verify).

Test, Learn, Adapt vs the Octalysis Framework

Octalysis names the motivational structure under the behavior; Test, Learn, Adapt names the empirical structure under the intervention. The two are complements. An Octalysis-grade design decides which Core Drive to bend; Test, Learn, Adapt decides whether the bend actually happened. Designers fluent in only one of these typically over-claim. Octalysis without Test, Learn, Adapt over-claims mechanism. Test, Learn, Adapt without Octalysis under-claims theory. The Crosswalk later in this post is the bridge.

Test, Learn, Adapt in the Real World

HMRC Tax-Letter Trials

The flagship demonstration. Between 2010 and 2014 the BIT and HMRC ran a series of randomised trials on the standard tax-payment-reminder letter. The most-cited version (Hallsworth, List, Metcalfe and Vlaev, Journal of Public Economics 148, 2017) tested social-norm framings against a business-as-usual control. The descriptive-norm variant (“9 out of 10 people in the UK pay their tax on time”) lifted the 23-day payment rate by 5.1 percentage points on a sample of more than 100,000 letters. Extrapolated to the relevant population, the intervention recovered hundreds of millions of pounds in accelerated payment.

What made the trial canonical was not the effect size. It was the protocol. Pre-registered hypothesis. Pre-specified population. Pre-specified primary outcome. Analysis plan written before the data came in. Result published in a peer-reviewed economics journal, not a glossy departmental brochure. That is the full Test, Learn, Adapt loop in one trial.

Court-Fine Text-Message Trial

The 2012 BIT court-fines SMS trial sent personalised text messages to people who had been issued court fines and not paid within ten days. The message named the recipient and the amount owed. The control group received the standard letter only. Payment within thirteen weeks rose from 33% to 53% in the named-SMS arm. Net savings on bailiff costs, projected across the full fines population, ran into tens of millions of pounds per year. The trial used cluster randomisation across court regions to handle delivery-system constraints.

Jobcentre Plus Job-Search Intervention

Reported by Owain Service and colleagues in BIT’s 2014 EAST report. The trial restructured the initial Jobcentre Plus claimant interview around an immediate-action set-up (commitment device, calendared job-search activities, follow-up scheduling) versus the standard advisory format. Off-benefit rates at 13 weeks improved by several percentage points in the treatment arm. The trial mattered less for the effect size than for proving that a Test, Learn, Adapt loop could operate inside the actual front-line caseload of a major operational department, not just in a research-friendly back office.

Loft Insulation Uptake

The BIT and the (then) UK Department of Energy and Climate Change ran an RCT on loft-insulation uptake. The treatment arm received a loft-clearance service bundled with the insulation offer. The control arm received the insulation offer alone. The bundled offer roughly tripled uptake. The lesson the team published was not “people respond to bundles” (everyone knew that) but “the binding constraint was clutter in the loft, not the cost of insulation.” Diagnosis (COM-B Opportunity), design (bundled offer), verification (RCT). Test, Learn, Adapt let the team distinguish a real binding constraint from a plausible-sounding one.

Organ-Donor Registration

The BIT and the UK Driver and Vehicle Licensing Agency ran one of the protocol’s most-cited trials on the organ-donor sign-up page that accompanies driving-licence renewal. Eight message variants were randomly assigned across roughly one million renewals: a control, a basic prompt, a reciprocity framing (“if you needed an organ transplant would you have one?”), a loss-framed variant, a logo-bearing variant, and three social-proof framings of varying specificity. The strongest variant (“Every day thousands of people who see this page decide to register”) lifted sign-ups by an estimated 96,000 additional registrations per year against the control. The weakest social-proof variant (an image of a group of people accompanying the same text) actually underperformed the plain-text social-proof variant, against the team’s prior expectation.

The trial mattered for three reasons beyond the headline number. First, it ran in production traffic on a high-volume page, demonstrating that the protocol scales to administrative channels not normally treated as research surfaces. Second, the “obvious” prediction was wrong: the imagery variant was supposed to win and lost. Third, the team published the underperforming variants alongside the winner, modelling the publish-everything discipline the protocol depends on.

GP Antibiotic-Prescribing Letters

Hallsworth, Chadborn, Sallis and colleagues reported in The Lancet 387 (10029), 2016, a cluster-randomised trial of letters sent to the top 20% highest-prescribing general practitioners in England. The intervention letter was signed by England’s Chief Medical Officer and told the recipient practice “the great majority of practices in [your area] prescribe fewer antibiotics per head than yours.” Antibiotic prescribing in the treated practices fell by 3.3% over the following six months, equivalent to roughly 73,000 fewer prescriptions in absolute terms, against no change in the control arm. The trial cost approximately £4,000 in postage and produced an estimated £92,000 in NHS savings, plus the public-health benefit of reduced antibiotic resistance. The Test, Learn, Adapt loop closed twelve months later when the letter was rolled out as standard practice, with the trial design serving as the protocol for evaluating subsequent variants.

The Elephant in the Room

Test, Learn, Adapt is a morally neutral protocol. The same nine steps that prove a beneficial pension nudge can prove an exploitative subscription dark pattern. The same allocation-concealment discipline that protects an HMRC tax letter from staff gaming protects a gambling operator from staff who would intervene to slow a problem customer down. The protocol is not on anyone’s side. It improves the rate at which the team learns what works. Whether what works is good for the population is a question the protocol does not answer.

This is the recurring elephant in the Behavioral Framework Library, and the reason every V6 pillar imports the Libertarian Paternalism publicity test as its ethics check. Before the trial runs, the designer should be able to write the press release for both result halves — the trial worked, and the trial worked too well — and be willing to publish either under their own name to the affected population. The trial worked: did the population benefit, or did the funder? The trial worked too well: did the additional efficiency improve lives, or did it accelerate a transfer from the population to the funder?

An RCT-validated intervention that fails the publicity test is more dangerous than an unvalidated one, because the evidence base around it lends moral cover to the rollout. “We tested it” becomes the answer to ethical objections it does not actually address. The protocol’s contribution to the field is enormous and the protocol’s responsibility for the use of its outputs is zero, which is the load-bearing contradiction the discipline has not yet honestly resolved.

How to Apply Test, Learn, Adapt with the Octalysis Framework

The Octalysis Framework, my own contribution to behavioral design, identifies eight Core Drives that motivate behavior. The Test, Learn, Adapt protocol describes how to test interventions empirically. The two are complements. Below is the Crosswalk: which Core Drive sits under each step of the protocol, and which one applied teams most reliably skip.

Octalysis Framework with Game Techniques around each Core Drive — Yu-kai Chou

Step 1-2 (Define and Outcome): Core Drive 3 (CD3): Empowerment of Creativity & Feedback

The hypothesis-honesty step. Identifying the right contrast and the right outcome is a creative act — the designer must articulate, before any data exists, the specific behavior change worth testing and the specific signal that would confirm it. Applied teams skip this when they let the funder write the hypothesis. The CD3 fix is to ring-fence the analyst’s right to specify the trial in a way that can produce a null. A trial that cannot produce a null is a marketing exercise, not a test.

Step 3 (Randomisation Unit): Core Drive 5 (CD5): Social Influence & Relatedness

The cluster question is a social-architecture question. Treatments leak across people who talk to each other, share a manager, attend the same school, work the same shift. Choosing the randomisation unit honestly requires modelling the social structure of the population the intervention will reach. Individual randomisation in a tightly-networked population produces contamination that masks the effect. Cluster randomisation respects the network at the cost of statistical power. CD5 is the lens that names the trade-off correctly.

Step 4 (Power): Core Drive 8 (CD8): Loss & Avoidance

The power calculation is loss avoidance in pure form. An underpowered trial loses twice: once because the data cannot answer the question, again because the political capital spent on running it is gone whether the trial succeeded or not. CD8 in trial design is the discipline of declaring, in advance, the loss the trial is allowed to absorb. A trial powered for an effect smaller than the smallest acceptable real-world effect is a structurally guaranteed loss. The CD8 move is to refuse to run.

Step 5-6 (Allocate and Deliver): Core Drive 6 (CD6): Scarcity & Impatience

The operational window. Random allocation and intervention delivery happen in a calendar window before the broader policy rollout, and that window is scarce. Operational staff under time pressure are the single largest threat to allocation concealment and fidelity. CD6 in trial delivery is the discipline of treating the window as the precious resource it is, and protecting it institutionally with monitoring, audit, and an escalation path that the operational lead cannot quietly override.

Step 7-8 (Measure and Update): Core Drive 1 (CD1): Epic Meaning & Calling

The truth-commitment step. Pre-registering the analysis plan, executing it as written, and updating belief according to the data rather than the politics is a CD1 move, not a CD3 one. The designer is committing to a purpose larger than the immediate result: the integrity of the institution’s evidence base, the next trial that depends on the credibility of this one, the field’s capacity to learn at all. Without CD1 fuel, the temptation to read the data optimistically is overwhelming, because every other Core Drive points the other way.

Step 9 (Adapt): Core Drive 4 (CD4): Ownership & Possession

The result has to be routed back into someone’s portfolio for the loop to close. The team that ran the trial has to take possession of the finding, communicate it inside the organisation, advocate for the rollout (or the kill), and write the next hypothesis. CD4 in policy is the named, accountable owner of the evidence stream. Without a CD4 owner the trial result sits in a published report and the policy machine carries on as if it had never run.

Throughout (Protocol Maturation): Core Drive 2 (CD2): Development & Accomplishment

The team itself is the meta-subject. Each trial is a unit of mastery. A team on its first trial moves slowly. A team on its twentieth trial moves quickly. The Test, Learn, Adapt loop is a CD2 progression for the team as a craft community. Departments that have run hundreds of trials (the BIT itself, the US Office of Evaluation Sciences, the New South Wales BI Unit) display the difference: trial preparation that would have taken six months in 2012 now takes six weeks.

Throughout (Genuine Experimentation): Core Drive 7 (CD7): Unpredictability & Curiosity

The protocol depends on the team being willing to be surprised. CD7 is the Core Drive of genuine experimentation: running a trial whose outcome you cannot confidently predict, and treating the result as information. A team that always predicts the result correctly is running too few trials of the right kind. A team that is genuinely surprised once in three trials is calibrated. CD7 in trial-design culture is the discipline of seeking out the trials whose outcomes are most uncertain and least obvious, because those are the trials that produce the most learning per pound spent.

Practical Steps: The Test, Learn, Adapt × Octalysis Audit

A six-step audit you can run on any proposed behavioural-insights intervention before it goes to trial. Each step is the operational form of the Crosswalk above.

  1. Elicit before designing. Run eight to twelve structured user conversations on the binding constraint, not the proposed intervention. Most failed BIT-style trials are failures at this step dressed as failures elsewhere. The constraint you find is rarely the constraint the funder named.
  2. Pre-register a hypothesis with a mandatory null prediction. Write the press release for the null outcome before opening the field. If you cannot live with the null version, your prior is strong enough that the trial is decorative.
  3. Choose the randomisation unit honestly. Model the social architecture of the population. Individual randomisation if no contamination; cluster if any; stepped wedge if withholding is politically untenable. Declare the choice in advance and defend it before the data.
  4. Power for the effect size that would change policy. If you would not act on a 1% lift, do not run a trial that can only see a 1% lift. The honest version of this calculation usually shrinks the menu of trial-able interventions, which is healthy.
  5. Pre-commit to the reporting protocol. Null published, headline restored if exaggerated, uncertainty intervals reported at every step. The commitment must be institutional, not personal. A culture that punishes null publication will not publish nulls regardless of what any individual wants.
  6. Apply the publicity test before scaling. Could the team publish the full intervention design (population, lever pulled, funder, intended behavior change, beneficiaries) under their own name to the affected population? If not, the RCT evidence does not authorise the rollout.

Test, Learn, Adapt Was the Beginning, Not the End

The 2012 paper opened the protocol. The intervening decade has filled in the gaps. Pre-registration platforms (AsPredicted, the Open Science Framework, BIT’s own trial registry) made the hardest discipline mechanical. Stepped-wedge designs (Hussey and Hughes 2007 statistical foundation; Hemming, Haines, Chilton, Girling and Lilford 2015 reporting standard) gave political cover to interventions that could not ethically be withheld. The “What Works” centres across the UK government extended the protocol to education, policing, early intervention, ageing, well-being, and local economic growth.

The next decade’s frontier is mechanism. Knowing that an intervention worked is the 2012 win. Knowing why it worked, well enough to predict whether it will work in the next context, is the next mountain. The pairing of trial result with explicit mechanism statement — the heart of the Cartwright / Deaton critique — is the move the field still owes itself. Octalysis-style frameworks that name the motivational mechanism under the intervention are one path. Causal-inference frameworks that name the structural mechanism under the data are another. Neither alone is sufficient. The combined practice will be.

The protocol’s deepest gift to behavioural design is not the trial. It is the loop. Public-sector and private-sector institutions both routinely confuse a single answer for knowledge. The Test, Learn, Adapt insight is that knowledge is the rate at which you compound answers. Protect the loop, and the questions that matter eventually get asked. Break the loop, and the answers you already have decay into folklore. Everything else in the protocol is engineering around that single observation.

Frequently Asked Questions

What is Test, Learn, Adapt in one sentence?

A 2012 UK Cabinet Office paper by Haynes, Service, Goldacre and Torgerson that established a nine-step protocol for using randomised controlled trials to test public-policy interventions, learn from the results, and adapt the policy accordingly.

Who wrote Test, Learn, Adapt?

Laura Haynes (BIT senior advisor), Owain Service (BIT’s first managing director), Ben Goldacre (medical doctor and author of Bad Science), and David Torgerson (director of the York Trials Unit). The paper was published under the UK Cabinet Office imprint with a foreword by David Halpern, founder of the Behavioural Insights Team.

How is Test, Learn, Adapt different from MINDSPACE?

MINDSPACE is a 2010 BIT catalogue of nine intervention levers (Messenger, Incentives, Norms, Defaults, Salience, Priming, Affect, Commitments, Ego). Test, Learn, Adapt is the protocol for empirically testing which of those levers actually moves the dial in your context. MINDSPACE generates hypotheses; Test, Learn, Adapt evaluates them.

How is Test, Learn, Adapt different from EAST?

EAST (Easy, Attractive, Social, Timely) is the post-2014 BIT operational simplification of MINDSPACE for front-line teams. It is a design heuristic. Test, Learn, Adapt is a verification protocol. EAST helps you draft the intervention; Test, Learn, Adapt decides whether the intervention worked.

What is the most common failure mode of a Test, Learn, Adapt deployment?

Selective publication. A team runs the trial, gets a null or negative result, and quietly buries it. Six months later the same intervention is proposed again, the prior literature looks favourable (because the nulls were never published), and the policy goes ahead. The protocol’s pre-registration and registry requirement exists specifically to prevent this. Units that import the protocol without enforcing publication revert to selective reporting within two years.

Are randomised controlled trials always appropriate in public policy?

No. The Cartwright and Deaton critiques are correct that RCTs answer a narrow question well and a wider question poorly. They require a manipulable treatment, a measurable outcome, and a population that can be randomly assigned. That excludes most of macro policy. Inside the band of small, well-specified treatments on individual or local-cluster populations, RCTs are the strongest method available. Outside it, they are silent.

What is the relationship between Test, Learn, Adapt and COM-B?

COM-B (Capability, Opportunity, Motivation, Behavior; Michie, van Stralen, West 2011) is the diagnostic framework that names which of three constraints is binding on a target behavior. Test, Learn, Adapt is the verification framework that decides whether the chosen intervention actually moved the outcome. The chain is COM-B (diagnose) → MINDSPACE or EAST (design) → Test, Learn, Adapt (verify).

How do you handle ethical concerns about randomly withholding an intervention?

Three moves. First, the stepped-wedge design rolls the intervention out to all units in a randomised order, so every unit eventually gets it; the test is on the order, not on whether anyone receives the treatment. Second, the existing policy (which we cannot prove works) is itself a treatment being delivered to one arm by default — the trial’s “control” is what people would have received anyway. Third, the publicity test from Libertarian Paternalism: would the designer be willing to publish the trial design under their own name to the affected population? If yes, the ethics check is satisfied; if no, the trial should not run regardless of design.

Can a small organisation run a Test, Learn, Adapt loop without a dedicated trials unit?

Yes, with caveats. The full protocol scales down. A six-person team can run a two-arm randomised email trial on its own customer base inside a month. The constraint at small scale is statistical power. If the base rate of the outcome is low and the population is small, the trial will be too underpowered to detect anything actionable. The honest answer is to combine trials across many small interventions over time, treating the loop itself as the learning unit, rather than running one large trial.

What is the canonical first trial for a team adopting Test, Learn, Adapt?

An A/B test of two reminder messages on a behavior the team already has administrative data for. The trial is mechanically simple, the cost is near zero, and the population is captive. The team learns the protocol mechanics on a low-stakes problem before it has to defend a randomisation decision on a high-stakes one. The HMRC tax-letter trial was structurally this: an existing letter, an existing population, an existing outcome metric, randomised wording.

References

  1. Haynes, L., Service, O., Goldacre, B., & Torgerson, D. (2012). Test, Learn, Adapt: Developing Public Policy with Randomised Controlled Trials. Cabinet Office Behavioural Insights Team.
  2. Halpern, D. (2015). Inside the Nudge Unit: How Small Changes Can Make a Big Difference. WH Allen.
  3. Dolan, P., Hallsworth, M., Halpern, D., King, D., & Vlaev, I. (2010). MINDSPACE: Influencing Behaviour through Public Policy. Cabinet Office & Institute for Government.
  4. Service, O., Hallsworth, M., Halpern, D., Algate, F., Gallagher, R., Nguyen, S., et al. (2014). EAST: Four Simple Ways to Apply Behavioural Insights. Behavioural Insights Team.
  5. Hallsworth, M., List, J. A., Metcalfe, R. D., & Vlaev, I. (2017). The behavioralist as tax collector: Using natural field experiments to enhance tax compliance. Journal of Public Economics, 148, 14-31.
  6. Cartwright, N. (2007). Are RCTs the gold standard? BioSocieties, 2(1), 11-20.
  7. Cartwright, N., & Hardie, J. (2012). Evidence-Based Policy: A Practical Guide to Doing It Better. Oxford University Press.
  8. Deaton, A., & Cartwright, N. (2018). Understanding and misunderstanding randomized controlled trials. Social Science & Medicine, 210, 2-21.
  9. Duflo, E., Glennerster, R., & Kremer, M. (2007). Using randomization in development economics research: A toolkit. NBER Technical Working Paper 333.
  10. Banerjee, A. V., & Duflo, E. (2009). The experimental approach to development economics. Annual Review of Economics, 1, 151-178.
  11. Imbens, G. W., & Rubin, D. B. (2015). Causal Inference for Statistics, Social, and Biomedical Sciences. Cambridge University Press.
  12. Hussey, M. A., & Hughes, J. P. (2007). Design and analysis of stepped wedge cluster randomized trials. Contemporary Clinical Trials, 28(2), 182-191.
  13. Hemming, K., Haines, T. P., Chilton, P. J., Girling, A. J., & Lilford, R. J. (2015). The stepped wedge cluster randomised trial: rationale, design, analysis, and reporting. BMJ, 350, h391.
  14. Manzi, J. (2012). Uncontrolled: The Surprising Payoff of Trial-and-Error for Business, Politics, and Society. Basic Books.
  15. John, P. (2018). How Far to Nudge? Assessing Behavioural Public Policy. Edward Elgar.
  16. Cochrane, A. L. (1972). Effectiveness and Efficiency: Random Reflections on Health Services. Nuffield Provincial Hospitals Trust.
  17. Sanders, M., Snijders, V., & Hallsworth, M. (2018). Behavioural science and policy: where are we now and where are we going? Behavioural Public Policy, 2(2), 144-167.
  18. Linos, E., & Reinhard, J. (2015). A head for hiring: The behavioural science of recruitment and selection. CIPD.
  19. Michie, S., van Stralen, M. M., & West, R. (2011). The behaviour change wheel: A new method for characterising and designing behaviour change interventions. Implementation Science, 6, 42.
  20. Sunstein, C. R. (2014). Why Nudge? The Politics of Libertarian Paternalism. Yale University Press.
  21. Chou, Y. (2015). Actionable Gamification: Beyond Points, Badges, and Leaderboards. Octalysis Media.

WOULD YOU LIKE YU-KAI CHOU TO WORK WITH YOUR ORGANIZATION?

Yukaichou.com Main Contact Form

Continue your training

Reading is XP. Now test what drives you — or pick a quest path.

Keep exploring

Related articles