Adaptive Experiments Lab
Many nonprofits now run a programme through WhatsApp, an app or a call centre, and can test two versions of a message on real users within a week. This Lab shows how such tests work, using simulated users whose true response rates you set yourself. You will run a fixed A/B test, run a bandit on the same users, see what a bandit costs you in statistical accuracy, and check whether this kind of experiment suits your programme.
Improving a running programme and measuring its impact are separate questions
Both use random assignment. They are asked by different people, for different decisions, and the answer to one cannot be used as the answer to the other.
A programme team that sends early-childhood activities to parents over WhatsApp might ask whether a reminder at 7 in the evening gets more parents to do the day's activity than one at 10 in the morning. The programme is already running, both versions are acceptable, and the team wants to run it better. Random assignment answers that question cheaply, because the team can send the two reminders to two random halves of its parents and count what happens within days.
A funder deciding whether to pay for the same programme in another state asks something else: would these children be better prepared for school than children whose parents received nothing? That needs a comparison with people who do not get the programme at all, an outcome measured months later, and usually an independent evaluator. The Impact Evaluation Designer and the RCT Readiness Diagnostic in these Labs cover that question.
Fast experiments of the first kind are now within reach of small teams. Evidential, a free and open-source experiment engine built by IDinsight and the Agency Fund, lists among the questions its early partners have tested whether a casual tone with younger users raises completion rates and what time of day to send messages. Its documentation also says plainly that it is not built for full-scale randomised controlled trials.
Sort these six questions
For each one, choose which kind of question it is. The explanation appears when you choose.
A fixed A/B test
Every version gets an equal share of users from start to finish, and you compare them at the end.
In a real test nobody knows the true response rate of each version; that is why the test is run. In this simulator you set the true rates, so that you can see how close each method gets to them and what each method costs. A response here means whatever you count: a reply, an opened link, a completed activity. Each simulated user responds at random with the probability you set for the version they receive.
Simulated data Nothing on this page is a result from a real programme.
A bandit moves users towards the version that is doing better
This one uses Thompson sampling, one of the standard rules for bandits.
The first batch of 100 users is split equally. After each batch, the simulator looks at the responses so far and, for every version, draws a plausible response rate from what the data allow (a Beta distribution, which is what the counts of responses and non-responses imply). It does this 200 times and counts how often each version comes out on top. The next batch is split in those proportions. A version that is clearly ahead gets most of the next batch; a version that is uncertain still gets some users, so the bandit can change its mind.
William Thompson proposed the rule in 1933 for allocating patients between two treatments. Kasy and Sautmann (2021) adapted it into a rule they call exploration sampling, built for experiments whose purpose is to pick one policy for a later rollout, and used it in Odisha to choose among six ways of recruiting farmers to an agricultural extension service, working with Precision Agriculture for Development.
The bandit below uses the rates, number of users and seed you set in the previous step, so the two methods face the same population.
The usual estimates go wrong when allocation follows the results
Run the same experiment hundreds of times and count how often the usual 95% interval contains the true difference.
A version that has bad luck in the first batches gets fewer users from then on, so its low early estimate is corrected slowly. A version that has good luck gets more users, and its high early estimate is corrected quickly. Low errors persist and high errors are washed out, so the estimated rates come out too low on average, and the ordinary 95% interval for the difference between versions contains the truth less often than 95 times in 100. Hadad, Hirshberg, Zhan, Wager and Athey (2021) set out the problem and give estimators that correct it.
The simulation below runs two versions many times, once with a fixed equal split and once with the bandit, and computes the ordinary interval each time. Start with two versions that are equally good, which is where the problem is largest.
A fast experiment finds the version that wins on what you count fast
Replies and clicks arrive within hours. The change a programme exists to make takes longer to see.
Suppose a team tests two messages. Version A is short and ends with a one-tap quiz. Version B is longer and asks the parent to do a ten-minute activity with the child and send a photo. Parents reply to A more often. Whether the parent then does the activity with the child is the outcome the programme is for, and it can only be checked through the photo or a later call. Change the numbers below and watch which version each choice of metric picks.
Illustrative numbers Invented for teaching.
Evidential's documentation makes the same point about the metrics an experiment should target: they "should be the primary metrics that the organization holds itself accountable to." Before you use a quick measure as the target, check on past data whether it moves with the outcome you care about. If you cannot check, run the test long enough to observe the outcome, or treat the quick result as a lead to test properly.
Does this kind of experiment suit your programme?
Answer for one change you would like to test. The verdict counts your answers and names what is missing.
One tool, as an example: Evidential
Evidential was built by IDinsight and the Agency Fund and is released under the Apache 2.0 licence. It is free; an organisation can use the managed service or run it on its own servers, and it reads outcome data from the organisation's own data warehouse or event logs. It supports frequentist A/B tests, Bayesian tests, and multi-armed and contextual bandits, and integrations with messaging and survey platforms are listed as coming soon. Its documentation thanks Rocket Learning as its earliest design partner, and Noora Health, myAgro, Jacaranda Health and Saajha for running experiments with it.
Its published analysis page, read on 9 October 2026, describes the A/B analysis: power calculated for a two-sample t-test at 80% power and a 5% significance level, a default target effect of 10% of the baseline, a balance check, and an ordinary least squares estimate of the intent-to-treat effect with pre-specified strata and pre-treatment outcomes as controls. It does not describe how results from its bandits are analysed. If you use a bandit on any platform and plan to report an effect size, ask how its intervals allow for adaptive allocation.
Evidential at IDinsight · Evidential documentation
Sources
- Thompson, W. R. (1933). On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika, 25(3/4), 285–294.
- Russo, D., Van Roy, B., Kazerouni, A., Osband, I. and Wen, Z. (2018). A tutorial on Thompson sampling. Foundations and Trends in Machine Learning, 11(1), 1–96.
- Kasy, M. and Sautmann, A. (2021). Adaptive treatment assignment in experiments for policy choice. Econometrica, 89(1), 113–132.
- Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S. and Athey, S. (2021). Confidence intervals for policy evaluation in adaptive experiments. Proceedings of the National Academy of Sciences, 118(15), e2014602118.
- Evidential documentation, docs.evidential.dev: Welcome page and Data Analysis page, read 9 October 2026.
Where next
If your question is whether the programme works at all, go to the RCT Readiness Diagnostic and then the Impact Evaluation Designer.