Randomised evaluation and the small-question turn
Grand theories of development cannot be tested, but specific questions about specific programmes can be, by assigning them at random and comparing, and enough answered small questions is a better guide to policy than any answered large one.
Esther Duflo, Abhijit Banerjee, Michael Kremer · Credits and sources ↓
What it says
The starting position is deliberately modest and slightly rude about everything else on this shelf. Nobody can settle whether aid works, whether markets or states develop a country, or whether institutions cause growth, because the units are countries and there are not enough of them and they differ in every respect at once. But whether a second teacher raises learning in this district, whether bed nets are used when they are free rather than sold, whether a nudge raises immunisation, are questions with an answer that can be obtained.
The method borrows from medicine. Assign the programme at random across schools, villages or households, so that treated and untreated groups differ only by the assignment, and any difference in the outcome afterwards is caused by the programme. The apparatus around it, done properly, is large: pre-analysis plans, power calculations, tracking of attrition, and replication in other places.
The intellectual pay-off is not the individual estimate but what accumulates. Poor Economics argues from many trials at once that poor households behave in ways that are intelligible once the constraints are correctly described, and that the constraints are frequently not the ones policy assumed: not that people do not value schooling, but that the return is concentrated at levels they do not reach; not that they do not want immunisation, but that the marginal cost of the trip on the day exceeds a diffuse future benefit.
The strongest single case is Indian. Trials with Pratham of teaching children at the level they are actually at, rather than at the level the syllabus assumes, produced large learning gains, survived replication, and were then adopted by state governments at a scale the method's critics said it would never reach.
Drawn one step at a time
The theory as a graph, revealed a layer at a time. Use the buttons, or the left and right arrow keys. Each step adds the boxes that step introduces and the arrows into them.
- Starting condition
- Mechanism
- Outcome or policy
Four placements, and the reason for each
The four scores are editorial. They run from -3 to +3, they were assigned by the ImpactMojo editorial team from the theory's own texts, and each theory page shows the sentence that justifies its placement so the placement can be argued with. They are a way of arranging a shelf, not a measurement.
Agnostic by design, and criticised for it. The method evaluates an intervention without needing a position on who should be running the economy, which is a strength for a trial and a limitation for a theory.
Poverty is the subject, and the interventions are mostly redistributive in effect, but the approach has no position on aggregate growth at all.
Caste and gender enter as variables to be measured, sometimes with sharp results, rather than as structures the method has anything to say about.
Researchers, governments and implementing organisations. This is the most expert-driven entry on the shelf, and the criticism about whose questions get asked follows directly from that.
The theory against the record
Each entry takes one claim the theory makes and reports what the evidence says about it, with a named source and a year. This section is the reason the library exists; a catalogue of positions without it is a reading list.
The claim A trial result can reach national scale.
The clearest case is Indian. Remedial teaching at the child's actual level, tested with Pratham in Mumbai and Vadodara and reported in 2007, produced substantial learning gains, was replicated across several further trials, and was then adopted by state governments including Haryana and Bihar and by governments outside India. It is the strongest available answer to the charge that the method cannot scale.
Banerjee, Cole, Duflo and Linden, Quarterly Journal of Economics; subsequent Pratham and state government programmes · 2007
The claim Randomisation settles the causal question.
For the trial's own population, largely yes, and that is a real achievement. Angus Deaton's objection is about what follows: an unbiased average effect can conceal that the programme helped some and harmed others, tells you nothing about why it worked, and therefore provides little basis for predicting the effect anywhere else. Deaton and Nancy Cartwright set out the argument in detail, and their position is not that trials are useless but that they carry less authority than the field claims.
Angus Deaton, 'Instruments, Randomization, and Learning about Development', Journal of Economic Literature; Deaton and Cartwright, Social Science and Medicine · 2018
The claim Enough small answers add up to a guide for policy.
Partial, and the shortfall is systematic rather than random. What can be randomised is what can be assigned by an implementer, which excludes exchange rates, trade policy, labour law, land reform, and the structure of the state. Lant Pritchett's version of the objection is that the biggest measured improvements in poor people's lives have come from national growth episodes that no trial could have evaluated.
Lant Pritchett and Justin Sandefur on external validity and policy relevance; the wider debate following the 2019 Nobel · 2019
The claim The method is ethically neutral about the people it studies.
Not conceded by everyone, and the objection is worth stating from within the field rather than against it. Trials assign a programme by lottery to people who did not design the question, in countries whose researchers are frequently the field staff rather than the authors, and the Indian debate on this, conducted in the Economic and Political Weekly among others, has been sharper than the one abroad.
The debate in Economic and Political Weekly and elsewhere following the 2019 Nobel award · 2019
What this does not settle
The method's authority and its scope have never been reconciled, and both sides of the argument are correct about the half they hold. A well-run trial does establish what a programme did in a place, and nothing else in economics establishes that as cleanly. It also cannot reach the questions on which the largest differences in human welfare have turned, and it selects for questions that an implementing organisation can act on, which is a filter nobody chose and everybody works inside. What has not been tested, because it cannot be, is the field's implicit claim that the sum of what is randomisable is where the leverage is.
How it landed here
India is where more of these trials have been run than anywhere else, which is worth noticing as a fact about the method as well as about India: the country combines large administrative units, capable implementing organisations and long-standing research relationships, so it supplies what a trial needs. The teaching at the right level result is the method's best Indian outcome. The standing Indian criticism is that a country with the Annual Status of Education Report, the National Sample Survey and the National Family Health Survey has known the shape of its problems for decades, and that the binding constraint on acting was never the absence of a causal estimate.
One that agrees, one that does not
Capabilities and human development
Both refuse to take income as the measure and both look at what a specific person can actually do. One arrives at it philosophically, the other from field data.
Dependency and unequal exchange
The clearest disagreement about what a theory is for. One holds that the structure is the thing to explain and cannot be experimented on; the other holds that a claim you cannot test is not doing work.
Whose theory this is
Esther Duflo born 1972
Argued that the economist working on poverty should think of the job as plumbing: the details of how a policy is installed determine whether it works, and they cannot be deduced.
Abhijit Banerjee born 1961
Much of the Indian field work, including the education trials with Pratham that became the most-scaled result the method has produced.
Michael Kremer born 1964
Ran the first of these trials in Kenyan schools in the 1990s, which is where the approach starts.
What ImpactMojo added
The causal diagram, the four placements and the notes justifying them, and the evidence section: what each claim predicted and what the record shows, with a named source and year for every entry.
ImpactMojo · content CC BY-NC-ND 4.0 · code MIT
Start with these
- Abhijit Banerjee and Esther Duflo, Poor Economics (2011). The accumulated picture rather than the method. Organised by constraint rather than by sector, which is the argument in the structure.
- Angus Deaton, Instruments, Randomization, and Learning about Development (2010). Journal of Economic Literature. The most careful statement of what a randomised estimate does and does not license.
- Banerjee, Cole, Duflo and Linden, Remedying Education: Evidence from Two Randomized Experiments in India (2007). Quarterly Journal of Economics. The Pratham trials, and the origin of teaching at the right level.
Open access, in Development Discourses:
- Randomization and Social Policy Evaluation Revisited — James J. Heckman (2020)
- Experimentation at Scale — Karthik Muralidharan, Paul Niehaus (2017)
- Impact Evaluation in Practice, Second Edition — Paul J. Gertler, Sebastian Martinez, Patrick Premand, Laura B. Rawlings, Christel M. J. Vermeersch (2016)