ImpactMojo ImpactMojo
Premium
Theories of Development · Experimentalist · 1997–present

Randomised evaluation and the small-question turn

Grand theories of development cannot be tested, but specific questions about specific programmes can be, by assigning them at random and comparing, and enough answered small questions is a better guide to policy than any answered large one.

Esther Duflo, Abhijit Banerjee, Michael Kremer · Credits and sources ↓

The argument

What it says

The starting position is deliberately modest and slightly rude about everything else on this shelf. Nobody can settle whether aid works, whether markets or states develop a country, or whether institutions cause growth, because the units are countries and there are not enough of them and they differ in every respect at once. But whether a second teacher raises learning in this district, whether bed nets are used when they are free rather than sold, whether a nudge raises immunisation, are questions with an answer that can be obtained.

The method borrows from medicine. Assign the programme at random across schools, villages or households, so that treated and untreated groups differ only by the assignment, and any difference in the outcome afterwards is caused by the programme. The apparatus around it, done properly, is large: pre-analysis plans, power calculations, tracking of attrition, and replication in other places.

The intellectual pay-off is not the individual estimate but what accumulates. Poor Economics argues from many trials at once that poor households behave in ways that are intelligible once the constraints are correctly described, and that the constraints are frequently not the ones policy assumed: not that people do not value schooling, but that the return is concentrated at levels they do not reach; not that they do not want immunisation, but that the marginal cost of the trip on the day exceeds a diffuse future benefit.

The strongest single case is Indian. Trials with Pratham of teaching children at the level they are actually at, rather than at the level the syllabus assumes, produced large learning gains, survived replication, and were then adopted by state governments at a scale the method's critics said it would never reach.

The causal chain

Drawn one step at a time

The theory as a graph, revealed a layer at a time. Use the buttons, or the left and right arrow keys. Each step adds the boxes that step introduces and the arrows into them.

  • Starting condition
  • Mechanism
  • Outcome or policy
Where it sits

Four placements, and the reason for each

The four scores are editorial. They run from -3 to +3, they were assigned by the ImpactMojo editorial team from the theory's own texts, and each theory page shows the sentence that justifies its placement so the placement can be argued with. They are a way of arranging a shelf, not a measurement.

Who allocates

Agnostic by design, and criticised for it. The method evaluates an intervention without needing a position on who should be running the economy, which is a strength for a trial and a limitation for a theory.

What comes first

Poverty is the subject, and the interventions are mostly redistributive in effect, but the approach has no position on aggregate growth at all.

Where hierarchy sits

Caste and gender enter as variables to be measured, sometimes with sharp results, rather than as structures the method has anything to say about.

Who moves

Researchers, governments and implementing organisations. This is the most expert-driven entry on the shelf, and the criticism about whose questions get asked follows directly from that.

What happened

The theory against the record

Each entry takes one claim the theory makes and reports what the evidence says about it, with a named source and a year. This section is the reason the library exists; a catalogue of positions without it is a reading list.

The claim A trial result can reach national scale.

The clearest case is Indian. Remedial teaching at the child's actual level, tested with Pratham in Mumbai and Vadodara and reported in 2007, produced substantial learning gains, was replicated across several further trials, and was then adopted by state governments including Haryana and Bihar and by governments outside India. It is the strongest available answer to the charge that the method cannot scale.

Banerjee, Cole, Duflo and Linden, Quarterly Journal of Economics; subsequent Pratham and state government programmes · 2007

The claim Randomisation settles the causal question.

For the trial's own population, largely yes, and that is a real achievement. Angus Deaton's objection is about what follows: an unbiased average effect can conceal that the programme helped some and harmed others, tells you nothing about why it worked, and therefore provides little basis for predicting the effect anywhere else. Deaton and Nancy Cartwright set out the argument in detail, and their position is not that trials are useless but that they carry less authority than the field claims.

Angus Deaton, 'Instruments, Randomization, and Learning about Development', Journal of Economic Literature; Deaton and Cartwright, Social Science and Medicine · 2018

The claim Enough small answers add up to a guide for policy.

Partial, and the shortfall is systematic rather than random. What can be randomised is what can be assigned by an implementer, which excludes exchange rates, trade policy, labour law, land reform, and the structure of the state. Lant Pritchett's version of the objection is that the biggest measured improvements in poor people's lives have come from national growth episodes that no trial could have evaluated.

Lant Pritchett and Justin Sandefur on external validity and policy relevance; the wider debate following the 2019 Nobel · 2019

The claim The method is ethically neutral about the people it studies.

Not conceded by everyone, and the objection is worth stating from within the field rather than against it. Trials assign a programme by lottery to people who did not design the question, in countries whose researchers are frequently the field staff rather than the authors, and the Indian debate on this, conducted in the Economic and Political Weekly among others, has been sharper than the one abroad.

The debate in Economic and Political Weekly and elsewhere following the 2019 Nobel award · 2019

What this does not settle

The method's authority and its scope have never been reconciled, and both sides of the argument are correct about the half they hold. A well-run trial does establish what a programme did in a place, and nothing else in economics establishes that as cleanly. It also cannot reach the questions on which the largest differences in human welfare have turned, and it selects for questions that an implementing organisation can act on, which is a filter nobody chose and everybody works inside. What has not been tested, because it cannot be, is the field's implicit claim that the sum of what is randomisable is where the leverage is.

In India

How it landed here

India is where more of these trials have been run than anywhere else, which is worth noticing as a fact about the method as well as about India: the country combines large administrative units, capable implementing organisations and long-standing research relationships, so it supplies what a trial needs. The teaching at the right level result is the method's best Indian outcome. The standing Indian criticism is that a country with the Annual Status of Education Report, the National Sample Survey and the National Family Health Survey has known the shape of its problems for decades, and that the binding constraint on acting was never the absence of a causal estimate.

Read next

One that agrees, one that does not

Closest to it

Capabilities and human development

Both refuse to take income as the measure and both look at what a specific person can actually do. One arrives at it philosophically, the other from field data.

Furthest from it

Dependency and unequal exchange

The clearest disagreement about what a theory is for. One holds that the structure is the thing to explain and cannot be experimented on; the other holds that a claim you cannot test is not doing work.

Credit where it is owed

Whose theory this is

Esther Duflo born 1972

Argued that the economist working on poverty should think of the job as plumbing: the details of how a policy is installed determine whether it works, and they cannot be deduced.

Abhijit Banerjee born 1961

Much of the Indian field work, including the education trials with Pratham that became the most-scaled result the method has produced.

Michael Kremer born 1964

Ran the first of these trials in Kenyan schools in the 1990s, which is where the approach starts.

What ImpactMojo added

The causal diagram, the four placements and the notes justifying them, and the evidence section: what each claim predicted and what the record shows, with a named source and year for every entry.

ImpactMojo · content CC BY-NC-ND 4.0 · code MIT

Start with these

  • Abhijit Banerjee and Esther Duflo, Poor Economics (2011). The accumulated picture rather than the method. Organised by constraint rather than by sector, which is the argument in the structure.
  • Angus Deaton, Instruments, Randomization, and Learning about Development (2010). Journal of Economic Literature. The most careful statement of what a randomised estimate does and does not license.
  • Banerjee, Cole, Duflo and Linden, Remedying Education: Evidence from Two Randomized Experiments in India (2007). Quarterly Journal of Economics. The Pratham trials, and the origin of teaching at the right level.

Open access, in Development Discourses: