ImpactMojoCausal Inference 101www.impactmojo.in
ImpactMojo 101 Series · Free Forever
Causal
Inference
101
What would have happened otherwise: the logic of causal claims for development practitioners, from the counterfactual to randomisation, difference-in-differences, regression discontinuity and instrumental variables
Evaluation MethodsSouth Asia Focus100 SlidesFree Forever
ImpactMojoCausal Inference 101www.impactmojo.in
What we cover
01
Why causal questions are hard
Slides 3–10
02
Potential outcomes and the counterfactual
Slides 11–19
03
Selection bias and confounding
Slides 20–27
04
Causal diagrams
Slides 28–35
05
Randomisation
Slides 36–45
06
Natural experiments and difference-in-differences
Slides 46–54
07
Regression discontinuity
Slides 55–62
08
Instrumental variables
Slides 63–70
09
Matching and synthetic control
Slides 71–78
10
External validity and heterogeneity
Slides 79–86
11
Reading a causal claim in practice
Slides 87–94
12
Summary and next steps
Slides 95–99
ImpactMojoCausal Inference 101www.impactmojo.in
01
Section One
Why causal questions are hard
ImpactMojoCausal Inference 101www.impactmojo.in
Describe, predict, or explain what a change would do
Most questions a programme team asks fall into one of three kinds, and each needs a different kind of evidence. A descriptive question asks what is the case. A predictive question asks what we expect to observe next. A causal question asks what would change if we, or someone, did something differently. Only the third tells a ministry whether to spend money, and it is the hardest of the three, because the answer compares the world we see with a world we never see.
KindExample questionWhat answers it
DescriptiveWhat share of births in Bihar happen in a health facility?A good sample survey such as NFHS, with weights
PredictiveWhich blocks are likely to report the most malnourished children next year?A model that fits past data well; it need not explain anything
CausalDid a cash incentive to mothers raise facility births, and by how much?A comparison that stands in for the births that would have happened without the incentive
Causal, at scaleWould the same incentive work if every state ran it through its own health department?Causal evidence plus an argument about context and implementation
This course is about the third and fourth rows. The first two are covered in Data Literacy 101 and Survey Design 101.
ImpactMojoCausal Inference 101www.impactmojo.in
Two things that move together may have no effect on each other
An association is a pattern in data: where one thing is higher, another tends to be higher or lower too. A causal effect is a claim about what happens to one thing when another is changed. The two separate whenever something else drives both, or when the arrow runs the other way. Development data are full of both situations because programmes are placed on purpose and people choose whether to join them.
A common cause (Illustrative)
Villages with a bank branch have higher household incomes. A planner concludes that branches raise incomes. Banks, though, open branches where incomes, roads and markets already exist. The same prosperity that attracted the branch would show up in incomes even if the branch had never opened.
Reverse causation (Illustrative)
Districts that spend more on tuberculosis treatment report more tuberculosis cases. Spending does not cause the disease. High caseloads draw the budget, and better funded programmes also find and notify more cases, so the reported numbers rise with the spending that was meant to bring them down.
ImpactMojoCausal Inference 101www.impactmojo.in
A wrong causal answer costs money and lives
Governments in South Asia run some of the largest social programmes in the world, and each one is a bet about cause and effect. A school feeding scheme assumes meals raise attendance and learning. A conditional cash transfer for institutional delivery assumes the cash moves births into facilities and that facility births save newborns. If the first link holds and the second does not, the scheme spends its budget and the mortality it was built to reduce stays where it was.
Causal effect
The difference between the outcome a unit (a person, a household, a village) has with a treatment and the outcome the same unit would have had, at the same time, without it. Everything in this course is a way of estimating that difference when half of it can never be observed.
Causal evidence will not tell you what to value. It tells you what a decision is likely to do, which is the part that can be checked.
ImpactMojoCausal Inference 101www.impactmojo.in
Two Nobel prizes for getting the comparison right
2019
Abhijit Banerjee, Esther Duflo and Michael Kremer, "for their experimental approach to alleviating global poverty"
Nobel Prize in Economic Sciences 2019, nobelprize.org
2021
Joshua Angrist and Guido Imbens, "for their methodological contributions to the analysis of causal relationships" (shared with David Card)
Nobel Prize in Economic Sciences 2021, nobelprize.org press release
Both prizes rewarded the same shift. From the 1980s, economists stopped trusting a regression because it had many control variables and started asking where the variation in the treatment came from. Was it a lottery, a rule, a border, a date? If the answer was "people chose it", the estimate inherited every reason they chose. Much of the experimental work honoured in 2019 was run in India, including the remedial education trials in Mumbai and Vadodara described in Section 5. Card's half of the 2021 prize was for empirical contributions to labour economics; his minimum wage study with Alan Krueger appears in Section 6.
ImpactMojoCausal Inference 101www.impactmojo.in
Four routes from a real pattern to a false conclusion
  • Selection. The people who take up a programme differ from those who do not, in ways that also shape the outcome. Self-help group members were often already more connected.
  • Confounding. A third factor drives both treatment and outcome. Land ownership raises both access to credit and farm income.
  • Reverse causation. The outcome drives the treatment. Sick children are taken to clinics, so clinic visits correlate with illness.
  • Chance and choice of analysis. With enough outcomes and subgroups, something will look significant. Reporting only that result manufactures an effect.
01
PATTERN: participants do better
→
02
QUESTION: better than whom?
→
03
CHECK: why did they participate?
→
04
DESIGN: find variation they did not choose
→
05
CLAIM: effect, with its assumptions stated
Each design later in this course is a defence against one or more of these four. Knowing which defence a study used tells you which failure to look for.
ImpactMojoCausal Inference 101www.impactmojo.in
One cash transfer, two published answers
India launched the Janani Suraksha Yojana (JSY) in 2005, paying women who give birth in a health facility (Lim et al., The Lancet 2010). Two careful studies of the same scheme reached different conclusions on the outcome that mattered most.
Lim and colleagues, The Lancet, 2010
District Level Household Surveys of 2002–04 and 2007–09, analysed by matching, a with-versus-without comparison and difference-in-differences. JSY raised antenatal care and facility births. In the matching analysis, JSY payment was associated with 3.7 fewer perinatal deaths per 1,000 pregnancies and 2.3 fewer neonatal deaths per 1,000 live births.
Powell-Jackson, Mazumdar and Mills, J. Health Economics, 2015
Difference-in-differences exploiting variation in how intensively districts implemented JSY. Cash incentives were associated with more use of maternity services, but there was no strong evidence of a reduction in neonatal or early neonatal mortality. They also found more pregnancies and a shift away from private providers.
We return to this pair in Section 9. By then you will be able to say why the two designs could disagree and which assumptions each one needs.
ImpactMojoCausal Inference 101www.impactmojo.in
What this course covers, and where to go for the rest
Covered here
  • The counterfactual and the vocabulary of treatment effects
  • Selection bias, confounders, mediators and colliders, with diagrams
  • Why randomisation works and what still goes wrong in trials
  • Difference-in-differences, regression discontinuity, instrumental variables, matching and synthetic control, each with an Indian or South Asian study
  • External validity, heterogeneity and scale
  • A checklist for reading any causal claim
Covered elsewhere
  • Running regressions and reading output: Econometrics 101
  • Designing and commissioning an evaluation end to end: Impact Evaluation 101
  • Sample size and power calculations: Impact Evaluation 101
  • Questionnaires and survey sampling: Survey Design 101
  • Pooling many studies: Systematic Reviews 101
  • Formal estimation in code: the flagship Causal Inference course
No mathematics beyond averages and subtraction is needed. Where a formula appears it is written out in words beside it.
ImpactMojoCausal Inference 101www.impactmojo.in
02
Section Two
Potential outcomes and the counterfactual
ImpactMojoCausal Inference 101www.impactmojo.in
Every unit has two possible outcomes, and we see one
Potential outcomes
For each unit there is an outcome if treated, written Y(1), and an outcome if untreated, written Y(0). The causal effect for that unit is Y(1) minus Y(0). Both exist in principle before treatment is assigned; assignment decides which one we get to observe.
Donald Rubin set out this framework for observational and randomised studies in the Journal of Educational Psychology in 1974 (vol. 66, no. 5, pp. 688–701), and it is often called the Rubin causal model. Its value is that it forces a precise statement. "The scholarship improved girls' schooling" becomes "for these girls, years of schooling with the scholarship minus years of schooling without it, at the same time and place, is positive on average". Once the claim is written that way, the missing half is obvious: no girl was both given and denied the scholarship.
Every method in this course is a way of filling in the missing potential outcome with something defensible.
ImpactMojoCausal Inference 101www.impactmojo.in
You cannot observe the same person treated and untreated
Paul Holland named this the fundamental problem of causal inference in "Statistics and Causal Inference", Journal of the American Statistical Association, 1986 (vol. 81, no. 396, pp. 945–960). It is a problem of missing data, and no amount of extra data of the same kind solves it. Surveying a million households that joined a programme still tells you nothing about what those households would have done had they not joined.
The way out is to give up on individual effects and estimate averages. If we can find a group whose average untreated outcome equals what the treated group's average untreated outcome would have been, the difference in averages is an average causal effect. The rest of the course is about finding that group.
The two ways to fill the gap
  • Different units, same time. Compare treated households with untreated ones that would have fared the same.
  • Same units, different time. Compare households with their own past, assuming nothing else changed.
  • Most real designs combine the two, and each brings its own assumption.
ImpactMojoCausal Inference 101www.impactmojo.in
Five women, ten outcomes, five observed
Illustrative. Monthly earnings in rupees for five women offered a tailoring course. In real data only the shaded column for each woman exists.
WomanTook course?Y(1): earnings if trainedY(0): earnings if notIndividual effectWhat we observe
AshaYes₹6,000₹5,000+₹1,000₹6,000
BinaYes₹7,000₹6,500+₹500₹7,000
ChandniYes₹5,500₹4,000+₹1,500₹5,500
DeviNo₹3,500₹2,500+₹1,000₹2,500
EshaNo₹3,000₹2,000+₹1,000₹2,000
True average effect for the three who trained: ₹1,000. True average for all five: ₹1,000.
Naive comparison of observed means: ₹6,167 minus ₹2,250 = ₹3,917. Nearly four times too large, because the women who chose the course already earned more.
ImpactMojoCausal Inference 101www.impactmojo.in
ATE, ATT and LATE answer different policy questions
EstimandPlain meaningPolicy question it answers
ATE: average treatment effectAverage of Y(1) minus Y(0) over everyone in the populationWhat if we gave it to everybody?
ATT: average effect on the treatedAverage effect among those who actually received itWas it worth it for the people who got it?
ATU: average effect on the untreatedAverage effect among those who did not receive itWhat would expansion to the rest achieve?
ITT: intention to treatEffect of being offered or assigned, whether or not the unit took it upWhat does announcing the programme achieve, given real take-up?
LATE: local average treatment effectAverage effect among units whose take-up was changed by the offer or instrumentWhat does it do for people who respond to the nudge?
These are equal only when effects are the same for everyone, which they rarely are. A microcredit programme can have a large ATT among existing entrepreneurs who borrowed and a small ATE across all households, most of whom never wanted a loan. A report that says "the effect" without saying which one has left out the most useful sentence.
ImpactMojoCausal Inference 101www.impactmojo.in
The comparison group is the counterfactual
Every causal estimate is a comparison, and the comparison group is the study's guess at the missing potential outcome. Asking "compared with whom?" is the single most useful question a reader can put to a claim, because the answer tells you what had to be true for the number to be right.
Weak comparisons
  • Participants against eligible people who declined
  • Programme districts against districts the state chose to leave out
  • The same village before and after, in a year with a better monsoon
  • Beneficiaries against the state average
Stronger comparisons
  • Applicants who lost a lottery for places
  • Districts just below and above an eligibility score
  • Late-phase districts, before their phase began, tracked alongside early ones
  • Villages randomly assigned to wait a year
The stronger comparisons share one property: something other than the units' own choices or the implementer's judgement decided who was treated.
ImpactMojoCausal Inference 101www.impactmojo.in
Before-after comparisons credit the programme with everything else
A before-after comparison uses each unit's own past as its counterfactual. It assumes the outcome would have stayed exactly where it was. In a fast-changing economy that assumption fails almost by default: wages, prices, rainfall, roads, phones and other schemes all move between the two survey rounds.
Illustrative: a district reports that average yields rose 18% in the two years after a soil health card drive. If the second year had a normal monsoon after a drought, much of the rise would have happened anyway. The before-after number measures the card drive plus the rain plus every other change, and cannot tell them apart.
When before-after is all you have
  • Plot several years of the outcome before the programme, if they exist
  • Look at the same outcome in places without the programme over the same years
  • Name the other changes in the period and say which way each would push
  • Report the result as a change over time; call it an effect only with a comparison group
ImpactMojoCausal Inference 101www.impactmojo.in
No spillovers and one version of treatment
Stable unit treatment value assumption (SUTVA)
Two conditions behind every simple estimate. First, one unit's outcome does not depend on whether other units were treated (no interference). Second, the treatment is the same thing for everyone who receives it (no hidden versions).
Interference in practice
Deworming one child lowers infection risk for classmates. A public works programme that hires many labourers can raise wages for workers who never joined it, which Imbert and Papp (2015) found for India's employment guarantee. If untreated units are affected, comparing them with treated units misstates the total effect.
Hidden versions in practice
"Received the training" can mean a three-day course run well in one block and a half-day session run badly in another. Averaging the two produces an effect for a treatment that nobody received. Record what was delivered, where, and how much.
Randomising whole villages or schools, with everyone inside a unit in the same arm, is a common way to contain spillovers inside the unit.
ImpactMojoCausal Inference 101www.impactmojo.in
A naive difference equals the effect plus selection bias
Observed difference in means = ATT + selection bias, where selection bias = average Y(0) of the treated minus average Y(0) of the untreated.
Read the equation in words. The gap we see between participants and non-participants has two parts. One is what the programme did for participants. The other is how different the two groups would have been with no programme at all. In the tailoring table, the observed gap was ₹3,917, the effect on the trained was ₹1,000, and the remaining ₹2,917 was selection bias: the trained women would have earned ₹5,167 on average without the course, against ₹2,250 for the others.
This identity, set out in Angrist and Pischke's Mostly Harmless Econometrics (Princeton University Press, 2009), explains the whole course in one line. Randomisation makes the second term zero in expectation. Difference-in-differences removes the part of it that is fixed over time. Regression discontinuity makes it small near a cutoff. Matching removes the part explained by what you measured. Each method attacks the same term.
ImpactMojoCausal Inference 101www.impactmojo.in
03
Section Three
Selection bias and confounding
ImpactMojoCausal Inference 101www.impactmojo.in
Who joins a programme is never an accident
Participation is a decision, made by people with reasons, and the reasons usually relate to the outcome. Farmers who adopt a new seed tend to have irrigation, credit and information. Parents who enrol a child in an after-school class tend to care a great deal about schooling. Women who join a self-help group in its first year tend to be the ones with time, a network and some savings.
Positive selection
The better-off or more motivated join, so participants would have done well anyway. The naive estimate is too large. Most voluntary training, credit and technology programmes are at risk of this.
Negative selection
The worse-off are targeted or join out of need, so participants would have done worse anyway. The naive estimate is too small and can even flip sign. Nutrition rehabilitation centres admit the most malnourished children, so their graduates can look worse than children who never needed the centre.
Targeting is a form of selection. A well-targeted programme is, for the evaluator, a programme whose beneficiaries differ systematically from everyone else.
ImpactMojoCausal Inference 101www.impactmojo.in
Microcredit in Bangladesh: the estimate that did not survive reanalysis
Pitt and Khandker, JPE 1998
Studied participation in Grameen Bank and two other group-based credit programmes, using a quasi-experimental survey design to correct for unobserved individual and village-level differences. Reported that household consumption rose 18 taka for every 100 taka borrowed by women, against 11 taka for men. Roodman and Morduch later called it the most influential study of microcredit impacts. (Journal of Political Economy 106(5): 958–996.)
Roodman and Morduch, J. Development Studies 2014
Replicated and reanalysed the same data. The poverty results disappeared after dropping outliers or using a less fragile estimator, and assumptions the original analysis relied on, such as normally distributed errors, were contradicted by the data. Their conclusion: questions about impact cannot be answered in these data. Pitt published a reply in the same issue. (50(4): 583–604.)
The lesson concerns selection, and it applies well beyond microcredit. When participation is chosen, the estimate rests on modelling assumptions, and those assumptions can be tested and can fail. A randomised trial of group lending in Hyderabad (Banerjee et al., AEJ: Applied 2015) found higher business investment but no significant rise in consumption, health or education.
ImpactMojoCausal Inference 101www.impactmojo.in
A confounder causes both the treatment and the outcome
Confounder
A variable that affects whether a unit is treated and also affects the outcome through another path. Leaving it out mixes its effect into the estimated effect of the treatment.
Illustrative. A study finds that children in households with a toilet are taller for their age. Household wealth raises the chance of having a toilet and, separately, buys better food and health care. Mother's schooling does both too. Caste and village location shape all of these at once. Each is a confounder of the toilet-height relationship. Some, like wealth, can be measured roughly. Others, such as a family's attention to hygiene, are hard to measure at all.
Adjusting for measured confounders helps. It cannot remove confounding by things that were never recorded, and a long list of controls in a table gives no information about what is missing from it.
WealthToiletChild height
ImpactMojoCausal Inference 101www.impactmojo.in
The direction of bias can be reasoned out in advance
When a confounder is left out of a comparison or regression, the estimate is biased by an amount that depends on two signs: how the confounder relates to the treatment, and how it relates to the outcome. You can usually reason about both before seeing any results, which tells you whether a reported effect is more likely too big or too small.
Confounder linked to treatmentConfounder linked to outcomeBias in naive estimateExample (Illustrative)
PositivelyPositivelyUpward, effect overstatedMotivated farmers adopt drip irrigation and also manage crops better
PositivelyNegativelyDownward, effect understatedSicker patients are more likely to get the new drug and more likely to die
NegativelyPositivelyDownward, effect understatedRicher households are less likely to use a ration shop and have better nutrition
NegativelyNegativelyUpward, effect overstatedRemote villages get fewer schools and also have lower learning for other reasons
Use this when reading a report: if every plausible omitted factor pushes the estimate up, a modest reported effect may be close to zero.
ImpactMojoCausal Inference 101www.impactmojo.in
Sometimes the outcome is choosing the treatment
Reverse causation means the arrow points from the outcome to the treatment. It is common wherever programmes respond to need, which in public policy is most of the time. Drought relief goes to districts with poor harvests. Police are posted where crime is high. Remedial teachers are sent to the weakest classes. In each case a simple correlation between the programme and the outcome has the wrong sign: relief is associated with low yields, police with crime.
Timing helps but is not decisive. Measuring the treatment before the outcome rules out the outcome causing the treatment directly, yet a third factor can still drive both, and anticipation can make the outcome move first.
Simultaneity
In markets, quantity and price are set together, so regressing one on the other recovers neither supply nor demand. The same holds for women's earnings and household bargaining power, or for school quality and enrolment. When two things cause each other, only variation that shifts one of them from outside, such as a rule or an instrument, separates the two directions. Section 8 returns to this.
ImpactMojoCausal Inference 101www.impactmojo.in
Attrition and measurement can create bias after assignment
Even a well-chosen comparison can be spoiled by what happens to the data afterwards. The two most common problems in South Asian field studies are people who cannot be found at follow-up and outcomes that are measured differently in the two groups.
Differential attrition
If successful migrants leave the village and cannot be traced, and the programme encouraged migration, the treated group loses its most successful members. The follow-up compares stayers with everyone. Report attrition by arm, test whether leavers differ, and use bounds (for example Lee bounds) when they do.
Measurement that differs by group
Participants in a hygiene programme learn which answers are expected and report more handwashing. Enumerators who know which villages were treated probe differently. Use outcomes that are hard to game (test scores, administrative records, observed behaviour), and keep enumerators blind to treatment status where possible.
Attrition and measurement bias apply to randomised trials as much as to observational studies. Random assignment protects only the moment of assignment.
ImpactMojoCausal Inference 101www.impactmojo.in
LaLonde's test: adding controls did not recover the experimental answer
Robert LaLonde took a United States job training programme that had been run as a randomised experiment, so the true effect was known. He then threw away the experimental control group and asked what the standard non-experimental methods of the day would have estimated, using comparison groups drawn from national surveys and the usual econometric adjustments.
Many of the methods failed to reproduce the experimental result, and they disagreed with each other. Specifications that looked reasonable gave answers far from the truth. ("Evaluating the Econometric Evaluations of Training Programs with Experimental Data", American Economic Review 76(4), 1986, pp. 604–620.)
Dehejia and Wahba (JASA 1999) reanalysed the same data with propensity score matching. Smith and Todd (Journal of Econometrics 125, 2005) found matching estimates on these data highly sensitive to the variables in the score and the sample used, and concluded that matching is a useful tool with no general claim to solve the evaluation problem. Section 9 picks this up.
Why it matters
A benchmark nobody usually has
In most evaluations there is no experiment to check against, so a wrong non-experimental estimate looks exactly as credible as a right one. LaLonde's study is valuable because it measured how wrong the usual tools could be.
ImpactMojoCausal Inference 101www.impactmojo.in
04
Section Four
Causal diagrams
ImpactMojoCausal Inference 101www.impactmojo.in
Draw your assumptions before you estimate anything
Directed acyclic graph (DAG)
A diagram of variables (boxes) joined by arrows, where an arrow from A to B means A is assumed to cause B, and no chain of arrows leads back to where it started. Missing arrows are assumptions too: they say one variable does not directly affect another.
Causal diagrams were developed by Judea Pearl and others in computer science and epidemiology. Economists came to them later, but they are now standard teaching because they turn a vague worry ("there might be confounding") into a precise question: which variables must be held fixed, and which must be left alone, to isolate one arrow.
A DAG does not estimate anything. It tells you what your assumptions imply about what to control for. Hernán and Robins's free book Causal Inference: What If (miguelhernan.org/whatifbook) teaches the method with health examples, and the free web tool DAGitty (dagitty.net) lets you draw a diagram and lists the adjustment sets it implies.
Drawing a diagram with a programme team takes an hour and surfaces disagreements about how the programme works that would otherwise appear only after the data are in.
ImpactMojoCausal Inference 101www.impactmojo.in
Forks, chains and colliders
ShapeArrowsName of middle variableAdjust for it?
ForkT ← Z → YConfounderYes: it opens a non-causal path
ChainT → M → YMediatorNo, if you want the total effect of T
ColliderT → K ← YColliderNo: adjusting opens a false path
Every larger diagram is made of these three shapes. A path between treatment and outcome carries association unless it is blocked. A fork or chain is blocked by adjusting for its middle variable. A collider is blocked as long as you leave it, and its consequences, alone. Adjusting for a collider unblocks it.
TreatmentConfounderOutcomeCollider
The common habit of controlling for every available variable is safe only for forks. For chains and colliders it introduces bias.
ImpactMojoCausal Inference 101www.impactmojo.in
Controlling for a mediator removes part of the effect
Illustrative. A cash transfer to mothers is meant to improve child height. Most of the effect is expected to work through more spending on food. An analyst regresses height on the transfer and, to be careful, adds household food expenditure as a control. The transfer's coefficient falls close to zero and the report concludes the transfer had no effect on height.
The conclusion is wrong. Holding food spending fixed asks what the transfer did for children whose food spending did not change, which removes the main channel by design. The coefficient is now an estimate of the effect through every other route, and even that is biased if anything unmeasured affects both food spending and height.
Rule: variables measured after treatment that could have been changed by it are not controls. Decompose mechanisms with methods built for mediation, and treat the results as weaker than the total effect.
TransferFood spendHeight
ImpactMojoCausal Inference 101www.impactmojo.in
The birth weight paradox: selecting on a collider
Among United States infants born in 1991, babies of mothers who smoked had higher risks of both low birth weight and infant death. But among low birth weight babies, those born to smokers had lower mortality (relative rate 0.79). For years this was read as a puzzle, even as a sign smoking protected small babies.
Hernández-Díaz, Schisterman and Hernán showed with causal diagrams that the pattern appears with no protective effect at all. Low birth weight is caused by smoking and also by other things, such as birth defects, that raise mortality. Among small babies, those whose smallness is not explained by smoking are more likely to be small for a deadlier reason. (American Journal of Epidemiology 164(11), 2006, pp. 1115–1120.)
SmokingDefectsLow weightDeath
Restricting a sample to one level of a collider is the same as adjusting for it.
ImpactMojoCausal Inference 101www.impactmojo.in
Samples selected on an outcome invite collider bias
Collider bias often arrives through who is in the dataset, with no control variable involved. Any sample defined by something the treatment and the outcome both affect is conditioned on a collider.
  • Surviving firms. Studying only enterprises still operating two years after a loan programme. Loans and good management both keep a firm alive, so among survivors borrowers may look less well managed.
  • Admitted students. Among students admitted to a selective college, entrance score and family income can look negatively related even if they are unrelated in the population.
  • Facility records. Studying only women who reached a hospital. Distance and complications both shape who arrives.
What to do
  • Draw how units got into the data, as an arrow in the diagram
  • Prefer population samples (NFHS, PLFS, a census frame) to programme or facility records when the question allows
  • Where only a selected sample exists, say so and reason about the direction of bias
  • Do not add a control just because it is available
ImpactMojoCausal Inference 101www.impactmojo.in
Block every backdoor path, open no new ones
A backdoor path is any path from the treatment to the outcome that begins with an arrow pointing into the treatment. Such paths carry association that is not the treatment's effect. Pearl's backdoor criterion says a set of variables is enough to adjust for if it blocks every backdoor path and contains no variable caused by the treatment.
In plain terms: control for common causes, leave alone anything the treatment might have changed, and do not condition on colliders. If an unmeasured variable sits on a backdoor path and nothing else blocks it, no regression with observed controls identifies the effect, and you need a design from Sections 5 to 8.
01
LIST: every variable that could affect treatment or outcome
→
02
DRAW: arrows the team believes exist, with reasons
→
03
TRACE: all paths from treatment to outcome
→
04
MARK: backdoor paths and colliders
→
05
CHOOSE: a measured set that blocks the backdoors
→
06
ADMIT: paths no measured set can block
ImpactMojoCausal Inference 101www.impactmojo.in
Drawing the diagram for Bihar's bicycles for girls
Bihar's Cycle programme gave girls who continued to secondary school a bicycle (Muralidharan and Prakash 2017). Draw the question "did the scheme raise girls' secondary enrolment?" before choosing a method.
  • Treatment: being in a cohort of girls eligible for the bicycle.
  • Outcome: age-appropriate enrolment in secondary school.
  • Common causes over time: rising incomes, new roads, more secondary schools, other state schemes, all of which raised enrolment for boys too.
  • Common causes across places: Bihar differs from neighbouring states in many ways that also shape schooling.
  • Mediator: distance to school becoming manageable. Do not control for it.
What the diagram points to
Neither a before-after comparison (confounded by time) nor a Bihar-versus-other-states comparison (confounded by place) works alone. Comparing girls with boys, before and after, in Bihar and in a neighbouring state, blocks both sets of backdoor paths. That is the triple difference Muralidharan and Prakash used, covered in Section 6.
ImpactMojoCausal Inference 101www.impactmojo.in
05
Section Five
Randomisation
ImpactMojoCausal Inference 101www.impactmojo.in
A lottery makes the two groups alike in expectation
When assignment is decided by a lottery, nothing about a unit can influence whether it is treated. Motivation, wealth, caste, distance, the officer's preferences: none of them can predict the coin. So the treatment and control groups have the same expected values of every characteristic, measured or not, and the selection bias term from Section 2 is zero in expectation.
"In expectation" matters. In any one draw the groups differ by chance, and with few units they can differ a lot. That chance difference is what the standard error measures, and it shrinks as the number of randomised units grows.
What randomisation buys
  • Balance on unmeasured characteristics, which no other design can promise
  • A simple estimator: the difference in mean outcomes
  • Transparent inference: the uncertainty comes from a known random process
  • Credibility with sceptical audiences, including finance ministries
It does not buy external validity, freedom from attrition, or a correct theory of why the programme worked.
ImpactMojoCausal Inference 101www.impactmojo.in
Checking that randomisation did its job
Illustrative baseline balance table from a school-level trial with 120 schools randomised to a remedial reading programme. A balance table reports baseline means by arm and the difference. It tests the implementation of the lottery; it does not test whether the design is valid, which randomisation already guarantees.
Baseline characteristicTreatment (60 schools)Control (60 schools)Differencep-value
Grade 3 reading score (standardised)0.02−0.010.030.71
Pupils enrolled per school14213840.64
Share of pupils from SC or ST households0.310.34−0.030.42
Teacher attendance on unannounced visit0.780.760.020.58
Distance to block headquarters (km)14.112.91.20.33
With many characteristics, about one in twenty will differ at the 5% level by chance. Decide in advance which baseline variables you will control for, and stratify the lottery on the most important ones (for example block, or baseline score) so the arms are balanced on them by construction.
ImpactMojoCausal Inference 101www.impactmojo.in
Individuals, schools, villages or districts
Unit randomisedWhen it fitsCost to the study
IndividualTreatment delivered person by person with little spillover: a scholarship, an SMS reminderCheapest in sample size; spillovers inside households or villages contaminate the control
HouseholdTransfers, asset grants, counsellingEffects on neighbours still possible
School or clinicAnything a whole institution delivers: a teaching method, staffingOutcomes within a school are correlated, so more pupils add less information
Village or wardInfrastructure, community mobilisation, local marketsNeed many villages; tens are rarely enough
Block, district or subdistrictAdministrative reforms rolled out by governmentFew units, so power is low unless the effect is large
Randomise at the level where the treatment is delivered and where spillovers stop. The cost is clustering: pupils in the same school share a teacher, so 40 pupils from one school carry much less information than 40 pupils from 40 schools. Standard errors must be clustered at the level of randomisation, and the power calculation must use the number of clusters. Impact Evaluation 101 works through the arithmetic.
ImpactMojoCausal Inference 101www.impactmojo.in
Balsakhis and computers in Vadodara and Mumbai
0.28 SD
rise in average test scores in schools with the remedial balsakhi programme
Banerjee, Cole, Duflo & Linden, QJE 2007
0.47 SD
rise in maths scores from computer-assisted learning
Banerjee, Cole, Duflo & Linden, QJE 2007
~0.10 SD
what remained one year after the programmes ended
Banerjee, Cole, Duflo & Linden, QJE 2007
The balsakhi programme hired young women from the community to teach children who were falling behind in basic literacy and numeracy. Schools were randomly assigned. Most of the gain came from children at the bottom of the distribution, the group the programme targeted. ("Remedying Education: Evidence from Two Randomized Experiments in India", Quarterly Journal of Economics 122(3): 1235–1264.)
Two lessons beyond the headline. A cheap input, a community teacher with little training, produced effects comparable to costlier ones. And gains faded once the programme stopped, which is a finding about persistence that only a follow-up could reveal.
ImpactMojoCausal Inference 101www.impactmojo.in
Randomising inside government in Andhra Pradesh
Teacher performance pay
Muralidharan and Sundararaman randomised a teacher performance pay programme across a representative sample of government rural primary schools in Andhra Pradesh. After two years, pupils in incentive schools scored 0.27 standard deviations higher in maths and 0.17 in language than pupils in control schools, with no evidence of adverse consequences. (Journal of Political Economy 119(1), 2011, pp. 39–77.)
Biometric Smartcards
Muralidharan, Niehaus and Sukhtankar evaluated biometrically authenticated payments for the employment guarantee and social pensions by randomising the rollout over 157 subdistricts covering 19 million people. Payments became faster, more predictable and less corrupt without reducing access. (American Economic Review 106(10), 2016, pp. 2895–2929.)
The Smartcards study shows a practical route to randomisation inside a government programme: the state could not roll out everywhere at once, so the order of rollout was randomised. Units waiting their turn served as the control. A phased rollout that has to happen anyway is often the cheapest and most ethical trial available.
ImpactMojoCausal Inference 101www.impactmojo.in
When you cannot assign the treatment, assign the offer
Mindspark lottery, urban India
Muralidharan, Singh and Ganimian studied a personalised, technology-aided after-school programme for middle-school pupils. Free places were allocated by lottery. Lottery winners scored 0.37 SD higher in maths and 0.23 SD higher in Hindi after 4.5 months. Gains were similar in absolute terms for all pupils, and much larger relative to their starting point for weaker ones. (AER 109(4), 2019, pp. 1426–1460.)
Migration incentive, Bangladesh
Bryan, Chowdhury and Mobarak randomly offered households in rural Bangladesh an incentive of US$8.50 to temporarily out-migrate during the pre-harvest lean season. The offer induced 22% of households to send a seasonal migrant, raised consumption at home, and treated households were 8–10 percentage points more likely to migrate again 1 and 3 years later. (Econometrica 82(5), 2014, pp. 1671–1748.)
Both designs randomise an offer. The comparison of winners with losers is the intention-to-treat effect. Section 8 shows how to turn it into the effect of actually attending or migrating.
ImpactMojoCausal Inference 101www.impactmojo.in
What South Asian trials found, on one scale
Effects on test scores in standard deviations, randomised studies
Banerjee et al., QJE 2007; Muralidharan & Sundararaman, JPE 2011; Muralidharan, Singh & Ganimian, AER 2019; Andrabi, Das & Khwaja, AER 2017
The studies differ in duration (4.5 months to two years), grade and test, so the bars are not a league table. Andrabi, Das and Khwaja randomised report cards across half of the sample villages in Pakistan; scores rose 0.11 SD, private school fees fell 17% and primary enrolment rose 4.5%.
ImpactMojoCausal Inference 101www.impactmojo.in
Randomisation protects assignment, nothing after it
ThreatWhat happensDefence
Non-complianceSome assigned units refuse; some control units find the treatment elsewhereReport intention to treat; estimate effect on compliers by IV (Section 8)
AttritionUnits lost at follow-up differ by armTrack hard; report rates by arm; bound the estimate
SpilloversControl units are affected by treated neighboursRandomise larger clusters; measure spillovers with buffer or partial-treatment designs
Hawthorne and John Henry effectsTreated units change because they are observed; controls competeMeasure outcomes unobtrusively; give controls a placebo contact
Specification searchingMany outcomes and subgroups tested, the significant ones reportedPre-register; write a pre-analysis plan; adjust for multiple outcomes
Implementation failureThe treatment was not delivered as designedMonitor delivery; report what was actually delivered
Banerjee, Duflo, Glennerster and Kinnan's Hyderabad microfinance trial (AEJ: Applied 7(1), 2015) shows one such issue plainly: the lender entered 52 randomly chosen neighbourhoods, but take-up of microcredit rose only 8.4 percentage points, and two years later the control areas had also gained access to microcredit.
ImpactMojoCausal Inference 101www.impactmojo.in
Running a fair trial in South Asia, as of October 2026
When a lottery is fair
  • Demand exceeds supply, so someone must be left out anyway
  • There is real uncertainty about whether the programme helps (equipoise)
  • The rollout is phased, so the control group is treated later
  • No one is denied an entitlement they hold in law
  • Participants give informed consent, and an ethics committee has approved the protocol
Registration and data
  • Register social science trials before baseline on the American Economic Association's RCT Registry; register clinical and health trials in India on the Clinical Trials Registry of India (CTRI)
  • Write a pre-analysis plan naming primary outcomes and the main specification
  • Personal data of participants in India falls under the Digital Personal Data Protection Act 2023 and the DPDP Rules 2025, whose duties apply from 13 May 2027. From that date section 17(2)(b) exempts processing for research and statistical purposes that meets its conditions, including that no decision specific to the person is taken. Plan studies running past that date to comply
Research Ethics 101 and Data Protection & the DPDP Act 101 cover consent, ethics review and data handling in detail.
ImpactMojoCausal Inference 101www.impactmojo.in
06
Section Six
Natural experiments and difference-in-differences
ImpactMojoCausal Inference 101www.impactmojo.in
When a rule, a lottery or history does the randomising
Natural experiment
A situation in which treatment was assigned by something outside the units' control, and plausibly unrelated to their outcomes, without the researcher designing it: a constitutional rotation, an eligibility cutoff, a phased rollout, a border, a date of birth.
Reserved seats for women pradhans
In West Bengal in 1998, one third of village council pradhan posts were randomly selected to be reserved for women. Chattopadhyay and Duflo used this randomised policy, with data on 265 village councils in West Bengal and Rajasthan. In West Bengal, women leaders invested more in water, fuel and roads, the goods rural women had asked for, and men more in education. (Econometrica 72(5), 2004; NBER Working Paper 8615.)
What makes it credible
Reservation was assigned at random, unrelated to a village's needs or politics, so reserved and unreserved councils were comparable. The researcher's task shifts from building a comparison to documenting that the rule really was followed and could not be manipulated. Always ask who could have influenced the "natural" assignment.
ImpactMojoCausal Inference 101www.impactmojo.in
Difference-in-differences in a two-by-two table
Illustrative. A state starts a free school bus in District A in 2024. District B gets nothing. Girls' attendance rates in both districts, before and after:
2023 (before)2025 (after)Change
District A (bus)71%82%+11 points
District B (no bus)64%70%+6 points
Difference A minus B7 points12 points+5 points
The first difference (after minus before, within A) removes everything fixed about District A. The second (A's change minus B's change) removes everything that changed over time for both districts, such as a statewide fee waiver. The +5 points is the estimate.
Neither single comparison would do. Before-after in A says +11, crediting the bus with the statewide trend. After-only, A against B, says +12, crediting the bus with A's head start. The DiD keeps only the part of A's change that B did not share.
ImpactMojoCausal Inference 101www.impactmojo.in
The one assumption: the same trend without the programme
Girls' attendance (%), Illustrative
Illustrative; dashed line is the assumed counterfactual for District A
DiD assumes that, without the programme, the treated group's outcome would have moved in parallel with the comparison group's. The levels can differ; the trends must match. The assumption is about a counterfactual, so it can never be proved.
What can be checked is whether the trends were parallel before treatment. Several pre-programme years moving together make the assumption more believable. Diverging pre-trends are a warning that the groups were already on different paths.
Parallel pre-trends are supporting evidence and fall short of proof. A shock that hits only the treated area at the same time as the programme still breaks the design.
ImpactMojoCausal Inference 101www.impactmojo.in
Two natural experiments every evaluator should know
Minimum wages, New Jersey, 1992
On 1 April 1992 New Jersey raised its minimum wage from $4.25 to $5.05 an hour. Card and Krueger surveyed 410 fast-food restaurants in New Jersey and eastern Pennsylvania, where the minimum wage did not change, before and after. Comparing employment growth across the border, they found no indication that the rise reduced employment. (American Economic Review 84(4), 1994.) The debate it started ran for years and sharpened how DiD studies are checked.
School construction, Indonesia, 1973–78
Global canon from Indonesia. Duflo combined differences across regions in the number of INPRES schools built with differences across birth cohorts in exposure. Each school built per 1,000 children raised education by 0.12 to 0.19 years and wages by 1.5% to 2.7%, implying returns to schooling of 6.8% to 10.6%. (American Economic Review 91(4), 2001, pp. 795–813.)
Both use the same logic: compare the change for those exposed with the change for those not exposed. The INPRES study shows how age cohorts can serve as the "before" and "after".
ImpactMojoCausal Inference 101www.impactmojo.in
The employment guarantee's phased rollout as a natural experiment
The National Rural Employment Guarantee Scheme came into force in February 2006 in the first 200 districts, reached a further set of districts in April 2007, and covered all remaining districts from 1 April 2008 (Zimmermann, IZA Discussion Paper 6858, 2012; Ministry of Rural Development). The phases were set by a backwardness ranking, so early districts were poorer.
Imbert and Papp compared trends in districts that got the programme earlier with those that got it later. Public works hiring crowded out private work and raised private sector wages, and they calculate that the wage gains for poor households were large relative to the gains of participants alone. (AEJ: Applied Economics 7(2), 2015, pp. 233–263.)
As of October 2026
The Mahatma Gandhi National Rural Employment Guarantee Act 2005 was repealed from 1 July 2026 and replaced by the Viksit Bharat G RAM G Act 2025, with a guarantee of 125 days. Studies of the earlier scheme describe that scheme. Any before-after comparison spanning July 2026 will mix the new law's effect with everything else that changed that year, which is exactly the problem a comparison group is needed to solve.
ImpactMojoCausal Inference 101www.impactmojo.in
Bihar's bicycles: boys and Jharkhand as two comparison groups
+32%
girls' age-appropriate enrolment in secondary school, cohorts exposed to the Cycle programme
Muralidharan & Prakash, AEJ: Applied 2017
−40%
reduction in the corresponding gender gap
Muralidharan & Prakash, AEJ: Applied 2017
+18%
girls appearing for the secondary school certificate exam
Muralidharan & Prakash, AEJ: Applied 2017
A simple DiD comparing Bihar girls before and after would credit the scheme with everything that raised enrolment in Bihar at the time. Comparing Bihar girls with Bihar boys removes Bihar-wide changes, but boys and girls might have been on different trends anyway. So the authors also used the neighbouring state of Jharkhand, carved out of Bihar on 15 November 2000, and computed the girl-boy gap's change in Bihar minus its change in Jharkhand. ("Cycling to School", AEJ: Applied Economics 9(3), 2017, pp. 321–350.)
The increases took place mostly in villages farther from a secondary school, the pattern you would expect if the bicycle worked by cutting the time and safety cost of getting there. A mechanism that fits the effect's shape makes the claim more believable.
ImpactMojoCausal Inference 101www.impactmojo.in
When units start at different times, the usual regression can mislead
Many Indian programmes roll out district by district over years. The usual estimator, a regression with unit and time fixed effects (two-way fixed effects), was long assumed to give a sensible average. Work published from 2020 onward showed that, when effects differ across groups or grow over time, it can compare newly treated units against units treated earlier, giving some comparisons negative weight.
PaperMain pointWhere
Goodman-Bacon (2021)The two-way fixed effects estimate is a weighted average of all two-by-two DiDs, including early-versus-late comparisonsJournal of Econometrics 225(2): 254–277
de Chaisemartin & D'Haultfœuille (2020)Weights can be negative; the coefficient can be negative while every group's effect is positiveAmerican Economic Review 110(9): 2964–2996
Callaway & Sant'Anna (2021)Estimate effects for each adoption cohort and period, then aggregate as you chooseJournal of Econometrics 225(2): 200–230
Sun & Abraham (2021)Event-study leads and lags are contaminated in the same way; an interaction-weighted fixJournal of Econometrics 225(2): 175–199
When reading a staggered-rollout study published before about 2021, ask whether its results hold with one of these estimators. Many authors have since re-run their own work this way.
ImpactMojoCausal Inference 101www.impactmojo.in
Event studies, placebos and the questions to ask
Event study
Estimate the treated-minus-comparison gap for each year before and after the start. Before the start the estimates should be close to zero with no trend. After, they show how the effect builds or fades. A plot of these coefficients is the single most informative figure in a DiD paper; read it before the main table.
Placebo tests
Pretend the programme started two years earlier and re-estimate: the effect should be zero. Or use an outcome the programme cannot plausibly affect. A non-zero placebo means something other than the programme differs between the groups.
  • Who chose which units were treated first, and on what basis?
  • Were there several years of data before treatment, and do they move together?
  • Did anything else start in the treated units at the same time: a new collector, another scheme, a road?
  • Did people move into or out of treated areas because of the programme?
  • If rollout was staggered, was one of the newer staggered-timing estimators used?
  • Are standard errors clustered at the level treatment was assigned (state, district)?
ImpactMojoCausal Inference 101www.impactmojo.in
07
Section Seven
Regression discontinuity
ImpactMojoCausal Inference 101www.impactmojo.in
A cutoff rule creates a local experiment
Regression discontinuity (RD)
A design for programmes that assign treatment by a rule based on a score (the running variable) and a cutoff. Units just above and just below the cutoff are nearly identical, so a jump in outcomes exactly at the cutoff estimates the effect of treatment for units near it.
Indian administration is full of cutoffs: population thresholds for roads and schools, marks for admission and scholarships, land size for subsidies, backwardness scores for district programmes, income for ration cards, and vote shares that decide who wins a seat. Each is a potential RD.
The logic: a village of 990 people and one of 1,010 differ in nothing that matters except which side of a 1,000 rule they fall on. If outcomes jump at 1,000 and nowhere else, the rule's treatment is the obvious explanation. RD is often called the most credible design short of a lottery.
The price is that the estimate applies to units near the cutoff. The effect on a village of 300 people may be quite different.
ImpactMojoCausal Inference 101www.impactmojo.in
When the rule is followed exactly, and when it is not
Sharp RDFuzzy RD
RuleEveryone above the cutoff is treated; no one below isCrossing the cutoff raises the chance of treatment but not from 0 to 1
Example (Illustrative)A scholarship paid automatically to every student scoring 80% or moreA road programme that prioritises villages above a population size, though some below get roads and some above do not
EstimateJump in outcome at the cutoffJump in outcome divided by jump in treatment probability
Who it describesUnits at the cutoffUnits at the cutoff whose treatment was changed by the rule (compliers)
Close cousinA simple comparison of means near the lineInstrumental variables, with the cutoff as instrument
Most government rules are fuzzy, because officers use discretion, records are wrong, and programmes are expanded over time. A fuzzy RD is still credible, but it needs the jump in take-up at the cutoff to be large, and the estimate is a ratio of two jumps, so it is noisier than a sharp one.
ImpactMojoCausal Inference 101www.impactmojo.in
Rural roads and population thresholds: PMGSY
The Pradhan Mantri Gram Sadak Yojana (PMGSY) is India's national rural road programme. Its rules prioritised villages by population, with eligibility thresholds at 500 and 1,000 people. Asher and Novosad used those thresholds in a fuzzy regression discontinuity design, combined with household and firm census microdata.
Four years after a road was built, the main effect was to move workers out of agriculture. There were no major changes in agricultural outcomes, income or assets, and employment in village firms expanded only slightly. ("Rural Roads and Local Economic Development", American Economic Review 110(3), 2020, pp. 797–823.)
Why the design convinces
  • The programme is large: the paper puts India's national rural road construction programme at $40 billion
  • The rule was set by the programme, so a village's own choices did not decide its priority
  • Road probability jumps at the threshold, which gives a usable first stage
  • The finding is a null for income and assets, which a reader can weigh against the claims made for roads
ImpactMojoCausal Inference 101www.impactmojo.in
The employment guarantee's backwardness ranking
Districts entered the employment guarantee in phases based on a Planning Commission index of backwardness built from data of the early to mid-1990s. Districts were ranked within each state, and the poorest received the programme first. Laura Zimmermann used the rank cutoff between phases as a regression discontinuity. Rank data were available for 447 of 618 districts.
She found that private sector wages rose substantially for women but not for men, concentrated in the main agricultural season, with little evidence of lower private employment. (IZA Discussion Paper 6858, 2012.)
Two designs, one scheme
Imbert and Papp's DiD and Zimmermann's RD study the same programme with different assumptions. DiD needs early and late districts to have been on parallel trends; RD needs districts on either side of the rank cutoff to be comparable and the ranking not to have been manipulated. Both find that the programme raised private wages. When designs with different weaknesses agree, the finding is stronger than either alone.
Zimmermann argues manipulation was unlikely because the index used data predating the scheme.
ImpactMojoCausal Inference 101www.impactmojo.in
What an RD plot shows
Share of workers outside agriculture by village population, Illustrative
Illustrative binned means; not data from any study
A good RD paper shows outcomes averaged in bins of the running variable, with separate fitted lines on each side. The effect is the vertical gap where the lines meet the cutoff.
  • The jump should be visible to the eye in the plot as well as in the regression table
  • Baseline characteristics plotted the same way should show no jump
  • Results should hold for narrower and wider bandwidths
  • Placebo cutoffs (say at 900 or 1,100) should show nothing
ImpactMojoCausal Inference 101www.impactmojo.in
Sorting at the cutoff and the choice of bandwidth
Manipulation of the running variable
If people can push their score across the line, units just above differ from those just below in whatever made them able to push. Examiners may round marks of 39 up to the pass mark of 40; households may under-report income to qualify for a card. The sign is bunching: too many units just above the cutoff. McCrary's density test checks for it (Journal of Econometrics 142(2), 2008, pp. 698–714).
Bandwidth and functional form
A narrow window around the cutoff keeps the comparison clean but leaves few units; a wide one adds units that are less alike. High-order polynomials fitted across the whole range can create spurious jumps. Current practice uses local linear regression with a data-driven bandwidth and bias-corrected confidence intervals (Calonico, Cattaneo and Titiunik, Econometrica 82(6), 2014), available as free software for R and Stata.
Also check whether other programmes use the same cutoff. If a population of 1,000 also triggered a school or a health sub-centre, the jump measures all of them together.
ImpactMojoCausal Inference 101www.impactmojo.in
Who wins a narrow race is close to a coin toss
Elections give a running variable: the winning margin. Among constituencies decided by a fraction of a percentage point, which candidate won is close to random. Comparing places where a type of candidate barely won with places where that type barely lost estimates the effect of having that type of representative.
Asher and Novosad used close elections in India to ask whether being represented by a member of the ruling party affects local economic activity. Favouritism led to higher private sector employment, higher share prices of firms and more output measured by night lights, and they present evidence that politicians work mainly through control over how regulation is implemented. (AEJ: Applied Economics 9(1), 2017, pp. 229–273.)
Practitioner use
A design you can borrow
Close-election RD needs only results by constituency, which the Election Commission of India publishes, and an outcome measured for the same constituencies, such as night lights, firm counts or programme spending. Check that close races are not won disproportionately by incumbents or the ruling party, which would break the coin-toss logic.
ImpactMojoCausal Inference 101www.impactmojo.in
08
Section Eight
Instrumental variables
ImpactMojoCausal Inference 101www.impactmojo.in
An instrument moves the treatment and touches nothing else
When people choose whether to be treated, find something that shifts the choice for reasons unrelated to the outcome. That something is an instrument. The variation in treatment it causes is as good as random, and the effect of the treatment can be estimated from that variation alone. A lottery for an offer is the cleanest instrument there is.
ConditionIn plain wordsCan you test it?
RelevanceThe instrument changes the treatment, substantiallyYes: look at the first stage and its F-statistic
IndependenceThe instrument is as good as randomly assigned, unrelated to anything else that affects the outcomePartly: check balance on observed characteristics
Exclusion restrictionThe instrument affects the outcome only through the treatment, by no other routeNo: it must be argued from knowledge of the setting
MonotonicityThe instrument pushes everyone the same way (no one is less likely to take the treatment because of it)Rarely; argued from the setting
Most failed instrumental variables studies fail on the exclusion restriction, which is the one condition data cannot check.
ImpactMojoCausal Inference 101www.impactmojo.in
State it in one sentence, then try to break it
The exclusion restriction: the instrument has no effect on the outcome except through its effect on the treatment.
Write the sentence out for the study in front of you, replacing the words. "Winning the Mindspark lottery affects test scores only by changing whether the child attends Mindspark." Then list every other way the instrument could reach the outcome. Could winners' parents have changed tuition or effort because they won? If so, the restriction fails and the estimate absorbs that route too.
Lotteries usually pass, because the only thing that differs between winners and losers is the offer. Instruments taken from geography, weather or history usually face many rival routes.
Rainfall: a cautionary instrument
Illustrative. An analyst uses district rainfall as an instrument for household income to estimate the effect of income on children's schooling. Rainfall does change income. It also changes disease, the demand for child labour on farms, migration and whether roads to school are passable. Each is a route from rainfall to schooling that bypasses income, so the exclusion restriction fails. A variable that affects many things is a poor instrument for any one of them.
ImpactMojoCausal Inference 101www.impactmojo.in
Dams in India: river gradient as an instrument
Large dams are not built at random. They go where the state expects benefits, where politics favours them, and where land can be acquired. Comparing districts with and without dams mixes these factors with the dams' effects.
Duflo and Pande used the fact that river gradient affects how suitable a district is for a dam. Gradient therefore predicts where dams were built. The exclusion restriction is that, once the authors' other controls are included, gradient affects agricultural output and poverty only through dams.
Results: downstream districts gained agricultural production and lower vulnerability to rainfall shocks; the district where the dam stood saw no significant production gain, more volatility, and higher rural poverty. ("Dams", Quarterly Journal of Economics 122(2), 2007, pp. 601–646.)
Try to break it
  • Does gradient affect farming directly, through soil or drainage?
  • Does it affect road building or settlement patterns?
  • Read the paper's own tests of these routes, and ask what else slope could change in your setting
ImpactMojoCausal Inference 101www.impactmojo.in
An instrument estimates the effect for those it moves
Imbens and Angrist (Econometrica 62(2), 1994) and Angrist, Imbens and Rubin (JASA 91(434), 1996) showed what an IV estimate means when effects differ across people. It is the average effect for compliers: units whose treatment was changed by the instrument. This is the local average treatment effect, or LATE.
TypeIf offeredIf not offeredIn a lottery for tuition places (Illustrative)
Always-takersTreatedTreatedWould have paid for similar classes anyway
Never-takersUntreatedUntreatedWould not attend even if free
CompliersTreatedUntreatedAttend only when the place is free; the LATE is their effect
DefiersUntreatedTreatedAssumed not to exist (monotonicity)
LATE is a real effect for a real group, but a policy that reaches different people, such as making attendance compulsory, may have a different average effect.
ImpactMojoCausal Inference 101www.impactmojo.in
From the effect of the offer to the effect of attending
LATE = effect of the instrument on the outcome ÷ effect of the instrument on the treatment
Illustrative arithmetic. A lottery offers free remedial classes. Winners' scores are 0.12 SD higher than losers' (the intention-to-treat effect). 60% of winners attend and 0% of losers do, so the offer raised attendance by 0.6. The effect of attending, for compliers, is 0.12 ÷ 0.6 = 0.20 SD.
The division assumes the offer helped only by causing attendance (exclusion) and that winners who did not attend gained nothing.
In the Mindspark study
Muralidharan, Singh and Ganimian report the lottery effect (0.37 SD in maths and 0.23 SD in Hindi over 4.5 months) and use the lottery as an instrument for days attended. Their IV estimates imply that attending for 90 days would raise maths and Hindi scores by 0.6 SD and 0.39 SD. The lottery effect answers "what does offering a place do?"; the IV estimate answers "what does using the programme do?" (AER 2019).
Policy reports often quote the larger IV number. If take-up in a scaled programme is low, the offer effect is the relevant one.
ImpactMojoCausal Inference 101www.impactmojo.in
A weak first stage magnifies every small violation
The IV estimate divides by the first stage. If the instrument moves the treatment only a little, a small direct effect of the instrument on the outcome (a slight exclusion failure) is divided by a small number and becomes a large bias. Weak instruments also make conventional confidence intervals far too narrow.
For decades practitioners pre-tested with a rule of thumb of a first-stage F-statistic above 10, drawn from Staiger and Stock (Econometrica 65(3), 1997) and Stock and Yogo (2005). Lee, McCrary, Moreira and Porter showed that conventional t-tests can be badly distorted even then. Across 61 American Economic Review papers, for a quarter of specifications the corrected standard errors were at least 49% larger than the usual ones at the 5% level. (AER 112(10), 2022, pp. 3260–3290.)
What to look for
  • The first-stage estimate and its F-statistic reported prominently
  • The reduced form (instrument on outcome), which needs no exclusion argument to interpret as the effect of the instrument
  • Confidence intervals that stay valid with weak instruments, such as Anderson-Rubin or the tF adjustment
  • An IV estimate many times larger than the simple comparison, which often signals a weak or invalid instrument
ImpactMojoCausal Inference 101www.impactmojo.in
Five questions for any instrumental variables claim
01
NAME: what is the instrument, and why does it move the treatment?
→
02
FIRST STAGE: how strongly, and is the F-statistic reported?
→
03
EXCLUSION: every other route from instrument to outcome, ruled out how?
→
04
COMPLIERS: who did the instrument move, and do they resemble the policy's target group?
→
05
COMPARE: how does the IV estimate relate to the simple and the reduced-form estimates?
Instruments that usually convince
Randomised offers or encouragements; lotteries for places; fuzzy cutoff rules; random assignment of cases to judges or officers with different tendencies.
Instruments that need hard scrutiny
Rainfall, distance, historical events, lagged values of the treatment, and averages of other people's choices in the same village. Each can reach the outcome by many routes.
If the report cannot state its exclusion restriction in one sentence that a programme officer could check against what they know of the place, treat the estimate as a correlation.
ImpactMojoCausal Inference 101www.impactmojo.in
09
Section Nine
Matching and synthetic control
ImpactMojoCausal Inference 101www.impactmojo.in
Compare each participant with a non-participant who looked the same
Selection on observables (conditional independence)
The assumption that, among units with the same values of the measured characteristics, treatment is as good as random. Matching, regression adjustment and weighting all rest on it.
Matching pairs each treated unit with one or more untreated units that had the same or similar baseline characteristics: age, caste category, landholding, education, village. The effect is the average outcome difference within matched pairs. It makes the comparison transparent: you can see exactly who stands in for whom, and check whether the matched groups are alike.
Its weakness is the assumption. Matching removes differences in what you measured. It does nothing about motivation, information, social networks or an officer's judgement, if those drove participation and are missing from the data. A matched comparison that looks perfectly balanced on every column can still be confounded by something that has no column.
Matching is the right tool when you know how selection happened and measured it, for example when officers applied a written eligibility checklist whose items are all in the data.
ImpactMojoCausal Inference 101www.impactmojo.in
Reducing many characteristics to one probability
With many characteristics, exact matches are impossible. Rosenbaum and Rubin showed that matching on the propensity score, the probability of treatment given the observed characteristics, balances those characteristics as well as matching on all of them would. ("The central role of the propensity score in observational studies for causal effects", Biometrika 70(1), 1983, pp. 41–55.)
The score is a convenience for balancing what you measured. It adds no protection against what you did not measure, and a well-fitting model of participation says nothing about whether the assumption holds.
01
MODEL: participation as a function of pre-treatment characteristics
→
02
SCORE: each unit's predicted probability of participating
→
03
TRIM: drop units with no counterpart (no common support)
→
04
MATCH or WEIGHT: pair or reweight units with similar scores
→
05
CHECK: balance on each characteristic after matching
→
06
ESTIMATE: outcome difference, with sensitivity analysis
ImpactMojoCausal Inference 101www.impactmojo.in
Matching only works where both groups exist
If almost every landless Dalit household in a block joined a programme and almost no landed household did, there is nobody to match the participants with. Estimates for regions of the data with no overlap are extrapolations from a model, whatever the software reports. Good practice is to show the distributions of the score in both groups and say how many units were dropped.
DiagnosticWhat good looks likeWhat it reveals if poor
Overlap of propensity scoresBoth groups spread across the same rangeThe comparison rests on a few untreated units or on extrapolation
Standardised differences after matchingBelow about 0.1 for each characteristicMatching did not balance what it was meant to
Number of units discardedReported, with who they wereThe estimate describes a subgroup the reader did not expect
Sensitivity analysisHow strong an unobserved confounder would have to be to erase the effectFragile results that a modest omitted factor would overturn
Report the effect for the population that was actually matched, and say who that is.
ImpactMojoCausal Inference 101www.impactmojo.in
Combine matching with differencing, then test sensitivity
Matching plus difference-in-differences
If baseline and follow-up data exist, match on pre-programme characteristics and then compare changes in outcomes. Differencing removes unmeasured differences that are fixed over time, such as a household's location or a village's history, which matching alone leaves in place. On the LaLonde data, Smith and Todd (Journal of Econometrics 2005) found the difference-in-differences matching estimator performed best among those they studied, because it removed time-invariant sources of bias such as geographic mismatch between participants and non-participants.
Sensitivity analysis
Ask how strong an unmeasured confounder would have to be to change the conclusion. Cinelli and Hazlett (Journal of the Royal Statistical Society Series B 82(1), 2020) give measures that can be computed from standard regression output, including the minimum strength of association hidden confounding would need, with both treatment and outcome, to overturn the result. Compare that strength with the strongest confounder you did measure, such as landholding or education.
A matching study that reports neither a pre-programme comparison nor a sensitivity analysis is asking the reader to accept its main assumption on trust.
Lim and colleagues' JSY evaluation used three approaches (matching, with-versus-without and difference-in-differences) for the same reason: agreement across methods with different weaknesses is more informative than any one estimate.
ImpactMojoCausal Inference 101www.impactmojo.in
Why matching and DiD could disagree about newborn deaths
Lim et al., Lancet 2010Powell-Jackson et al., J. Health Econ. 2015
Variation usedIndividual women: who did and did not receive JSY paymentDistricts: differences in how intensively JSY was implemented
Main designMatching of women who received JSY payment to similar women who did not; also with-versus-without and DiDDifference-in-differences on how intensively districts implemented JSY
ComparisonSimilar women in the same period who did not receive paymentDistricts with lower implementation, before and after
Assumption that must holdReceiving payment is as good as random among women with the same observed characteristicsHigh- and low-implementation districts would have had parallel trends
Use of maternity careAntenatal care and facility births increasedUptake of maternity services increased
Neonatal mortalityReduction of 2.3 per 1,000 live births (matching)No strong evidence of a reduction
Women who received JSY payment had to give birth in a facility, and women who reach a facility may differ in health and access in ways surveys do not record. That is the opening for selection in the matching estimate. The DiD avoids individual selection but relies on district trends. Both papers are careful; the disagreement is a lesson in how assumptions shape answers.
ImpactMojoCausal Inference 101www.impactmojo.in
Building a comparison unit from a weighted mix of others
Synthetic control
For a single treated unit (a state, a country), a weighted average of untreated units chosen so that it tracks the treated unit's outcome closely in the years before treatment. Its path after treatment stands in for the treated unit's counterfactual.
The Basque Country
Abadie and Gardeazabal built a synthetic Basque Country from other Spanish regions. After terrorism began in the late 1960s, per capita GDP declined about 10 percentage points relative to the synthetic control. (American Economic Review 93(1), 2003, pp. 113–132.)
California's tobacco programme
Abadie, Diamond and Hainmueller formalised the method and applied it to California's tobacco control programme, comparing California with a weighted mix of other US states. (Journal of the American Statistical Association 105(490), 2010, pp. 493–505.)
Synthetic control suits the policy changes Indian states make one at a time, such as a state-specific prohibition, a new welfare scheme, or a change in land law, where there is one treated state and many possible comparisons.
ImpactMojoCausal Inference 101www.impactmojo.in
Choosing donors and testing with placebos
Illustrative. Suppose one state introduces a monthly cash transfer to women in 2024 and an analyst wants its effect on female labour force participation, measured each year in the Periodic Labour Force Survey. The donor pool is the other states. The method picks weights (say 40% one neighbouring state, 35% another, 25% a third) so the weighted average matches the treated state's participation rate for every pre-2024 year.
If the fit before 2024 is close, the gap after 2024 is the estimate. If the fit is poor, no estimate should be reported.
How inference works
  • Re-run the method pretending each donor state was the treated one (in-space placebos)
  • If the real state's post-treatment gap is unusually large compared with the placebo gaps, the effect is unlikely to be chance
  • Re-run with a fake treatment year (in-time placebo)
  • Drop donors that had their own similar reform: they are not untreated
  • Check whether results change when any one donor is dropped
With one treated unit, there is no standard error in the usual sense. Placebo distributions are the evidence.
ImpactMojoCausal Inference 101www.impactmojo.in
10
Section Ten
External validity and heterogeneity
ImpactMojoCausal Inference 101www.impactmojo.in
Right for this sample, and right for the next place
Internal validity
Whether the study's estimate is a correct causal effect for the units and period it studied. Sections 2 to 9 are about this.
External validity
Whether the effect would be similar for other people, places, times, implementers or scales. It is a separate question and needs separate evidence.
A perfectly run trial in 120 schools in one district of Uttar Pradesh tells you what the programme did there, delivered by that organisation, in those years. Whether it would do the same in Tamil Nadu, delivered by the state education department to 30,000 schools, depends on whether the things that made it work are present there too. Those things include the children's starting level, teachers' incentives, the quality of supervision, and what the control group was getting instead.
A common error is to treat a well-identified effect as a constant of nature. It is a measurement of one programme in one context.
ImpactMojoCausal Inference 101www.impactmojo.in
Effects vary more than most reports admit
Eva Vivalt assembled a large dataset of impact evaluation results across development interventions and asked how much effect sizes for the same type of programme vary from study to study. She found a large amount of heterogeneity. Effect sizes varied systematically with study characteristics, and government-implemented programmes had smaller effects than those implemented by academics or non-governmental organisations, even controlling for sample size. Taking study characteristics into account reduced the unexplained variation appreciably.
"How Much Can We Generalize From Impact Evaluations?", Journal of the European Economic Association 18(6), 2020, pp. 3045–3089.
For practitioners
Expect less at scale
When a ministry plans to adopt a programme tested by an NGO, the honest planning assumption is a smaller effect than the trial reported, and the budget case should survive that. Ask for the evidence from government-run versions before assuming the trial number.
ImpactMojoCausal Inference 101www.impactmojo.in
Same contract, different employer, different result
Contract teachers in Kenya
Bold, Kimenyi, Mwabu, Ng'ang'a and Sandefur embedded a randomised trial in a nationwide reform of teacher hiring in Kenyan government primary schools. New teachers on a fixed-term contract offered by an international NGO significantly raised test scores. Teachers on identical contracts offered by the Kenyan government produced zero impact. Bureaucratic and political opposition to the reform led to delays and a different interpretation of the same contract terms. (Journal of Public Economics 168, 2018, pp. 1–20.)
This is a global example, but its lesson applies directly in South Asia, where many tested programmes are run by NGOs and scaled by state governments. The study held the treatment fixed on paper and changed only who delivered it. The effect vanished.
  • Record who implemented a tested programme, and how
  • Ask whether the implementer at scale has the same incentives and capacity
  • Treat "the programme" as including its delivery system
ImpactMojoCausal Inference 101www.impactmojo.in
Teaching at the Right Level: iterating towards scale in India
Pratham's Teaching at the Right Level reorganises instruction around children's actual learning levels in place of the prescribed syllabus. It had worked well when Pratham ran it. Banerjee, Banerji, Berry, Duflo, Kannan, Mukerji, Shotland and Walton describe what happened when it was moved into government schools.
Some designs evaluated by randomised trials produced no impact within the regular school system. Those failures shaped later versions, and two models were eventually developed that raised children's learning levels in government schools at scale. ("From Proof of Concept to Scalable Policies: Challenges and Solutions, with an Application", Journal of Economic Perspectives 31(4), 2017, pp. 73–102.)
What scaled was a process
  • Test the idea where it can work (proof of concept)
  • Test it inside the system that will run it
  • Expect failures and learn why they happened
  • Re-design and test again
  • Report the null results along with the successes
ImpactMojoCausal Inference 101www.impactmojo.in
The graduation approach in Bangladesh and six other countries
21,000+
households in 1,309 villages in BRAC's Bangladesh trial, surveyed four times over seven years
Bandiera et al., QJE 2017
6
countries in the multi-site trial: Ethiopia, Ghana, Honduras, India, Pakistan, Peru
Banerjee et al., Science 2015
10,495
households across the six country sites
Banerjee et al., Science 2015
BRAC's programme for the ultra-poor gives the poorest women livestock assets and training. In Bangladesh it let poor women move from seasonal casual wage labour into livestock rearing, raising their labour supply and earnings and leading to asset accumulation and poverty reduction ("Labor Markets and Poverty in Village Economies", QJE 132(2), 2017, pp. 811–870). A six-country trial of the same multi-part approach, including sites in India and Pakistan, found lasting progress for the very poor (Science 348(6236), 2015).
Running the same design in several countries, with common outcome measures, is the most direct evidence on external validity there is.
ImpactMojoCausal Inference 101www.impactmojo.in
The average can hide who gains and who does not
StudyAverage findingWhat the split revealed
Cash and capital grants to microenterprises, Sri Lanka (de Mel, McKenzie & Woodruff, QJE 2008)High average returns to capital, above market interest ratesReturns varied with ability and household wealth; no positive return in enterprises owned by women
Mindspark, urban India (Muralidharan, Singh & Ganimian, AER 2019)0.37 SD in mathsSimilar absolute gains for all pupils; much larger relative gains for academically weaker pupils
Report cards, Pakistan (Andrabi, Das & Khwaja, AER 2017)Scores up 0.11 SD, fees down 17%Effects differed by schools' initial scores, consistent with better information
Balsakhi, Vadodara and Mumbai (Banerjee et al., QJE 2007)0.28 SDMost of the gain among children at the bottom of the distribution
Subgroup results are useful and fragile. Credible ones were named in a pre-analysis plan, have a mechanism behind them, and are reported with all the other subgroups tested. A single significant subgroup out of twenty, found after the fact, is most likely chance.
ImpactMojoCausal Inference 101www.impactmojo.in
Effects can change when everyone gets the programme
Prices and wages move
A small trial changes nothing in the wider economy. A national programme can. Imbert and Papp found that India's employment guarantee raised private sector wages, so non-participants who hired labour paid more and labourers who never joined earned more. A trial comparing participants with non-participants in the same labour market would miss both effects.
Displacement and competition
If a skills programme helps its trainees win jobs that other workers would otherwise have filled, the trainees gain and the total number of jobs may not change. If one lender enters a market, others may follow; within two years the Hyderabad microfinance trial's control areas had access to microcredit too. The partial effect measured in a trial and the total effect of a national policy can differ in size and even in sign.
Designs that randomise at the level of whole markets, or vary the share of people treated across villages, are the way to measure these effects. Smartcards (157 subdistricts) and the BRAC trial (1,309 villages) were built at that scale.
ImpactMojoCausal Inference 101www.impactmojo.in
11
Section Eleven
Reading a causal claim in practice
ImpactMojoCausal Inference 101www.impactmojo.in
Ten questions to put to any causal claim
#QuestionA good answer looks like
1What exactly is the claimed effect, on what outcome, of what treatment?A number with units, a time period and a defined treatment
2Compared with whom?A named comparison group and the reason it is credible
3Who decided who was treated, and how?A lottery, a rule or a rollout schedule nobody could game
4What is the key assumption, in one sentence?Parallel trends, no sorting at the cutoff, an exclusion restriction, stated plainly
5What evidence supports that assumption?Pre-trends, balance tables, density tests, placebo results
6Which effect: offer or receipt, average or for compliers?ITT, ATT or LATE named
7Were outcomes and analysis fixed in advance?A registration and pre-analysis plan
8How many were lost, and from which group?Attrition by arm, with bounds if it differs
9Who ran it, where, and when?Implementer, setting and dates
10Do other studies with different designs agree?Citations to replications or a systematic review
ImpactMojoCausal Inference 101www.impactmojo.in
Which design fits the situation you are in
Your situationDesign to considerAssumption you will have to defend
Programme not yet started; more demand than placesRandomised lottery or randomised phase-inLittle beyond good implementation and low attrition
Eligibility decided by a score with a cutoffRegression discontinuityNo manipulation of the score; nothing else changes at the cutoff
Rolled out to some areas first, with data before and afterDifference-in-differences (with a staggered-timing estimator if rollout was staggered)Parallel trends between early and late areas
Offer is random but take-up is voluntaryIntention to treat, then IV for the effect of take-upExclusion: the offer matters only through take-up
One state or district treated, many untreated, long pre-periodSynthetic controlA close pre-treatment fit and no shocks unique to the treated unit
Selection decided by a known, recorded checklistMatching or weighting on that checklistNo unrecorded reasons for participation
None of the above, only participants' dataDescribe outcomes; do not claim an effectNone, because no effect is claimed
The best time to choose a design is before the programme starts, when a lottery or a phase-in can still be built into the rollout.
ImpactMojoCausal Inference 101www.impactmojo.in
A state scholarship for girls: choosing the design
Illustrative. A state launches a scholarship of ₹10,000 a year for girls from households below an income limit who pass Class 10 and enrol in Class 11. It begins in 12 districts in 2025 and is to cover all 38 districts by 2027. The finance department asks: does it raise girls' Class 11 enrolment, and is it worth extending?
Designs to reject
  • Recipients against non-recipients: recipients chose to enrol, which is the outcome itself
  • Statewide enrolment before and after: other changes in 2025–26 are mixed in
  • Pilot districts against the rest, after only: pilot districts were picked for a reason
Designs to propose
  • Randomise the order in which the remaining 26 districts join in 2026 and 2027, and run DiD with an estimator built for staggered timing
  • Use the income limit as an RD cutoff, if incomes were certified before the scheme was announced
  • Use boys in the same districts as an additional comparison (triple difference)
Ask early: who certifies income, and could families or officers adjust it once the scheme was known? If yes, the RD option is weak.
ImpactMojoCausal Inference 101www.impactmojo.in
Reading the numbers that come back
Illustrative results after the 2026 phase. Girls' Class 11 enrolment as a share of girls who passed Class 10:
2024 (before)2026 (after)Change
Districts randomly assigned to join in 202658%67%+9 points
Districts randomly assigned to join in 202757%61%+4 points
Difference-in-differences+5 points
Same comparison for boys62% → 66%61% → 65%0 points
Reading: the scholarship raised girls' enrolment by about 5 points, an intention-to-treat effect for all girls who passed Class 10 in those districts. Boys show no difference, which supports the comparison. Report the confidence interval, clustered by district; with 26 districts it will be wide.
For the finance department: divide the total cost by the number of additional enrolments implied (5 points times eligible girls). Dividing by all scholarship recipients would count girls who would have enrolled anyway. Cost Effectiveness 101 explains the calculation.
ImpactMojoCausal Inference 101www.impactmojo.in
Words that make a claim causal, and what they require
Causal verbs
"Led to", "resulted in", "increased", "reduced", "improved", "because of", "impact", "effect of", "attributable to", "thanks to the programme". Each claims a counterfactual. A report that uses them about a participant survey with no comparison group is making a causal claim without causal evidence.
Descriptive verbs
"Was associated with", "participants reported", "rose over the period", "was higher among", "coincided with". These are honest descriptions of patterns. They become misleading when the executive summary turns them back into causal verbs, which happens often between the results chapter and the press release.
  • Check whether the summary's verbs match the design described in the methods section
  • Look for the comparison group in the methods; if it is absent, read every causal verb as "was associated with"
  • Watch for percentages of participants who "benefited": a share of satisfied users is not an effect size
ImpactMojoCausal Inference 101www.impactmojo.in
Recent Indian reforms that will tempt before-after claims
Several large legal changes took effect close together. Any outcome measured across these dates will move for many reasons at once, so a simple before-after comparison attributed to one of them should be read with care. Facts as of October 2026.
ChangeDate in forceWhy it complicates causal claims
Four Labour Codes (Wages; Industrial Relations; Social Security; Occupational Safety, Health and Working Conditions)21 November 2025Definitions of wages, workers and establishments change, so administrative series may break
Income-tax Act 2025 replaces the Income-tax Act 19611 April 2026Tax data before and after follow different statutory definitions
Viksit Bharat G RAM G Act 2025 replaces MGNREGA 2005 (125 days)1 July 2026Rural employment and wage outcomes change for reasons beyond any one programme
Digital Personal Data Protection Act 2023, with the DPDP Rules 2025Board from 13 November 2025; duties and the s17(2)(b) exemption from 13 May 2027Changes what data evaluators can collect and how; research exemption under s17(2)(b) on conditions
Census of India, reference date 1 March 2027, including caste enumerationUpcomingNew baseline counts for sampling frames and running variables in future designs
ImpactMojoCausal Inference 101www.impactmojo.in
What to ask for when you commission an evaluation
  • Before the programme starts. Involve evaluators while rollout can still be randomised or phased. Afterwards, the options shrink to the weaker designs.
  • A written identification strategy. One page saying what the comparison is and which assumption it needs.
  • A pre-analysis plan. Primary outcomes, main specification and subgroups named in advance, and registered.
  • Power. A calculation showing the study can detect an effect large enough to matter for the budget decision.
  • Implementation data. Records of what was delivered, to whom, when.
  • Costs. Collected alongside outcomes, so cost-effectiveness can be computed.
  • Publication of results whatever they show. A null result is information the next state needs.
  • Data and code. Shared in a form another team can re-run, within the limits of the DPDP Act and consent.
Commissioners shape the evidence as much as researchers do. A contract that asks only for a final report at the end of the programme almost guarantees a before-after study of participants.
ImpactMojoCausal Inference 101www.impactmojo.in
12
Section Twelve
Summary and next steps
ImpactMojoCausal Inference 101www.impactmojo.in
What to remember
  • A causal effect compares the outcome with a treatment against the outcome the same units would have had without it. Half of that comparison is always missing.
  • A naive comparison equals the effect plus selection bias. Every design is a way of making the second term zero or small.
  • Draw a diagram first. Adjust for confounders; leave mediators and colliders alone.
  • Randomisation balances everything in expectation. It does not prevent attrition, spillovers or poor implementation.
  • DiD needs parallel trends; check pre-trends, and use modern estimators for staggered rollouts.
  • RD compares units just either side of a cutoff; check for sorting and report the plot.
  • IV needs an exclusion restriction you can state in one sentence, and gives the effect for compliers.
  • Matching handles only what was measured; synthetic control needs a close pre-treatment fit.
  • Effects vary with implementer, context and scale. Government delivery often yields less than an NGO trial.
  • Match a report's verbs to its design before repeating its claims.
The question that does most of the work, in every section: compared with whom, and why were they not treated?
ImpactMojoCausal Inference 101www.impactmojo.in
Terms used in this course
TermMeaning
CounterfactualThe outcome that would have occurred without the treatment, for the same units at the same time
Selection biasThe difference between groups that would exist even with no treatment
Confounder / mediator / colliderA common cause / a step on the causal path / a common effect
ITT / ATT / ATE / LATEEffect of the offer / on the treated / on everyone / on compliers
Parallel trendsDiD assumption that treated and comparison groups would have moved together
Running variableThe score that decides treatment in an RD design
Exclusion restrictionIV assumption that the instrument affects the outcome only through the treatment
Common supportThe range where treated and untreated units with similar characteristics both exist
Donor poolUntreated units from which a synthetic control is built
External validityWhether an effect holds in other places, times, implementers or scales
Pre-analysis planA document fixing outcomes and analysis before data are seen
ImpactMojoCausal Inference 101www.impactmojo.in
Books and guides for going further
ResourceWhy read itAccess
Angrist and Pischke, Mostly Harmless Econometrics (Princeton University Press, 2009)The economist's toolkit for randomisation, regression, IV, DiD and RDBook
Cunningham, Causal Inference: The Mixtape (Yale University Press, 2021)Readable, with code in R and Stata, including the newer DiD estimatorsFree online at mixtape.scunning.com
Huntington-Klein, The Effect: An Introduction to Research Design and Causality (2021)Built around causal diagrams; good for non-economistsFree online at theeffectbook.net
Hernán and Robins, Causal Inference: What IfPotential outcomes and diagrams from epidemiologyFree at miguelhernan.org/whatifbook
Gertler, Martinez, Premand, Rawlings and Vermeersch, Impact Evaluation in Practice, 2nd edition (World Bank)Practical guide for programme managers, with design chaptersFree from the World Bank Open Knowledge Repository
DAGittyDraw a diagram and get the adjustment setsFree at dagitty.net
For South Asian studies, the papers cited on each slide are the best starting points. Many have free working paper versions through NBER, J-PAL or the authors' sites.
ImpactMojoCausal Inference 101www.impactmojo.in
Continue with related courses
Go deeper on causal methods
The flagship Causal Inference course is the next step: it takes the designs in this deck into estimation, with worked exercises and a lexicon.

Impact Evaluation 101 covers the full evaluation cycle, including power and sample size. Econometrics 101 covers the regression tools these designs run on. Multivariate Analysis 101 covers working with many variables at once.
Related skills
Survey Design 101 for collecting the outcomes. Systematic Reviews & Evidence Synthesis 101 for combining many causal studies. Cost Effectiveness 101 for turning an effect into a funding decision. Theory of Change 101 for the causal story a programme assumes. Research Ethics 101 and Data Protection & the DPDP Act 101 for running studies responsibly.
Suggested order: this deck, then Impact Evaluation 101, then Econometrics 101, then the flagship course.
ImpactMojoCausal Inference 101www.impactmojo.in
Causal Inference 101 · Complete
Compared with whom?
Ask it every time.
CC BY-NC-ND 4.0·Free Forever·ImpactMojo 101 Series