fullscreen
ImpactMojoEconometrics 101www.impactmojo.in
ImpactMojo 101 Series · Free Forever
Econometrics
101
From Correlation to Credible Causation — a Foundational Course on Estimating Causal Effects for Development Practitioners in South Asia
Research-BackedSouth Asia Focus100 SlidesFree Access
ImpactMojoEconometrics 101www.impactmojo.in
What We Cover
01
What Econometrics Is
Slides 3–10
02
The Causal Question & the Counterfactual
Slides 11–19
03
OLS Regression, Properly Understood
Slides 20–28
04
Endogeneity & Omitted-Variable Bias
Slides 29–37
05
Randomised Experiments / RCTs
Slides 38–46
06
Instrumental Variables
Slides 47–55
07
Difference-in-Differences
Slides 56–64
08
Regression Discontinuity
Slides 65–73
09
Panel Data & Fixed Effects
Slides 74–82
10
Reading & Critiquing Results
Slides 83–92
11
Tools & Further Reading
Slides 93–99
ImpactMojoEconometrics 101www.impactmojo.in
01
Section One
What Econometrics Is
ImpactMojoEconometrics 101www.impactmojo.in
Econometrics, defined
Econometrics is where economics, statistics and real-world data meet. Its central ambition is not just to describe the world but to estimate causal effects — what would happen if we changed something.
Econometrics
The application of statistical methods to economic and social data in order to test theories and, above all, to measure the causal effect of a policy, programme or treatment on an outcome of interest.
You do not need heavy mathematics to think like an econometrician. You need to ask, relentlessly: compared to what?
Question typeNeedsExample
DescriptiveA representative sampleHow many households are below the poverty line?
PredictiveCorrelation and out-of-sample fitWhich households will fall below it next year?
CausalA credible counterfactualDid the transfer keep them above it?
Prediction and causation come apart completely. A model can forecast defaults perfectly using variables that cause nothing, and a causal estimate can predict badly — they answer different questions.
Most confusion in applied work comes from using a predictive model to answer a causal question. If your conclusion contains the word "should", you needed a design, not a fit.
ImpactMojoEconometrics 101www.impactmojo.in
Economics + statistics + data
Economics
A theory of how people and markets behave — what to look for and why
Statistics
Tools to separate signal from noise and quantify uncertainty
Data
Surveys, censuses, admin records, experiments — the evidence itself
Take away any one ingredient and you get something less: theory without data is speculation; data without theory is pattern-hunting.
IngredientSuppliesMissing it means
Economic theoryWhich variables matter and whyData mining with no interpretation
StatisticsUncertainty and inferencePoint estimates with no error bars
DataThe evidenceA model nobody can test
Theory does the most under-appreciated work. It tells you which confounders are plausible, whether an exclusion restriction is credible, and whether an estimate should transfer — none of which the data can answer.
That is also why purely algorithmic approaches struggle with causal questions. Prediction can be learned from data; identification requires assumptions the data cannot supply.
ImpactMojoEconometrics 101www.impactmojo.in
The question behind every study
Almost every econometric study is, at heart, answering one policy question: does X cause Y, and by how much?
01
Does a cash transfer raise school attendance?
02
Does a new road increase farm incomes?
03
Does microcredit lift consumption?
04
Does a midday meal cut child anaemia?
Each is a causal question. The whole discipline exists because answering it honestly is genuinely hard.
Policy questionThe design its variation supports
Does a cash transfer raise attendance?RCT, if roll-out can be randomised
Does a new road raise farm incomes?DiD across phased construction
Does microcredit lift consumption?RCT, or IV on branch placement rules
Does a scholarship raise completion?RD at the score cutoff
The design is dictated by how the programme was implemented, not by preference. A phased roll-out gives you DiD; a score threshold gives you RD; neither can be conjured if the implementation had neither.
Which is why evaluators should be involved before implementation. A small change to the assignment rule — a lottery among equally eligible applicants — can convert an unevaluable programme into an evaluable one at almost no cost.
ImpactMojoEconometrics 101www.impactmojo.in
Correlation is not causation
Two things moving together — districts with more bank branches having higher incomes — does not mean one caused the other. The link could run the other way, or a third factor could drive both.
Whenever two variables are correlated, at least four explanations are possible — and only one of them is 'X causes Y'.
— a working principle of causal inference
Correlation between X and YPossible explanation
X causes YThe claim being made
Y causes XReverse causality
Z causes bothConfounding
Selection into the sampleCollider bias — created by how you sampled
ChanceEspecially with many comparisons
The fourth is the least known and the hardest to spot. Conditioning on something both X and Y affect — a common sample restriction — creates a correlation between them that does not exist in the population.
The discipline is to name all five explanations before defending the first. A study that only argues against confounding has answered one of four rival accounts.
ImpactMojoEconometrics 101www.impactmojo.in
The comparison that fools you
Suppose villages with a microfinance branch have higher incomes than villages without one. Tempting conclusion: microfinance works. But branches were not placed at random — lenders chose more promising villages to begin with.
The income gap mixes the effect of microfinance with the pre-existing differences between the two kinds of village. Econometrics exists to pull these apart.
Why placement is not randomDirection of bias
Lenders choose promising villagesOverstates the effect
Programmes target the worst-offUnderstates the effect
Villages with better roads get everything firstOverstates
Politically connected villages selectedAmbiguous, and correlated with much else
Programme placement is a decision, and decisions have reasons that also affect outcomes. The direction of the resulting bias depends entirely on the rule, which is why you must know how units were selected.
Ask the implementer directly how sites were chosen. It is the cheapest piece of research in any evaluation and it usually determines which designs are available to you.
ImpactMojoEconometrics 101www.impactmojo.in
Two very different goals
Descriptive / predictive
Who is poor? Where is dropout highest? What will demand be next year? Correlations are enough here.
Causal
Will this programme reduce dropout? Here you must imagine a world without the programme — a counterfactual.
Most policy questions are causal. That is why this course spends most of its time on causal methods.
TaskCorrelation enough?
Targeting: who is poor?Yes
Forecasting demand next yearYes
Deciding whether to fund a programmeNo
Choosing between two designsNo
Describing where dropout is highestYes
Correlational work is not lesser work; it is different work. A well-built targeting model is enormously useful and answers no causal question, and problems arise only when it is read as one.
A quick test: does your sentence contain a hypothetical intervention? "If we did X, then Y" requires a counterfactual. "Where is Y highest" does not.
ImpactMojoEconometrics 101www.impactmojo.in
How the field changed
From the 1990s onwards, econometrics shifted from elaborate models toward research designs that mimic experiments — the 'credibility revolution'. The 2019 and 2021 Nobel prizes recognised RCTs and natural-experiment methods used heavily in development.
2019
Nobel: Banerjee, Duflo & Kremer for the experimental approach to poverty
Sveriges Riksbank Prize
2021
Nobel: Card, Angrist & Imbens for natural experiments & causal methods
Sveriges Riksbank Prize
Before the credibility revolutionAfter
Elaborate structural modelsTransparent research designs
Assumptions buried in functional formOne assumption, stated and debated
"Control for everything""Find exogenous variation"
Identification argued theoreticallyIdentification argued institutionally
The 2019 Nobel went to Banerjee, Duflo and Kremer for experimental approaches to poverty, and the 2021 prize to Card, Angrist and Imbens for natural experiments and causal methodology. Both recognised this shift.
The critique of the design-based approach is worth knowing too: it can favour questions that happen to have clean variation over questions that matter most, and LATE-type estimates over policy-relevant parameters.
ImpactMojoEconometrics 101www.impactmojo.in
02
Section Two
The Causal Question & the Counterfactual
ImpactMojoEconometrics 101www.impactmojo.in
Compared to what?
A causal effect is always a comparison: the outcome with the treatment versus the outcome that would have occurred without it, for the very same unit, at the same time.
Counterfactual
What would have happened to a treated unit had it not been treated. It is never observed — which is the whole difficulty of causal inference.
ComparisonValid counterfactual?
Treated units, before versus afterNo — other things changed too
Treated versus untreated, same periodOnly if they were comparable to begin with
Treated versus their own randomised controlYes
Treated versus a matched comparisonOnly on observables
Before-and-after is the most common evaluation in practice and the weakest. It attributes the entire trend — economic growth, other programmes, seasonality — to the intervention.
Matching is often presented as a design; it is adjustment with extra steps. It balances what you measured, which is precisely the confounding you were least worried about.
ImpactMojoEconometrics 101www.impactmojo.in
Two possible futures for each unit
The potential-outcomes framework imagines, for each person i, two outcomes:
Y₋(1)
outcome if person i receives the treatment
Y₋(0)
outcome if person i does NOT receive it
The individual causal effect is the difference Y₋(1) − Y₋(0). It is exactly what we want — and exactly what we can never see.
NotationMeaningObserved?
Y(1)Outcome if treatedOnly for the treated
Y(0)Outcome if untreatedOnly for the untreated
Y(1) − Y(0)The individual causal effectNever
E[Y(1)] − E[Y(0)]The average treatment effectEstimable, with a design
The framework’s value is that it makes the missing data explicit. Causal inference becomes a missing-data problem, and every design in this course is a strategy for filling in the blanks.
The framework carries an assumption usually left unstated — SUTVA: one unit’s treatment does not affect another’s outcome. Spillovers break it, and in village-level programmes they are common.
ImpactMojoEconometrics 101www.impactmojo.in
The fundamental problem of causal inference
For any one person we observe only one outcome: either the treated state or the untreated state, never both. The other is forever counterfactual.
The fundamental problem of causal inference: we can never observe both Y₋(1) and Y₋(0) for the same unit. Causal inference is, in essence, a missing-data problem.
ImpactMojoEconometrics 101www.impactmojo.in
We estimate averages, not individuals
Since the individual effect is unknowable, we aim instead for the average treatment effect across a group — the mean of Y(1) minus the mean of Y(0) for a population.
Average Treatment Effect (ATE)
The average of the individual causal effects across all units in a population: what the treatment does on average, even though no single person's effect is observed.
The trick is finding a credible stand-in for the unobserved counterfactual outcome of the treated group.
EstimandAnswers
ATEEffect if everyone were treated
ATTEffect on those who actually were
ATUEffect on those who were not
LATEEffect on those an instrument moved
These differ whenever effects vary across people, which is almost always. Asking "what is the effect" is under-specified — the useful question is "the effect on whom".
For a scale-up decision you usually want the ATE or the ATU: the effect on people not yet reached. Most designs deliver the ATT or a LATE, which are effects on people already served.
ImpactMojoEconometrics 101www.impactmojo.in
Why a simple group difference goes wrong
We compare treated people's outcomes with untreated people's outcomes. But the untreated are different people — their average Y(0) need not equal the treated group's Y(0).
Observed gap = True effect (ATT) + Selection bias
Selection bias = how treated & untreated differ in Y(0) before any treatment
ComponentWhat it isZero when
Observed gapWhat you compute from the data
ATTThe true effect on the treatedThe programme does nothing
Selection biasE[Y(0)|treated] − E[Y(0)|untreated]Groups are comparable in absence of treatment
Every design in this course is an argument that the last row is zero. RCTs make it zero by construction; the others argue it away with an assumption you have to evaluate.
Selection bias can exceed the true effect in size and go in either direction. That is why a naive comparison is not "approximately right" — it can have the wrong sign entirely.
ImpactMojoEconometrics 101www.impactmojo.in
The villain of the whole course
Selection bias
The systematic difference in (untreated) outcomes between those who get a treatment and those who do not, arising because they are not comparable to begin with.
Healthier people exercise more; richer farmers adopt new seeds first; motivated students attend coaching. Compare them with everyone else and you measure who they were, not what the treatment did.
SettingWho selects inNaive comparison suggests
MicrofinanceVillages with better prospectsMicrofinance works better than it does
Job trainingThe motivated and job-readyTraining works better than it does
Hospital useThe sickHospitals harm you
Fertiliser adoptionRicher farmers with better landFertiliser works better than it does
Three of four biases run the same way, which is why naive evaluations mostly overstate. Programmes tend to reach the already-advantaged, and their outcomes would have been better anyway.
The hospital row is the useful teaching case precisely because the answer is obvious. The identical logic is invisible when the topic is microcredit, and the bias is the same size.
ImpactMojoEconometrics 101www.impactmojo.in
Do hospitals make people sicker?
People who visited a hospital last year report worse health than people who did not. Does the hospital harm them? Of course not — sick people go to hospital. The comparison is contaminated by who selects in.
This cartoon example is the same logic that wrecks naive evaluations of training, microcredit and health camps. Selection is everywhere.
Naive comparisonSelection storyTrue direction
Hospital users are less healthySick people go to hospitalHospitals help
Trained workers earn moreThe employable enrolTraining helps less than shown
Insured people claim moreThe unwell buy insuranceAdverse selection, not moral hazard alone
Each row is obvious once stated and invisible in a results table. The discipline is to ask, before reading any comparison, how people ended up on each side of it.
Write the selection story explicitly for your own evaluations. If you cannot articulate why the untreated are untreated, you cannot assess whether the comparison means anything.
ImpactMojoEconometrics 101www.impactmojo.in
Every method is a counterfactual strategy
The rest of this course is a toolkit of research designs, each a different way to construct a credible counterfactual — a comparison group that plausibly shows what would have happened anyway.
01
RCTs: randomisation makes groups comparable
02
IV: an external nudge mimics random assignment
03
DiD: a comparison group tracks the trend
04
RD: units just above/below a cutoff are alike
DesignSource of the counterfactualLoad-bearing assumption
RCTRandomly assigned control groupRandomisation worked and held
IVVariation driven by the instrumentExclusion restriction
DiDComparison group’s change over timeParallel trends
RDUnits just the other side of a cutoffNo manipulation of the running variable
Fixed effectsThe unit’s own pastNo time-varying confounders
This table is the whole course. Each design buys a counterfactual with exactly one central assumption, and reading a paper means finding that assumption and asking whether it holds here.
None of these assumptions is testable directly. You can gather supporting evidence — balance tables, pre-trends, density tests — but the assumption itself is always an argument.
ImpactMojoEconometrics 101www.impactmojo.in
03
Section Three
OLS Regression, Properly Understood
ImpactMojoEconometrics 101www.impactmojo.in
Regression: fitting a line through data
Ordinary Least Squares (OLS) finds the straight line that minimises the sum of squared vertical distances between the line and the data points. It is the workhorse of applied economics.
Y = α + βX + ε
outcome = intercept + slope × predictor + error
OLS doesIt does not
Minimise squared vertical distancesEstablish which variable causes which
Give the best linear fit to the dataDetect a non-linear relationship
Provide standard errors under assumptionsValidate those assumptions
Summarise an associationDeliver a counterfactual
Regression is a description tool that becomes a causal tool only when a design makes it one. The arithmetic is identical either way, which is exactly why the distinction has to be argued rather than computed.
Squared distances mean outliers pull the line hard. A handful of extreme observations can drive a coefficient, which is why influence diagnostics belong in every regression workflow.
ImpactMojoEconometrics 101www.impactmojo.in
A fitted regression line
Years of schooling vs monthly wage — OLS line of best fit
Illustrative
The slope says how much wage rises, on average, per extra year of schooling — in this data. Whether that slope is causal is a separate question entirely.
ImpactMojoEconometrics 101www.impactmojo.in
The conditional expectation function
Conditional Expectation Function (CEF)
E[Y | X] — the average value of the outcome Y for each value of the predictor X. OLS gives the best straight-line approximation to this function.
Read a regression as a machine that answers: for units with this value of X, what is the average Y? Nothing more is guaranteed.
Reading E[Y|X]Example
Average Y among units with this XMean wage among those with 10 years of schooling
OLS approximates it with a lineBest straight-line fit to those averages
A good approximation whenThe true relation is roughly linear
A poor one whenThe relation bends, or has thresholds
Reading regression as a conditional-average machine removes most of the mysticism. It answers "what is the average outcome for units that look like this" — no more, and that is a useful thing to know.
Plot the data before fitting. A scatter with a clear bend, or a cluster at zero, will be summarised by a straight line that describes neither region well.
ImpactMojoEconometrics 101www.impactmojo.in
What a slope coefficient means
A coefficient β on X is the predicted change in Y associated with a one-unit increase in X, holding the other included variables fixed.
Two load-bearing words: 'associated' (not necessarily caused) and 'included' (only the variables you put in the model are held fixed — not the ones you left out).
Phrase in the coefficient’s meaningWhat it hides
"Associated with"Not necessarily caused by
"Holding other variables fixed"Only the ones you included
"On average"Effects may differ across the sample
"A one-unit increase"Units matter; check whether X is logged
"Included" is the load-bearing word. A regression holds fixed the variables in the model and nothing else, so the coefficient describes a comparison across units that differ in every unmeasured way.
Read every coefficient aloud with "among the units in this sample, comparing those with the same measured characteristics". If the sentence then sounds like a weak claim, that is the claim.
ImpactMojoEconometrics 101www.impactmojo.in
Controlling for other variables
Adding controls — Y = α + βX + γZ + ε — lets β describe the X–Y link among units with the same Z. This is how we try to compare like with like.
But you can only control for what you measure. The variables you cannot observe — ability, motivation, soil quality — are precisely the ones that cause trouble.
Adding a controlEffect
A genuine confounder you measuredReduces bias — the intended case
A variable that causes Y but not XImproves precision; no bias change
A mediator on the X→Y pathRemoves part of the effect you wanted
A colliderCreates bias where there was none
A proxy for an unobservableReduces bias partly; residual remains
Two of the five make things worse, which is why "add more controls" is not a strategy. Whether a control helps depends on where the variable sits in the causal structure, not on whether it is available.
Controlling for a post-treatment variable is the commonest version of the third row, and it is easy to do by accident when the dataset was assembled after the programme ran.
ImpactMojoEconometrics 101www.impactmojo.in
Levels, logs and dummies
FormCoefficient reads asCommon use
Y on X (levels)ΔY in Y-units per 1-unit ΔXMost variables
log Y on Xapprox. % change in Y per 1-unit ΔXWages, income
log Y on log Xelasticity: % ΔY per 1% ΔXDemand, output
Y on a dummy (0/1)gap in mean Y between the two groupsTreated vs control
Knowing the functional form tells you how to translate a coefficient into a sentence a programme officer understands.
SpecificationCoefficient reads asUse
Y on XΔY in units per 1-unit ΔXMost variables
log Y on XApproximately % ΔY per unit ΔXWages, income
log Y on log XElasticity: % ΔY per 1% ΔXDemand, output
Y on a 0/1 dummyDifference in means between groupsTreatment indicators
The log-Y approximation degrades above roughly 0.1. For a coefficient of 0.5, the percentage change is about 65%, not 50% — use exp(β) − 1 when the coefficient is large.
Logs also drop zeros, which matters when the outcome is income or expenditure in a poor sample. Inverse hyperbolic sine is the usual workaround and has its own interpretation problems.
ImpactMojoEconometrics 101www.impactmojo.in
Gauss–Markov & BLUE
Under a set of assumptions — the Gauss–Markov conditions — OLS is BLUE: the Best Linear Unbiased Estimator. Among all linear unbiased methods, it has the smallest variance.
  • Best: lowest variance among linear unbiased estimators
  • Linear: it is a linear function of the data
  • Unbiased: on average it hits the true value
  • Estimator: a recipe for guessing the parameter
Gauss–Markov assumptionFails whenConsequence
Linearity in parametersThe true relation is non-linearMisspecification
Random samplingThe sample is selectedBiased and unrepresentative
No perfect collinearityTwo regressors move identicallyCannot estimate
ExogeneityX correlates with the errorBiased — not causal
HomoskedasticityError variance varies with XWrong standard errors, not bias
Only one of these threatens causality. The last row affects inference and not the estimate, which is why robust standard errors solve it and nothing solves an exogeneity failure except a better design.
BLUE is a statement about efficiency within the class of linear unbiased estimators. It says nothing about whether the estimand is the causal effect — a distinction that trips up most first courses.
ImpactMojoEconometrics 101www.impactmojo.in
Unbiased ≠ causal
The key Gauss–Markov assumption is that the error term is uncorrelated with X (exogeneity). If something in the error — an omitted cause — correlates with X, OLS is biased and the coefficient is not the causal effect.
This single assumption is where most of applied econometrics lives or dies. The next section is entirely about how it breaks.
PropertyGuaranteed by Gauss–Markov?What it means
UnbiasedYes, given exogeneityRight on average across samples
Efficient among linear estimatorsYesLowest variance in that class
CausalNoRequires exogeneity to be true, not assumed
Correct standard errorsOnly under homoskedasticityUse robust errors otherwise
Unbiasedness is conditional on the assumption that fails most often. BLUE is a mathematical result about a model; whether that model describes your data is a separate empirical question.
This is the hinge of the whole course. Everything from here is about generating the exogeneity that Gauss–Markov assumes and observational data almost never supplies.
ImpactMojoEconometrics 101www.impactmojo.in
04
Section Four
Endogeneity & Omitted-Variable Bias
ImpactMojoEconometrics 101www.impactmojo.in
Endogeneity, defined
Endogeneity
A situation where the explanatory variable X is correlated with the error term — that is, with something that also affects Y but is left out of the model. It makes the OLS coefficient biased and non-causal.
When X is exogenous, OLS recovers the causal effect. When X is endogenous, it does not. Diagnosing endogeneity is the practitioner's core skill.
TermMeans
Exogenous XUncorrelated with the error — OLS is causal
Endogenous XCorrelated with the error — OLS is biased
The error termEverything affecting Y that is not in the model
BiasThe estimate is systematically wrong, not just noisy
Bias does not shrink with sample size. A biased estimate from a million observations is precisely wrong, and the tight confidence interval makes it more persuasive rather than less.
This is the single most important thing to carry out of the course: more data fixes noise, and only a better design fixes bias.
ImpactMojoEconometrics 101www.impactmojo.in
Where endogeneity comes from
Omitted variables
A common cause of both X and Y is left out
Reverse causality
Y also affects X — the arrow runs both ways
Measurement error
X is recorded with noise, biasing its coefficient
All three break the exogeneity assumption. We take them one at a time.
SourceMechanismTypical fix
Omitted variablesA common cause of X and Y is missingIV, RCT, fixed effects
Reverse causalityY also affects XIV, RD, timing
Measurement error in XNoise attenuates the coefficientBetter data, or IV
Selection into the sampleWho is observed depends on YModel the selection
All four produce the same symptom — a biased coefficient — and require different remedies. Diagnosing which one you face is what determines the design, so it cannot be skipped.
They also co-occur. A study of microcredit faces placement selection, reverse causality and mismeasured income at once, which is why single-fix approaches rarely suffice.
ImpactMojoEconometrics 101www.impactmojo.in
Omitted-variable bias, illustrated
Wage vs schooling — the same data split by (omitted) ability
Illustrative
Within each ability group the schooling slope is gentle. Pooled, a steep line appears — because higher-ability people get both more schooling and higher wages. OLS credits schooling for ability's effect.
OLS attributes to schoolingWhat is actually happening
The full pooled slopeHigher-ability people get more schooling and earn more anyway
The pooled line is steeper than any within-group line. That gap is the omitted-variable bias, visible in a picture in a way it never is in a coefficient table.
Ability bias in the returns to schooling is the classic case, and much of the IV literature — quarter of birth, distance to college, compulsory-schooling laws — exists to get around it.
ImpactMojoEconometrics 101www.impactmojo.in
Which way does OVB push?
The direction of omitted-variable bias depends on two signs: how the omitted variable Z relates to X, and how Z relates to Y.
Z → XZ → YBias on β
++Upward (too big)
Upward (too big)
+Downward (too small)
+Downward (too small)
Even when you cannot fix OVB, you can often reason about its direction — and so bound how wrong your estimate might be.
Z→XZ→YBias on β
PositivePositiveUpward — β too big
NegativeNegativeUpward
PositiveNegativeDownward
NegativePositiveDownward
You can sign the bias without measuring the omitted variable. Ability plausibly raises both schooling and wages, so both signs are positive and the OLS return to schooling is biased upward.
That gives you a bound rather than an estimate, and a bound is often enough. If the biased estimate is already below the threshold that would justify a programme, the bias direction settles the decision.
ImpactMojoEconometrics 101www.impactmojo.in
When the arrow runs both ways
Do more police cause more crime? Cross-section data often shows a positive correlation — because cities with more crime hire more police. Causation runs from Y to X.
01
More crime (Y)
02
leads cities to hire more police (X)
03
so X and Y correlate positively
04
even if police actually reduce crime
Sign of reverse causalityCheck
The outcome plausibly drives the treatmentAsk which is decided first
Cross-sectional correlation with a puzzling signMore police, more crime
Treatment responds to needProgrammes target the worst-off
Timing is unclear in the dataLook for a lag structure or a policy date
Reverse causality is why targeted programmes look ineffective in naive comparisons. Assistance flows to the struggling, so recipients do worse — and the programme takes the blame for the targeting.
Lagging the treatment variable is the common fix and rarely sufficient. If the assignment responded to an expected future outcome, a lag does not break the link.
ImpactMojoEconometrics 101www.impactmojo.in
Noise in X biases toward zero
When the predictor X is measured with random error — self-reported income, recalled expenditure — its coefficient is pulled toward zero. This is attenuation bias.
Counter-intuitive but important: messy measurement of X usually makes a real effect look smaller than it is, not larger. Noise in Y, by contrast, mainly inflates uncertainty.
Measurement error inEffect
X (the regressor)Attenuation — coefficient pulled toward zero
Y (the outcome)Larger standard errors; no bias
A control variableIncomplete control; residual confounding
Noise in X and noise in Y do completely different things, which is counter-intuitive and worth remembering: sloppy outcome measurement costs precision, sloppy treatment measurement costs the estimate.
Attenuation means self-reported income, recalled expenditure and mismeasured programme exposure all understate real effects — so a null result from noisy data is weak evidence of no effect.
ImpactMojoEconometrics 101www.impactmojo.in
More controls is not a cure
It is tempting to believe that adding enough control variables removes bias. But you can only control for what you observe and measure. Unobservables — motivation, ability, local governance — remain in the error.
Worse, controlling for the wrong variable — something caused by the treatment, or a collider — can introduce bias. Controls are a scalpel, not a sledgehammer.
Belief about controlsReality
"More controls means less bias"Only for genuine, well-measured confounders
"Controls fixed it — the coefficient barely moved"It may not move for unobservables either
"A rich dataset means we controlled for everything"Ability and motivation are in no dataset
"Adding controls is always safe"Colliders and mediators create bias
The second row is the subtle one. Stability of a coefficient across control sets is often presented as robustness; it can equally mean your controls are uninformative about the confounding that matters.
The honest version of a controlled regression is: "this is the association among units similar on the variables we happened to measure." Everything beyond that is design, not adjustment.
ImpactMojoEconometrics 101www.impactmojo.in
Design beats adjustment
Because we can never be sure we have controlled for every confounder, credible causal work relies on a research design that creates exogenous variation in X — variation unrelated to the unobservables.
The way to estimate a causal effect is not to control for everything, but to find variation in the treatment that is as good as random.
— the design-based view of econometrics
ApproachRelies onCredible when
Adjustment (controls)Having measured every confounderAlmost never, for behaviour
Design (RCT, IV, DiD, RD)Exogenous variation in XThe assumption is defensible
This is the credibility revolution in one line. The field moved from modelling the confounders to finding variation that was never confounded in the first place.
It also explains why so much applied work now looks like detective work about institutions — roll-out dates, eligibility rules, arbitrary thresholds. Those are where exogenous variation actually comes from.
ImpactMojoEconometrics 101www.impactmojo.in
05
Section Five
Randomised Experiments / RCTs
ImpactMojoEconometrics 101www.impactmojo.in
Randomisation solves selection
In a randomised controlled trial (RCT), units are assigned to treatment or control by a coin flip. On average the two groups are identical in everything — observed and unobserved — except the treatment.
Because assignment is independent of potential outcomes, the control group is a credible counterfactual. Selection bias is designed away.
Randomisation deliversWhich no control set can
Balance on observablesA rich dataset can approximate this
Balance on unobservablesNothing else does this
A known assignment mechanismYou do not have to guess how units selected in
A pre-specified analysisReduces scope for fishing
The second row is the entire argument for RCTs. Ability, motivation and local governance are balanced in expectation without anyone measuring them, which is what no regression can achieve.
Randomisation is a property of the assignment procedure, not of the resulting groups. A balance table is evidence the procedure worked; a chance imbalance in a small sample is possible and not a design failure.
ImpactMojoEconometrics 101www.impactmojo.in
Balance in expectation
Randomisation does not make any two specific people identical. It makes the groups statistically equivalent, so their average Y(0) is the same. The control group's outcome stands in for the treated group's missing counterfactual.
Effect = mean Y(treated) − mean Y(control)
and with randomisation, selection bias ≈ 0
Randomisation makesIt does not make
Groups equivalent on averageAny two individuals identical
E[Y(0)] equal across armsEvery covariate exactly balanced
The control a valid counterfactualThe result generalise beyond the sample
Balance holds in expectation, which is a statement about the procedure rather than about your particular draw. Small samples can be imbalanced by chance and the design is still sound.
Stratified or block randomisation forces balance on key covariates while preserving the design. It is standard practice in field experiments with modest sample sizes.
ImpactMojoEconometrics 101www.impactmojo.in
A balance table
Baseline characteristics — treatment vs control (should be similar)
Illustrative balance check
Good randomisation produces near-identical groups at baseline. A balance table is the first thing to check in any RCT — it is the evidence that the design worked.
Reading a balance tableConcern
One variable differs at p < 0.05 out of 20Expected by chance — not alarming
Several differ, all in one directionSuggests the randomisation was compromised
Baseline outcome differsMost serious — controls should be shown
Testing twenty variables at the 5% level produces one significant difference by construction. A single imbalance is not evidence against randomisation, and treating it as such is a misreading of the test.
What matters more is whether imbalance is systematic and whether it is on variables that predict the outcome. Those are the ones that would bias the estimate if unaddressed.
ImpactMojoEconometrics 101www.impactmojo.in
RCTs in development: J-PAL & the field
The Abdul Latif Jameel Poverty Action Lab (J-PAL), founded in 2003, popularised RCTs in development. Hundreds of trials — many in India — have tested deworming, remedial teaching, immunisation incentives, and more.
A landmark example: Pratham's 'Teaching at the Right Level' remedial-education model was refined and scaled through a sequence of RCTs across Indian states.
Development RCT findingConsequence
Remedial teaching by matched instructionPratham’s Teaching at the Right Level, scaled across states
Small incentives raise immunisation uptakeCamps plus lentils outperformed camps alone
Deworming and schoolingContested after replication and re-analysis
The third row is the honest one to include. Trials produce contested findings as well as durable ones, and a course that presents only the successes teaches the wrong lesson about evidence.
J-PAL was founded in 2003 and its India work is extensive. The scaled examples matter most: an effect that survived transfer from trial to government delivery is a stronger result than the trial alone.
ImpactMojoEconometrics 101www.impactmojo.in
The rise of development RCTs
Cumulative development RCTs registered (stylised, illustrative trend)
Illustrative — stylised to show the trend, not exact counts
The exact numbers here are illustrative, but the shape is real: development RCTs grew explosively after the mid-2000s.
ImpactMojoEconometrics 101www.impactmojo.in
Internal vs external validity
Internal validity
Is the estimated effect causally correct for this sample? RCTs are strong here — their headline virtue.
External validity
Will it hold elsewhere — other states, scales, populations? RCTs are often weak here.
A perfectly clean trial in one district may not generalise. 'It worked in Rajasthan' is not 'it will work in Bihar'.
Internal validityExternal validity
QuestionIs the estimate right here?Does it hold elsewhere?
RCTsStrong by designOften weak
Threatened byAttrition, spillovers, non-complianceContext, scale, population, implementer
Fixed byBetter design and executionReplication and theory
A perfectly internally valid trial can be irrelevant to your decision. An intervention that worked in one district, delivered by a specialist NGO at pilot scale, tells you little about a state-wide government roll-out.
The general-equilibrium problem is the sharpest version: a job-training programme that helps its participants may simply reallocate scarce jobs, and a trial at pilot scale cannot detect that.
ImpactMojoEconometrics 101www.impactmojo.in
What can still go wrong
ThreatWhat happensFix / response
AttritionTreated & control drop out differentlyTrack everyone; bound effects
SpilloversControl units affected by treatmentRandomise at cluster level
Non-complianceAssigned but don't take treatmentAnalyse by assignment (ITT)
Hawthorne effectsBeing watched changes behaviourBlinding where possible
Intention-to-treat (ITT) — analysing people by the group they were assigned to — preserves the randomisation even when compliance is imperfect.
ThreatSymptomResponse
AttritionDifferential dropout between armsTrack everyone; report bounds
SpilloversControl units affected by treatmentRandomise at cluster level
Non-complianceAssigned but untreated, or vice versaReport intention-to-treat
Hawthorne / John HenryBehaviour changes from being observedBlinding where possible
Intention-to-treat is the honest estimate under non-compliance. Comparing those who actually took up treatment reintroduces selection, undoing the randomisation entirely — a common and serious error.
Differential attrition is the threat that quietly destroys a trial. If the least successful participants drop out of the treatment arm, the surviving sample is no longer comparable to the control.
ImpactMojoEconometrics 101www.impactmojo.in
Is it ethical to randomise a benefit?
  • Randomise when there is genuine uncertainty about whether the programme works (equipoise)
  • Use waitlists or phased roll-outs so the control group eventually benefits
  • Never withhold a known, proven, life-saving treatment to run a trial
  • Secure informed consent and ethics-board (IRB) approval
Scarce budgets mean not everyone can be served at once anyway. A lottery for limited places can be both fair and a clean experiment.
Ethical conditionMeaning
EquipoiseGenuine uncertainty about whether it works
ScarcityYou could not have served everyone anyway
Informed consentParticipants know they are in a study
Ethics approvalIndependent review, before fieldwork
Eventual accessWaitlist or phased roll-out
Scarcity is the strongest practical justification. Where a programme can reach only a third of eligible villages, a lottery is arguably fairer than the usual alternatives — proximity, connections or convenience.
Never randomise away a known effective, life-saving treatment. Equipoise is a factual condition about the state of evidence, not a formality to assert in an ethics application.
ImpactMojoEconometrics 101www.impactmojo.in
06
Section Six
Instrumental Variables
ImpactMojoEconometrics 101www.impactmojo.in
Borrowing randomness from nature
Often you cannot run an experiment — the treatment already happened, or randomising is impossible. An instrumental variable (IV) finds a source of variation in X that is 'as good as random'.
Instrumental variable (instrument)
A variable Z that shifts the treatment X but affects the outcome Y only through X. It isolates the part of X that is unrelated to the confounders.
You cannot randomise becauseIV may still work if
The treatment already happenedSomething as-good-as-random shifted it
Randomising is unethicalA natural source of variation exists
The programme is universalExposure varied by an arbitrary rule
Assignment was politicalA separate quasi-random nudge exists
IV is what you reach for when the variation you need already exists somewhere in the world. The work is finding it and arguing that it is uncontaminated — not estimating anything difficult.
The estimation is mechanical; the credibility is entirely in the argument. This is why IV papers spend more space on institutional detail than on econometrics.
ImpactMojoEconometrics 101www.impactmojo.in
Use only the exogenous part of X
01
Instrument Z (as-good-as-random)
02
shifts treatment X
03
X changes the outcome Y
04
Z affects Y ONLY through X
IV throws away the endogenous, confounded variation in X and keeps only the clean variation driven by Z. That clean slice yields a causal estimate.
StageWhat is estimatedDiagnostic
First stageEffect of Z on XF-statistic — report it
Reduced formEffect of Z on YIf this is null, IV will be too
IV estimateReduced form divided by first stageSensitive to a weak denominator
Look at the reduced form first. If Z has no detectable effect on Y, no amount of scaling by the first stage will produce a credible effect — and a large IV estimate from a null reduced form is a warning sign.
IV is division, which is why a small first stage is dangerous: dividing by something near zero amplifies both noise and any small violation of the exclusion restriction.
ImpactMojoEconometrics 101www.impactmojo.in
An instrument must satisfy BOTH
1. Relevance
Z must actually shift X — a real, strong first-stage relationship between instrument and treatment. Testable in the data.
2. Exclusion restriction
Z must affect Y only through X — no other pathway, no direct effect, uncorrelated with the error. Untestable; argued, not proven.
BOTH are required. Relevance you can check; the exclusion restriction you must defend with theory and institutional knowledge — it is where most IV claims succeed or fail.
ConditionTestable?How to argue it
Relevance — Z shifts XYes — first-stage FReport the F-statistic
Exclusion — Z affects Y only via XNoInstitutional knowledge and argument
Independence — Z as good as randomPartlyBalance on covariates
Monotonicity — no defiersNoArgument about behaviour
The one you cannot test is the one that matters most. Exclusion is defended by a story about the world, and evaluating that story is the whole job of reading an IV paper critically.
Ask what else Z could plausibly affect. Rainfall affects income — and also health, migration and school attendance, any of which could reach conflict without going through income.
ImpactMojoEconometrics 101www.impactmojo.in
Rainfall as an instrument
To study whether economic downturns fuel conflict, researchers have used rainfall shocks as an instrument for agricultural income in rain-fed economies. Rain is plausibly random year to year.
  • Relevance: rainfall strongly affects farm income (first stage)
  • Exclusion: rainfall is argued to affect outcomes only via income — the part you must defend
  • Caveat: if rain also affects, say, mobility or disease directly, exclusion fails
Rainfall as an instrumentAssessment
RelevanceStrong — rainfall drives rain-fed farm income
IndependencePlausible — annual weather is close to random
ExclusionThe weak link
ObjectionsRain also affects health, migration, road access, mobilisation
Rainfall satisfies two conditions easily and the important one poorly. Any channel from weather to the outcome that bypasses income breaks the exclusion restriction, and there are several.
The lesson generalises: instruments that look wonderfully exogenous are often exogenous and multi-channel. Exogeneity is necessary and nowhere near sufficient.
ImpactMojoEconometrics 101www.impactmojo.in
Distance, sib-sex & quarter of birth
Instrument (Z)Treatment (X)Exclusion argument
Distance to a school/collegeYears of schoolingDistance affects wages only via schooling
Quarter of birthYears of schoolingBirth-month is arbitrary, tied to school-start laws
Sex composition of first 2 kidsHaving a 3rd childSex mix is random, shifts fertility
Each is clever — and each has been challenged on exclusion grounds. A good IV invites scrutiny of the one assumption you cannot test.
InstrumentExclusion argumentStandard objection
RainfallWeather affects income and nothing else relevantAlso affects health, migration, mood
Distance to schoolDistance only matters via schoolingFamilies choose where to live
Quarter of birthBirth month is arbitrarySeason of birth correlates with family type
Sibling sex compositionSex of a first child is quasi-randomNot where sex selection occurs
Every classic instrument has a well-known objection, and the last is specific to South Asia. Sibling-sex instruments assume the sex of births is random, which sex-selective abortion violates directly.
The literature has moved towards more institutional instruments — policy rules, administrative thresholds, roll-out timing — precisely because their exclusion arguments are easier to make concrete.
ImpactMojoEconometrics 101www.impactmojo.in
A local effect, for compliers
IV does not recover the average effect for everyone. It recovers the Local Average Treatment Effect (LATE) — the effect for the compliers, those whose treatment status is moved by the instrument.
So 'the IV estimate' answers a specific question: the effect on people the instrument actually nudged. Different instruments can give different — both correct — LATEs.
GroupDefinitionIncluded in LATE?
CompliersTreated because of ZYes — only these
Always-takersTreated regardless of ZNo
Never-takersUntreated regardless of ZNo
DefiersDo the opposite of ZAssumed not to exist
LATE is the effect on people you usually cannot identify. Compliers are defined by a counterfactual response, so you can characterise them statistically but never list them.
This matters for policy. If your instrument is distance to school, the compliers are people near the margin of attending — and a policy aimed at everyone will not have this effect.
ImpactMojoEconometrics 101www.impactmojo.in
The danger of a weak first stage
If Z only weakly predicts X (a weak instrument), IV estimates become wildly imprecise and can be more biased than plain OLS — even tiny exclusion violations get amplified.
Rule of thumb: report the first-stage F-statistic; a common (rough) threshold is F > 10. A weak instrument is worse than no instrument at all.
First-stage FInterpretation
Below 10Weak instrument — treat results with suspicion
Around 10The traditional rule of thumb, now considered too lenient
Well above 10Adequate relevance
Not reportedAsk why
The F > 10 rule is a rough convention and recent work argues it is too permissive. Weak-instrument-robust inference exists and should be used when the first stage is anywhere near the boundary.
With a weak instrument, IV can be more biased than plain OLS, and biased toward the OLS estimate — so a weak-IV result that "confirms" OLS is confirming nothing.
ImpactMojoEconometrics 101www.impactmojo.in
Three questions for any IV study
  • Is it relevant? Is the first stage strong (high F)?
  • Is exclusion plausible? What is the story for 'only through X', and what would break it?
  • Whose effect is it? Who are the compliers — and do you care about them?
A persuasive IV paper spends most of its words defending the exclusion restriction, not running the regression.
QuestionWhere the answer lives
Is it relevant?The first-stage F, reported in the table
Is exclusion plausible?The institutional argument in the text
Whose effect is this?The complier characterisation, often absent
Does the reduced form hold?Z on Y directly — ask for it
A persuasive IV paper spends most of its length on the second row. If the exclusion argument occupies one sentence, the authors have not done the work that makes the estimate credible.
Complier characterisation is increasingly expected and still often missing. Without it you have an effect on an unnamed subpopulation, which is hard to apply to any policy decision.
ImpactMojoEconometrics 101www.impactmojo.in
07
Section Seven
Difference-in-Differences
ImpactMojoEconometrics 101www.impactmojo.in
Before-and-after, with a comparison group
Difference-in-Differences (DiD) studies a policy that hits one group but not another. It compares the change in the treated group with the change in an untreated comparison group.
Difference-in-Differences
An estimator that subtracts the before–after change in a comparison group from the before–after change in the treated group, netting out both fixed group differences and common time trends.
BeforeAfterDifference
Treated groupABB − A
Comparison groupCDD − C
DiD estimate(B − A) − (D − C)
The first difference removes anything fixed about each group; the second removes anything that changed for both. What survives is what changed for the treated and not the comparison.
DiD requires only repeated cross-sections, not a panel of the same units — which is why it works with survey rounds where individuals cannot be linked over time.
ImpactMojoEconometrics 101www.impactmojo.in
Why subtract twice?
Difference 1
Treated group: after − before (removes fixed traits of the group)
Difference 2
Comparison group: after − before (captures what would have happened anyway)
The DiD estimate is Difference 1 − Difference 2. The comparison group's change is the counterfactual trend for the treated group.
Difference removesWhich handles
First: treated after minus beforeEverything fixed about the treated group
Second: comparison after minus beforeEverything that changed for both groups
What remainsWhat changed only for the treated
The second difference is the comparison group doing its job: it measures the common trend — growth, seasonality, national policy — that a before-and-after study would have credited to the programme.
This also shows what DiD cannot remove: anything that changed for the treated group alone, for reasons unrelated to the treatment. That is precisely the parallel-trends assumption.
ImpactMojoEconometrics 101www.impactmojo.in
A difference-in-differences plot
Outcome over time — treated vs comparison, policy at 'After'
Illustrative
The DiD effect is the gap between the treated group's actual outcome (58) and its counterfactual (46) — about 12 points. The dashed red line is the assumed parallel trend.
ImpactMojoEconometrics 101www.impactmojo.in
Parallel trends — not equal levels
DiD is valid only if, absent the policy, the two groups would have moved in parallel — the same trend over time. The groups need NOT start at the same level.
Common error: thinking DiD requires the groups to be identical before treatment. It does not. It requires their trends to be parallel — a statement about slopes, not levels.
RequirementDiD needs it?
Equal levels before treatmentNo — the commonest misunderstanding
Parallel trends absent treatmentYes — the whole assumption
Similar groupsHelpful, not required
The same units observed twiceNo — repeated cross-sections work
Levels can differ by any amount; it is the trajectories that must match. A comparison group at half the treated group’s baseline is fine if both were rising at the same rate.
Parallel trends is a counterfactual claim about a period that did not happen, so it can never be tested directly. Pre-trend evidence supports it; it does not establish it.
ImpactMojoEconometrics 101www.impactmojo.in
How to support parallel trends
  • Plot pre-treatment trends: did the groups move together before the policy?
  • Run a placebo / event-study check on pre-periods
  • Choose a comparison group as similar as possible to the treated one
  • Be honest: parallel trends is an assumption, never fully provable
Parallel pre-trends do not prove parallel counterfactual trends — but their absence is a serious warning sign.
Evidence for parallel trendsStrength
Plot several pre-treatment periodsThe minimum expected
Event-study coefficients, pre-period near zeroStronger
Placebo test on a fake earlier treatment dateStrong
A comparison group chosen for similaritySupportive, not evidence
One pre-period onlyAlmost none
A single pre-period cannot show a trend. Two points determine a line, and the assumption is about the line — so a paper with one pre-period has offered no evidence at all.
Parallel pre-trends make the assumption more plausible without proving it. A policy anticipated in advance can produce flat pre-trends and violated parallel trends simultaneously.
ImpactMojoEconometrics 101www.impactmojo.in
When policy creates the design
DiD shines with natural experiments — policies rolled out to some states/districts and not others, or at different times. The staggered roll-out supplies the treatment and comparison groups.
Indian examples: the phased district roll-out of NREGA (2006–08), or state-level reforms introduced in some states before others, are natural settings for DiD.
Natural experimentWhy it works
Phased district roll-outLater districts serve as comparisons for earlier ones
A policy in some states onlyNeighbouring states as comparison
An eligibility rule changeCohorts either side of the change
An unanticipated shockNobody could adjust in advance
The last row is the strongest and rarest. Anticipation is the main threat to the others: if districts knew they were next, behaviour shifts before the official date and contaminates the pre-period.
Ask how roll-out order was decided. If the earliest districts were chosen for readiness or political importance, the comparison groups differ systematically and parallel trends is doubtful.
ImpactMojoEconometrics 101www.impactmojo.in
Threats to a DiD design
ThreatWhat it doesWatch for
Diverging trendsGroups were drifting apart anywayNon-parallel pre-trends
Other shocksA second event hits only one groupConcurrent policies
Composition changeWho is in each group shifts over timeMigration, attrition
AnticipationBehaviour changes before the policyPre-period jumps
Recent methods literature also warns that staggered roll-outs with two-way fixed effects can mislead if effects vary over time — use modern DiD estimators.
ThreatWhat to look for
Diverging pre-trendsThe event-study plot before treatment
A concurrent shockWhat else happened to one group at that time?
Composition changeMigration, attrition, sample redefinition
AnticipationBehaviour shifting before the official date
Staggered adoptionWhether a modern estimator was used
The last row is a live methodological issue. With treatment timing that varies across units, the standard two-way fixed-effects estimator can produce a weighted average with negative weights and the wrong sign.
Estimators from Callaway and Sant’Anna, Sun and Abraham, and Borusyak and co-authors address this. A staggered DiD published without one of them, or an equivalent, warrants scepticism.
ImpactMojoEconometrics 101www.impactmojo.in
Questions for any DiD study
  • Did the authors show parallel pre-trends?
  • Is the comparison group genuinely comparable?
  • Could another shock have hit only one group at the same time?
  • With staggered timing, did they use an appropriate modern estimator?
DiD is powerful and intuitive — which is exactly why its one assumption deserves the hardest scrutiny.
Question for a DiD paperRed flag if
Are pre-trends shown?Only one pre-period
Is the comparison group defensible?Chosen with no stated rationale
Any concurrent shock?Not discussed
Staggered timing handled?Plain two-way fixed effects
Standard errors clustered?At the wrong level, or not at all
Two of these five are recent additions to the checklist. Staggered-adoption problems and clustering practice have both changed materially, so older papers can be sound work by the standards of their time and still need re-examination.
DiD is powerful and heavily used, which makes it heavily misused. It is the design where a plausible-looking result most often rests on an assumption nobody checked.
ImpactMojoEconometrics 101www.impactmojo.in
08
Section Eight
Regression Discontinuity
ImpactMojoEconometrics 101www.impactmojo.in
Assignment by an arbitrary cutoff
Many programmes use a threshold rule: a scholarship for scores above 60, a poverty scheme for those below a deprivation score. Regression Discontinuity (RD) exploits that sharp cutoff.
Regression Discontinuity (RD)
A design that compares units just above and just below a cutoff on a 'running variable'. Near the threshold, who lands on which side is essentially random, so the two sides are comparable.
RD needsCheck
A cutoff rule that determines eligibilityIs it actually enforced?
A continuous running variableNot a category dressed as a score
Enough observations near the cutoffBandwidth versus precision
No precise manipulationDensity test around the threshold
Eligibility rules are everywhere in South Asian programmes — BPL lines, SECC deprivation scores, exam cutoffs, population thresholds triggering a facility — which makes RD unusually available here.
The rule must actually bind. Where local officials routinely override the cutoff, you have a fuzzy design at best and no design at all if the overrides correlate with the outcome.
ImpactMojoEconometrics 101www.impactmojo.in
Just-above ≈ just-below
A student scoring 59 and one scoring 61 are, in every meaningful way, alike — ability, background, motivation. Yet one gets the programme and the other does not. The cutoff manufactures a local experiment.
The jump in the outcome at the threshold — a discontinuity that nothing else can explain — is the causal effect of the programme.
Why 59 and 61 are comparableWhy 40 and 80 are not
A two-point gap is mostly noiseA forty-point gap is real ability
Neither could control which side they landedThey differ systematically
Backgrounds are similar in expectationBackgrounds differ
The cutoff manufactures a local experiment out of an administrative decision. Nobody randomised anything; the arbitrariness of the threshold does the same work over a narrow window.
That local quality is both the strength and the limit. The comparison is airtight near the cutoff and says nothing about anyone far from it, which is a real constraint on policy use.
ImpactMojoEconometrics 101www.impactmojo.in
A jump at the cutoff
Outcome vs running variable — treatment assigned above the cutoff (50)
Illustrative
The vertical jump at the cutoff (50) — roughly 41 to 53 — is the estimated effect. The smooth slope on each side is the relationship that would hold without any jump.
ImpactMojoEconometrics 101www.impactmojo.in
Two flavours of RD
Sharp RD
Crossing the cutoff perfectly determines treatment — everyone above is treated, everyone below is not.
Fuzzy RD
Crossing the cutoff only raises the probability of treatment. The jump in take-up is used like an instrument.
Fuzzy RD is essentially IV at the threshold: the cutoff instruments for actual treatment.
Sharp RDFuzzy RD
Crossing the cutoffDetermines treatmentRaises its probability
EstimateJump in the outcomeJump in outcome / jump in take-up
Equivalent toIV, with the cutoff as instrument
InterpretationEffect at the cutoffLATE for compliers at the cutoff
Fuzzy RD inherits every IV caveat. It is division again, so a small jump in take-up at the threshold produces an unstable estimate for the same reason a weak instrument does.
Most real programmes are fuzzy: some eligible households never enrol and some ineligible ones do. Treating a fuzzy design as sharp overstates take-up and understates the effect.
ImpactMojoEconometrics 101www.impactmojo.in
RD gives a LOCAL effect
RD estimates the effect only at the cutoff — for units near the threshold. It says little about people far from it.
Key caveat: the RD effect is local. A scholarship's effect for students scoring 59–61 may differ entirely from its effect for those scoring 90. Do not over-generalise the jump.
RD estimates the effectIt does not estimate
At the cutoffThe effect far from the cutoff
For marginal unitsThe effect for clearly eligible units
Under the current ruleWhat would happen if the cutoff moved
The third row is the policy-relevant limitation. An RD tells you the effect at a threshold and not what happens if you shift the threshold — which is usually the decision being contemplated.
The local effect can still be exactly what is needed. Whether to extend a scholarship to those just below the line is precisely a question about marginal students.
ImpactMojoEconometrics 101www.impactmojo.in
Manipulation of the running variable
RD fails if people can precisely manipulate which side of the cutoff they land on — an examiner nudging a 59 to a 61, a household mis-reporting assets to qualify.
Diagnostic: check for bunching — a suspicious pile-up of cases just on the favourable side of the cutoff (a McCrary density test). Smoothness across the threshold is the credibility test.
Manipulation evidenceDiagnostic
Bunching just above the cutoffMcCrary density test
Covariates jump at the thresholdThey should be smooth
Suspicious rounding in scoresLook at the raw distribution
Discretion in who is assessedInstitutional knowledge
Manipulation is the one threat that destroys RD outright. If people can position themselves precisely, those just above the cutoff differ from those just below in exactly the ways that matter.
Partial manipulation is common and less fatal: examiners nudging borderline scores upward affects a narrow band, and a donut RD excluding the immediate neighbourhood of the cutoff is the standard response.
ImpactMojoEconometrics 101www.impactmojo.in
Bandwidth and functional form
  • Bandwidth: how wide a window around the cutoff to use — narrow is cleaner but noisier
  • Functional form: fit flexible curves each side; beware high-order polynomials that invent jumps
  • Covariate smoothness: other variables should NOT jump at the cutoff — a useful placebo check
Good RD work shows the estimate is robust to the bandwidth choice, not an artefact of one window.
ChoiceTrade-offGood practice
BandwidthNarrow is cleaner but noisierShow results across bandwidths
Functional formFlexibility versus invented jumpsLocal linear; avoid high-order polynomials
CovariatesShould not change the estimate muchReport with and without
Placebo cutoffsShould find nothingTest several
High-order global polynomials manufacture discontinuities that are not there — a well-documented failure. Local linear estimation with a data-driven bandwidth is the current standard.
A result that appears only at one bandwidth is not a result. The sensitivity plot across bandwidths is the single most informative figure in an RD paper.
ImpactMojoEconometrics 101www.impactmojo.in
RD in development practice
RD is ideal wherever eligibility hinges on a score or threshold: poverty-line targeting (BPL cutoffs, SECC deprivation scores), exam-based scholarships, population thresholds that trigger a facility or grant.
Because eligibility rules are everywhere in Indian welfare programmes, RD is often the most natural — and most credible — design available to an evaluator.
Indian eligibility ruleRD opportunity
BPL and SECC deprivation scoresHouseholds just either side of the cutoff
Exam-based scholarshipsStudents just above and below the mark
Population thresholds for a facilityVillages just either side
Land-holding ceilings for schemesFarms just either side
Threshold rules are pervasive in Indian programme design, which makes RD unusually available here. Any scheme with a score-based eligibility test is a potential design waiting for someone to use it.
Check enforcement before assuming a design exists. Where local discretion routinely overrides the score, you have fuzzy assignment at best and manipulation at worst.
ImpactMojoEconometrics 101www.impactmojo.in
09
Section Nine
Panel Data & Fixed Effects
ImpactMojoEconometrics 101www.impactmojo.in
What panel data buys you
Panel data tracks the same units — households, districts, firms — across multiple periods. This repeated observation lets us net out stable, unchanging differences between units.
Panel (longitudinal) data
Data on the same set of units observed at two or more points in time, combining a cross-section with a time dimension.
Panel data allowsCross-section does not
Netting out fixed unit characteristicsCannot — they are unobserved
Observing change within a unitOnly differences between units
Controlling for common time shocksNo time dimension
Studying dynamics and lagsNo
Repeated observation is what converts unobservable characteristics into something you can remove. You never learn what they are; you simply arrange the comparison so they cancel.
The cost is attrition. Panels lose units over time and the losses are rarely random, which reintroduces exactly the selection the design was meant to eliminate.
ImpactMojoEconometrics 101www.impactmojo.in
Each unit becomes its own control
With fixed effects, we compare each unit to itself over time. Anything about the unit that stays constant — and so cannot explain changes — is swept out of the comparison.
Yᵢₜ = βXᵢₜ + αᵢ + δₜ + εᵢₜ
αᵢ = unit fixed effect  δₜ = time fixed effect
Fixed effectAbsorbs
Unit (αᵢ)Everything constant about that unit — geography, culture, institutions
Time (δₜ)Everything affecting all units in that period — national shocks, inflation
BothThe standard two-way specification
Unit-specific trendsDifferent trajectories per unit — demanding of the data
Unit fixed effects control for unmeasured characteristics you never had to name, which is their power. They cost you any variable that does not change within a unit — you cannot estimate the effect of caste or geography.
Adding unit-specific trends is often proposed as a robustness check. It is a strong requirement and frequently leaves too little variation to identify anything.
ImpactMojoEconometrics 101www.impactmojo.in
Fixed effects use within-unit variation
The fixed-effects (or within) estimator subtracts each unit's own average from every observation, so β is identified only from how a unit changes relative to itself over time.
Differences between units — rich vs poor district, fertile vs arid land — are discarded. Only the within-unit story remains, and that is what removes time-invariant confounders.
Within estimator usesDiscards
How a unit changes over timeDifferences between units
Units that actually changeUnits with no variation — they contribute nothing
Variation after de-meaningAny time-invariant regressor
Units that never change drop out of the estimation entirely. If only a handful of districts switched treatment status, your result rests on those few, whatever the total sample size suggests.
Report how much identifying variation there is: how many units switched, when, and in which direction. A fixed-effects estimate on n = 50,000 driven by three districts is a different claim.
ImpactMojoEconometrics 101www.impactmojo.in
Fixed effects remove only TIME-INVARIANT confounders
Unit fixed effects control for everything about a unit that is constant over time — geography, culture, fixed institutions — even if you never measured it.
Critical caveat: they do nothing about confounders that change over time. A district-specific shock that moves with the treatment will still bias β. Fixed effects are not a magic exogeneity machine.
ConfounderFixed effects handle it?
District geography and climateYes — time-invariant
Persistent local institutionsYes
A district-level policy changeNo — time-varying
Local economic growthNo
Changing local leadershipNo
Fixed effects are frequently over-sold as solving confounding. They solve exactly one class of it, and the confounders that most often drive results in development — other things happening at the same time — are not in that class.
Ask what else changed in treated units during the study window. If a state adopted three reforms in the same year, unit fixed effects separate none of them.
ImpactMojoEconometrics 101www.impactmojo.in
Two ways to model the unit term
Fixed effectsRandom effects
AssumesUnit term may correlate with XUnit term uncorrelated with X
UsesWithin-unit variation onlyWithin + between variation
Robust toTime-invariant confoundingMore efficient if assumption holds
Safer whenYou fear omitted unit traitsStrong, often unrealistic
For causal work where you worry about unobserved unit traits, fixed effects is usually the safer default — it makes the weaker assumption.
Fixed effectsRandom effects
AssumesUnit term may correlate with XIt does not
UsesWithin-unit variation onlyWithin and between
EfficiencyLowerHigher, if the assumption holds
RobustnessRobust to time-invariant confoundingBiased if the assumption fails
Default to fixed effects in applied causal work. The random-effects assumption — that unobserved unit characteristics are uncorrelated with the treatment — is exactly what you doubted in the first place.
A Hausman test compares the two. Failing it means random effects are inconsistent; passing it is weak evidence, since the test has low power in the samples typical of development data.
ImpactMojoEconometrics 101www.impactmojo.in
A close cousin: differencing
With two periods, first-differencing — regressing the change in Y on the change in X — removes the fixed unit term just as fixed effects do. (DiD is exactly this idea with a comparison group.)
Fixed effects, first differences and DiD are a family: all exploit repeated observation to subtract away stable, unobserved differences between units.
EstimatorRelationship
First differencesRemoves the unit term by subtracting last period
Fixed effectsRemoves it by subtracting the unit mean
DiDFirst differences, with a comparison group
With two periodsFirst differences and fixed effects are identical
These are one family, not four methods. Each removes a unit-specific term by differencing something away, and they differ mainly in what they subtract and how many periods they use.
With more than two periods they diverge, and which is preferable depends on the error structure — first differences handle serially correlated errors better, fixed effects are more efficient otherwise.
ImpactMojoEconometrics 101www.impactmojo.in
Why you must cluster
In panel data, a unit's observations are correlated across time — this year looks like last year. Ignoring that makes standard errors far too small and 'significance' spurious.
Fix: cluster the standard errors at the unit level (e.g. by district or village). Clustering acknowledges that observations within a group are not independent — honest uncertainty, not inflated confidence.
Cluster atWhen
The unit of treatment assignmentAlmost always the right answer
Village or districtWhen treatment varies at that level
Too fine a levelStandard errors too small; false significance
Too few clusters (under ~40)Cluster-robust errors themselves unreliable
Cluster at the level of treatment assignment, not the level of observation. A village-randomised programme analysed with household-level standard errors will report significance that is not there.
With few clusters the correction fails in the other direction. Wild cluster bootstrap is the standard remedy, and a paper with twelve clusters and conventional standard errors is reporting fiction.
ImpactMojoEconometrics 101www.impactmojo.in
Questions for a fixed-effects study
  • Are both unit and time fixed effects included where needed?
  • Could a time-varying confounder still drive the result?
  • Are standard errors clustered at the right level?
  • Is β identified from credible within-unit variation, or a few odd cases?
Fixed effects buy a lot — but remember what they cannot buy: protection from confounders that move over time.
QuestionWeak answer
Are unit and time effects both included?Only one, with no reason given
Could a time-varying confounder drive this?Not discussed
Are errors clustered correctly?At the observation level
Where does the variation come from?Unreported
The second question is the one fixed-effects papers most often dodge. Absorbing time-invariant confounders is presented as if it addressed confounding in general, and it does not.
Ask what else changed in treated units during the window. In development settings, programmes arrive in bundles, and fixed effects separate none of them from one another.
ImpactMojoEconometrics 101www.impactmojo.in
10
Section Ten
Reading & Critiquing Results
ImpactMojoEconometrics 101www.impactmojo.in
Standard errors quantify sampling noise
Standard error
A measure of how much an estimate would vary across repeated random samples. It quantifies sampling uncertainty — how precisely the effect is pinned down — not whether the design is valid.
A small standard error says the number is precise. It says nothing about whether the number is right — a biased design gives precisely wrong answers.
A small standard error meansIt does not mean
The estimate is precisely measuredThe estimate is correct
Sampling noise is lowThe design is valid
A large sample, or low varianceThe absence of bias
Precision and validity are independent. A biased estimate from administrative data covering a whole population has essentially no sampling error and is still wrong.
This is why design questions come first when reading a paper. If the identification fails, the standard errors describe how precisely you have measured the wrong quantity.
ImpactMojoEconometrics 101www.impactmojo.in
Report a range, not just a point
A coefficient of 0.12 is shorthand. The honest version is a confidence interval — say [0.04, 0.20] — the range of effects consistent with the data at, usually, 95% confidence.
If a 95% interval comfortably includes zero, the data cannot rule out 'no effect'. Always read the interval, not just the point estimate or the stars.
Reported asTells you
β = 0.12A point estimate and nothing else
β = 0.12 (se 0.04)Precision; you can compute the interval
β = 0.12, 95% CI [0.04, 0.20]The range consistent with the data
β = 0.12***Only that it differs from zero
Stars are the least informative way to report a result and the most common. They tell you about a comparison with zero, which is rarely the comparison a decision-maker needs.
An interval that includes zero does not mean no effect. It means the data cannot distinguish the effect from zero — which with a small sample is entirely compatible with a large true effect.
ImpactMojoEconometrics 101www.impactmojo.in
Same point estimate, very different certainty
Three studies, all estimating a +4-point effect (95% intervals)
Illustrative
All three centre on +4, but only Study A rules out zero. Study C's interval spans negative values — its 'effect' is indistinguishable from no effect. The point estimate alone hides this.
StudyEstimateIntervalConclusion
A+4[+2, +6]Effect, precisely estimated
B+4[0, +8]Suggestive; cannot rule out zero
C+4[−6, +14]Uninformative
All three would be reported as "an effect of 4 points". Only the interval distinguishes a finding from noise, which is why point estimates quoted alone are close to meaningless.
Study C is also compatible with a large positive effect. A wide interval is not evidence of no effect — it is evidence that the study could not tell.
ImpactMojoEconometrics 101www.impactmojo.in
What a p-value does — and doesn't — say
p-value
The probability of seeing an estimate at least this extreme if the true effect were zero. A small p-value means the result is unlikely under 'no effect' — nothing more.
Statistically significant ≠ large, important, or causally valid. With a big sample a trivial effect can be 'significant'. Always ask: how big is the effect, and from a credible design?
A p-value isA p-value is not
P(data this extreme | no effect)P(no effect | data)
A statement about samplingA statement about importance
Dependent on sample sizeA measure of effect size
Meaningful only if the design is validA substitute for identification
The first row is the error almost everyone makes. A p-value of 0.03 does not mean a 3% chance the effect is zero — that is a different quantity requiring a prior.
With a large enough sample, a trivially small effect becomes significant. With a small one, an important effect may not be. Significance is a statement about precision, not about the world.
ImpactMojoEconometrics 101www.impactmojo.in
Is the effect big enough to matter?
An effect can be real, precise and significant — yet too small to justify the cost. Always translate a coefficient into something a decision-maker feels: rupees, percentage points, children, school days.
Pair statistical significance with practical significance and a cost comparison. 'Significant' is the start of the conversation, not the end.
CoefficientTranslated
0.12 log points on wagesAbout 12% higher earnings
0.08 standard deviations on test scoresA few weeks of additional learning
2.3 percentage points on enrolment23 additional children per 1,000
₹340 per household per yearCompare against the cost per household
Standard-deviation effects are the least interpretable and the most reported in education research. Always ask what they translate to in learning time, or you cannot judge whether the programme is worth its cost.
Pair every effect with its cost per unit. An intervention with a small but very cheap effect can beat a large expensive one, and neither statistic answers that alone.
ImpactMojoEconometrics 101www.impactmojo.in
Does the result survive poking?
  • Does the estimate hold across alternative specifications and control sets?
  • Does it survive different samples, sub-groups and outlier handling?
  • Are there placebo tests that should find nothing — and do?
  • Do the authors show the result is not knife-edge on one choice?
A finding that appears only under one precise specification is fragile. Credible results are robust ones.
Robustness checkWhat it would reveal
Alternative specificationsWhether the result depends on modelling choices
Alternative samplesWhether it is driven by a subgroup
Outlier handlingWhether a few observations carry it
Placebo testsWhether the design finds effects where none should exist
Multiple-hypothesis correctionWhether it survives testing many outcomes
Placebo tests are the most informative and the least performed. A design that detects an effect on an outcome the programme could not possibly affect has failed, whatever it reports elsewhere.
Beware the robustness section that shows twelve variants all significant at exactly 5%. Genuine robustness looks like a coefficient that moves a little; identical results across variants suggests selection.
ImpactMojoEconometrics 101www.impactmojo.in
The garden of forking paths
Try enough specifications, subgroups and outcomes and some will cross p < 0.05 by chance. Reporting only those — p-hacking — manufactures false findings that will not replicate.
Be wary of a lone 'significant' subgroup, an oddly specific specification, or many outcomes with one star. Ask what was tested but not reported.
Forking pathHow it inflates false findings
Many outcomes tested, one reportedOne in twenty crosses 0.05 by chance
Subgroup analysis after seeing resultsSubgroups multiply the comparisons
Specification chosen after the factSelection on the result
Outliers dropped once the answer is knownSelection again
None of this requires dishonesty. A researcher making each choice in good faith, informed by the data, produces the same inflation as one fishing deliberately — which is why pre-registration exists.
Be most suspicious of a single significant subgroup with no prior reason to expect it. That is the signature pattern, and such findings replicate poorly.
ImpactMojoEconometrics 101www.impactmojo.in
Tie your hands in advance
A pre-analysis plan — specifying hypotheses, outcomes and methods before seeing the data — removes the freedom to fish. Registries (e.g. the AEA RCT Registry) make the commitment public.
Pre-registration cannot make a bad design good, but it makes an honest design credible — readers know the result was not cherry-picked after the fact.
A pre-analysis plan fixesBefore
Primary and secondary outcomesData collection
The main specificationSeeing results
Subgroups to be examinedAny subgroup analysis
Multiple-testing correctionReporting
Pre-registration cannot rescue a bad design. It makes deviations visible, which is a different and more limited claim than the one often made for it — and still a substantial improvement.
Read the registered plan against the published paper. Outcomes that appear in one and not the other, or a primary outcome demoted to secondary, is the check the registry exists to enable.
ImpactMojoEconometrics 101www.impactmojo.in
Will it travel?
A clean estimate is internally valid for its setting. Whether it generalises — to another state, scale, time or population — is external validity, a separate and often harder question.
Ask: what was the context and the sample? Would the mechanism plausibly work elsewhere? Scaling up can itself change the effect (general-equilibrium and implementation effects).
Before transferring an estimate, askBecause
Who was in the sample?The effect may be specific to them
What was the context?Complementary conditions may be absent
Who implemented it?A specialist NGO is not a line department
At what scale?General-equilibrium effects appear at scale
What was the counterfactual?The control condition differs across settings
The last row is the most neglected. A programme evaluated against nothing looks better than the same programme evaluated against an existing service — and the comparison, not the programme, differs.
Theory is what makes transfer possible. If you understand the mechanism, you can predict where it will and will not hold; without one, you are extrapolating a number.
ImpactMojoEconometrics 101www.impactmojo.in
11
Section Eleven
Tools & Further Reading
To go further, you needStart with
Intuition before mathematicsAngrist & Pischke, Mastering ‘Metrics
The design-based referenceMostly Harmless Econometrics
Code alongside theoryCunningham, Causal Inference: The Mixtape, free online
Development applicationsJ-PAL and IPA policy briefs
Start with Mastering ‘Metrics if you have no econometrics background. It covers the same five designs with minimal mathematics and is the gentlest credible entry point.
The Mixtape is free and includes runnable code in both R and Stata, which makes it the fastest route from reading about a design to estimating one.
ImpactMojoEconometrics 101www.impactmojo.in
What practitioners actually use
ToolGood forNote
StataApplied micro-econometrics, panel, IV, RDIndustry standard; paid
RFree, flexible, reproducible analysis & graphicsRich causal-inference packages
Python (statsmodels, linearmodels)Automation, large data, MLFree, general-purpose
Excel / SheetsQuick description, not inferenceFine to start; outgrow it
The software matters far less than the research design. A clean design in Excel beats a flawed one in Stata.
ToolStrengthConsideration
StataThe applied-micro standard; panel, IV, RD packagesPaid licence
RFree; strong causal-inference ecosystemSteeper start
PythonAutomation, large data, integrationFewer specialised econometrics packages
Choose by what your collaborators use. Reproducibility across a team matters more than any feature difference, and code nobody else can run is a liability in a multi-year project.
Whichever you pick, keep the analysis as a script rather than a sequence of clicks. Every design in this course requires showing your working, and a script is what makes that possible.
ImpactMojoEconometrics 101www.impactmojo.in
Which design for which situation?
If you can…UseKey assumption
Randomise treatmentRCTSuccessful randomisation
Find an as-good-as-random nudgeIVRelevance + exclusion
Compare a treated & untreated group over timeDiDParallel trends
Exploit a cutoff ruleRDNo manipulation at cutoff
Follow units over timePanel fixed effectsConfounders time-invariant
Start from the variation you have, then pick the design — not the other way round.
If you can…UseAnd must defend
RandomiseRCTBalance, attrition, spillovers, compliance
Find as-good-as-random variationIVRelevance and exclusion
Compare groups over timeDiDParallel trends
Exploit a cutoffRDNo manipulation; local validity
Follow units over timeFixed effectsNo time-varying confounders
Read the right-hand column as the questions to ask of any paper. Identifying the design tells you immediately which single assumption the entire result rests on.
The designs are not ranked. An RD with a clean cutoff is often more credible than a poorly executed RCT with heavy differential attrition, and the hierarchy depends on execution.
ImpactMojoEconometrics 101www.impactmojo.in
Angrist & Pischke
  • Mostly Harmless Econometrics — Angrist & Pischke (the design-based bible; RCT, IV, DiD, RD)
  • Mastering 'Metrics — Angrist & Pischke (gentler, intuitive introduction — start here)
  • Causal Inference: The Mixtape — Scott Cunningham (free online, code-rich)
If you read one book, read Mastering 'Metrics. It teaches the five core designs in this course with humour and real studies.
ImpactMojoEconometrics 101www.impactmojo.in
Evidence in development
  • Poor Economics — Banerjee & Duflo (RCTs and the lives of the poor)
  • Running Randomized Evaluations — Glennerster & Takavarasha (a practical RCT field guide)
  • J-PAL & Innovations for Poverty Action (IPA) — policy briefs and evidence syntheses
Read the methods section of real studies critically — it is the fastest way to internalise the ideas in this deck.
ImpactMojoEconometrics 101www.impactmojo.in
If you remember five things
  • Always ask 'compared to what?' — causation needs a counterfactual
  • Correlation is not causation — suspect selection and confounding first
  • Design beats adjustment — RCT, IV, DiD, RD, fixed effects each build a comparison
  • Every method has ONE load-bearing assumption — know it, and interrogate it
  • Read the standard errors and the robustness — precise is not the same as right
TakeawayThe mistake it prevents
Ask "compared to what?"Treating a before-after change as an effect
Suspect selection firstReading programme placement as programme impact
Design beats adjustmentBelieving controls remove confounding
Find the load-bearing assumptionAccepting a result without knowing what it rests on
Precision is not validityTrusting a tight interval around a biased estimate
The fourth is the reading skill worth practising. For any empirical claim, name the design, name its assumption, and ask what would break it — three questions that separate credible work from the rest.
You do not have to run these methods to use them. Most practitioners consume econometrics rather than produce it, and knowing what to interrogate is the more valuable half.
ImpactMojoEconometrics 101www.impactmojo.in
Keep building
Econometrics is learned by doing. Take a real Indian dataset, pose a causal question, and ask which design its variation can support — then stress-test the assumption that design relies on.
Pair this deck with ImpactMojo's Data Literacy, Impact Evaluation and Exploratory Data Analysis 101 courses.
ImpactMojoEconometrics 101www.impactmojo.in
Econometrics 101 · Complete
Now go ask
'compared to what?'
CC BY-NC-ND 4.0·Free Forever·ImpactMojo 101 Series