| The question requires | Which most reports lack |
|---|---|
| A comparison group | Only the treated are measured |
| Measurement before and after | Only endline exists |
| A stated identification strategy | None is offered |
| A pre-specified outcome | Outcomes chosen afterwards |
| Monitoring | Impact evaluation | |
|---|---|---|
| Asks | Are we doing what we planned? | Did it cause the change? |
| Tracks | Inputs, activities, outputs | Outcomes vs a counterfactual |
| Timing | Continuous, routine | Periodic, designed in advance |
| Example | 1,200 toilets built | Did diarrhoea actually fall because of them? |
| Answers | Implementation | Attribution |
| Monitoring | Impact evaluation | |
|---|---|---|
| Question | Are we on plan? | Did we cause it? |
| Frequency | Continuous | Once or twice |
| Needs a comparison | No | Yes |
| Cost | Built into delivery | A separate budget |
| Can correct delivery | Yes | Usually too late |
| Level | Monitored by | Evaluated by |
|---|---|---|
| Inputs and activities | Routine systems | — |
| Outputs | Routine systems | Process evaluation |
| Outcomes | Periodic survey | Impact evaluation |
| Impact | Rarely | Impact evaluation, if at all |
| Outcomes improved because | Unless you rule it out |
|---|---|
| The programme worked | You cannot claim it |
| The monsoon was good | A trend, not your effect |
| Incomes were rising anyway | A secular trend |
| Another scheme arrived | A confounder |
| You selected the promising villages | Selection bias |
| With credible evidence | Without it |
|---|---|
| Scale what works | Scale a failure |
| Stop what does not | Keep funding it |
| Defend the budget | Defend it on anecdote |
| Learn the mechanism | Repeat the design |
| Do an impact evaluation when | Do something else when |
|---|---|
| The question is causal and open | It is already well established |
| A real decision hangs on it | Nobody will act on the answer |
| The programme is stable | It changes weekly |
| A credible comparison exists | Everyone is treated |
| You can afford it | Monitoring would serve better |
| Commissioning requires you to | Which needs |
|---|---|
| Frame the question | Knowing what decision it informs |
| Judge a design | Knowing what each design assumes |
| Read a power calculation | Understanding MDE |
| Interpret the finding | Distinguishing null from inconclusive |
| Activity | Question it answers |
|---|---|
| Needs assessment | What is the problem, and for whom? |
| Process evaluation | Was the programme delivered as designed? |
| Monitoring | Are we on track against the plan? |
| Impact evaluation | Did the programme cause the change? |
| Cost-effectiveness analysis | Was the impact worth the cost? |
| Activity | Answers | Timing |
|---|---|---|
| Needs assessment | What is the problem? | Before design |
| Process evaluation | Was it delivered as designed? | During |
| Monitoring | Are we on track? | Continuous |
| Impact evaluation | Did it cause the change? | After, designed before |
| Cost-effectiveness | Was it worth it? | With the impact estimate |
| The counterfactual is | It is not |
|---|---|
| What would have happened to the same people | What happened to different people |
| In the same period | Before the programme |
| Absent the programme | A national average |
| Unobservable, always | Something you can measure directly |
| We cannot observe | So we estimate |
|---|---|
| The same person, treated and untreated | A group average |
| An individual effect | An average treatment effect |
| The true counterfactual | A credible substitute for it |
| Wrong comparison | What it actually measures |
|---|---|
| Before versus after | The programme plus everything else that changed |
| Participants versus non-participants | Who chose to join |
| Treated villages versus national average | Differences in the villages |
| Completers versus dropouts | Who could complete |
| Selection can come from | Example |
|---|---|
| Self-selection | The motivated volunteer |
| Programme placement | Villages chosen for likely success |
| Administrative targeting | The needlest selected |
| Attrition | Who remains at endline |
| Volunteers differ in | Which also affects |
|---|---|
| Motivation | The outcome |
| Prior education | The outcome |
| A good comparison group | Test it by |
|---|---|
| Looks alike at baseline | A balance table |
| Shared the same prior trend | Pre-period data |
| Faces the same outside forces | Same region, same period |
| Differs only in the programme | Ruling out other schemes |
| Estimand | Answers | Relevant when |
|---|---|---|
| ATE | Effect if everyone were treated | Considering universal rollout |
| ATT | Effect on those actually treated | Assessing the current programme |
| LATE | Effect on those an instrument moved | IV designs |
| Design | The promise it makes |
|---|---|
| RCT | Chance made the groups comparable |
| RD | Units either side of the cutoff are alike |
| DiD | The trends would have been parallel |
| Matching | Observed characteristics capture the difference |
| IV | The instrument affects the outcome only through take-up |
| Without a theory of change you cannot | Because |
|---|---|
| Choose outcomes | You do not know what should move |
| Interpret a null | You cannot say which link failed |
| Time the measurement | You do not know when to expect it |
| Test mechanisms | There are none stated |
| Link in the chain | The assumption |
|---|---|
| Transfer to income | It arrives, and is retained |
| Income to spending on food | It is spent that way |
| Spending to nutrition | The food improves diet quality |
| Nutrition to stunting | The window is right; enough time passes |
| Failure type | What it means | What to do |
|---|---|---|
| Implementation failure | The chain broke in delivery | Fix delivery; the theory is untested |
| Theory failure | Delivery worked; the link did not | Change the design |
| Question type | Method that answers it |
|---|---|
| Did it work, and by how much? | Impact evaluation |
| Through which link? | Mechanism analysis; qualitative |
| Was it delivered as designed? | Process evaluation |
| Was it worth the money? | Cost-effectiveness |
| For whom did it work? | Heterogeneity analysis |
| A sharp question names | Example |
|---|---|
| The intervention | 18 months of monthly cash transfers |
| The population | Mothers of children under two |
| The outcome | Height-for-age z-score |
| The comparison | Versus no transfer |
| The timeframe | Measured at 24 months |
| Ask before scoping | To find |
|---|---|
| What decision does this inform? | Whether it is worth doing |
| Who will act on the answer? | Whether it will be used |
| By when do they need it? | The feasible design |
| What answer would change the decision? | The MDE |
| Do not commission when | Instead |
|---|---|
| The programme changes weekly | Stabilise first |
| The question is operational | Process evaluation |
| No comparison group is possible | Contribution analysis |
| The result cannot arrive in time | Rapid learning methods |
| Nobody will act on it | Do not do it |
| Evaluability check | Fails if |
|---|---|
| Clear theory of change | Nobody can state the causal chain |
| Measurable outcomes | The outcome is undefined |
| A credible comparison | Everyone is treated |
| Stable implementation | The design keeps changing |
| Someone will use it | The decision is already made |
| Randomisation gives you | Provided |
|---|---|
| Comparable groups in expectation | Enough units |
| Balance on unobservables | The randomisation is genuine |
| A clean causal claim | Compliance and no spillover |
| A simple estimator | Attrition is low and balanced |
| Balanced by randomisation | Not balanced by any other design |
|---|---|
| Age, income, education | Also balanced by matching |
| Motivation | Unobserved |
| Ability and family support | Unobserved |
| Everything you did not think of | Unobserved |
| In a balance table, check | A warning sign |
|---|---|
| Means on key characteristics | Large differences |
| Baseline outcome levels | Imbalance on the outcome itself |
| J-PAL | Detail |
|---|---|
| Based at | MIT, with regional offices including South Asia |
| Method | Randomised evaluations of anti-poverty programmes |
| Founders | Banerjee, Duflo, and colleagues |
| Recognition | The 2019 Nobel Memorial Prize, shared with Kremer |
| The rise brought | And a critique |
|---|---|
| Credible causal estimates | Narrow questions, chosen for tractability |
| A generalisable method | Weak external validity |
| Individual randomisation | Cluster randomisation | |
|---|---|---|
| Efficiency | Higher | Lower for the same n |
| Spillover risk | High if they interact | Lower |
| Sample needed | Smaller | Larger |
| Fits | Isolated individual treatments | Village or school-level programmes |
| Spillover type | Effect on the estimate |
|---|---|
| Positive — control benefits too | Understates the impact |
| Negative — control loses out | Overstates it |
| Learning across the boundary | Understates |
| Migration between groups | Contaminates both |
| Ethically defensible when | Not when |
|---|---|
| Genuine uncertainty exists | The programme is known to work |
| Resources cannot cover everyone | Denial is arbitrary and avoidable |
| Randomised phase-in is used | Some are permanently excluded |
| A committee has reviewed it | Nobody independent has |
| You cannot randomise when | Quasi-experimental option |
|---|---|
| The programme already ran | DiD, if you have pre-period data |
| Eligibility uses a cutoff | Regression discontinuity |
| Rollout was phased | DiD or staggered adoption designs |
| Take-up was nudged by something external | Instrumental variables |
| DiD removes | It does not remove |
|---|---|
| Fixed differences between groups | Differential trends |
| Common time shocks | Shocks hitting one group only |
| Level differences at baseline | Anything changing differently |
| Read the DiD plot for | Which tells you |
|---|---|
| Pre-period lines | Whether trends were parallel |
| The gap after treatment | The estimated effect |
| Support parallel trends by | Which is |
|---|---|
| Plotting several pre-periods | The standard evidence |
| Testing for pre-trends | Necessary, not sufficient |
| Choosing a similar comparison | Judgement |
| Event-study specification | Showing the timing |
| RD requires | And gives you |
|---|---|
| A sharp, enforced cutoff | A credible local comparison |
| No manipulation of the score | Units alike either side |
| Enough observations near it | Precision |
| A continuous running variable | A local effect only |
| RD gives | And does not give |
|---|---|
| A credible local effect | The effect far from the cutoff |
| A visual, checkable design | Precision without enough observations |
| A defence against selection | Protection if the score is manipulated |
| Matching assumes | Which fails when |
|---|---|
| Selection is on observables only | Motivation drives take-up |
| Common support exists | Treated units have no comparable untreated ones |
| The matching variables are the right ones | Something unmeasured matters |
| A valid instrument must | Which is |
|---|---|
| Predict take-up | Testable |
| Affect the outcome only through take-up | Untestable, and the crux |
| Be as good as random | Argued, not proven |
| Design | Exploits | Key assumption |
|---|---|---|
| Diff-in-differences | Before/after × treated/comparison | Parallel trends |
| Regression discontinuity | A sharp eligibility cutoff | Units similar across the cutoff |
| Matching / PSM | Observed similarity | No unobserved confounders |
| Instrumental variables | An external nudge into take-up | Relevance + exclusion restriction |
| Design | Assumption | How to probe it |
|---|---|---|
| DiD | Parallel trends | Pre-period plots |
| RD | No manipulation; smoothness | Density test at the cutoff |
| Matching | Selection on observables | Cannot be tested |
| IV | Exclusion restriction | Argument; falsification tests |
| Force | Pulls toward |
|---|---|
| Validity | Randomisation |
| Feasibility | Whatever the rollout allows |
| Ethics | Not withholding a known benefit |
| Timeliness | Faster, simpler designs |
| Internal validity is threatened by | Which design handles it |
|---|---|
| Selection bias | Randomisation |
| Confounding trends | DiD |
| Reverse causation | Timing and design |
| Differential attrition | No design — only good execution |
| Design | Counterfactual quality | Credibility |
|---|---|---|
| Randomised controlled trial | Strongest — balances unobservables | Highest |
| Regression discontinuity | Strong, but local to the cutoff | High |
| Difference-in-differences | Good if parallel trends hold | Medium-high |
| Instrumental variables | Depends on a defensible instrument | Conditional |
| Matching / PSM | Only as good as observed variables | Medium |
| Before-after / naive comparison | No real counterfactual | Lowest |
| Design | Counterfactual quality | Main vulnerability |
|---|---|---|
| RCT | Strongest | Execution: attrition, spillover, compliance |
| RD | Strong, local | Manipulation; local only |
| DiD | Good if trends parallel | Differential trends |
| IV | Depends on the instrument | Exclusion restriction |
| Matching | Weakest | Unobservables |
| Feasibility question | If the answer is no |
|---|---|
| Is there still a stage to assign or compare? | The design space has closed |
| Do baseline data exist? | Power costs rise |
| Is the sample large enough? | Reconsider the MDE |
| Will the result arrive in time? | Choose a faster method |
| Phased rollout gives you | At no extra cost |
|---|---|
| A comparison group | Those not yet reached |
| A defensible allocation rule | Order can be randomised |
| An ethical answer | Everyone eventually receives it |
| Baseline opportunity | Before the later phases |
| Quantitative answers | Qualitative answers |
|---|---|
| Did it work? | Why or why not |
| By how much? | For whom, and how |
| How confident are we? | What surprised us |
| On average | What the average conceals |
| If | Then |
|---|---|
| You can assign randomly, ethically | RCT |
| There is a sharp eligibility cutoff | RD |
| Phased rollout and baseline data exist | DiD |
| Only an external nudge to take-up | IV |
| None of these | Reconsider the question |
| Execution problem | Effect |
|---|---|
| High attrition | Groups no longer comparable |
| Differential attrition | Bias, in an unknown direction |
| Spillovers | Estimate biased toward zero, usually |
| Non-compliance | ITT and ToT diverge |
| Broken randomisation | It is now an observational study |
| Underpowered study produces | Which is read as |
|---|---|
| A wide confidence interval | "No significant effect" |
| An imprecise estimate | "The programme did not work" |
| A failure to detect | A finding of no effect |
| Power of 80% means | And implies |
|---|---|
| A real effect is detected 4 times in 5 | It is missed once in 5 |
| Conditional on the assumed effect size | A smaller true effect is missed more often |
| A convention, not a law | Higher power costs sample |
| MDE reframes the question | From | To |
|---|---|---|
| Sample | How many do we need? | What can we see with what we have? |
| Decision | Is it significant? | Is the MDE smaller than a useful effect? |
| Factor | Effect on required sample |
|---|---|
| Smaller expected effect | Much larger sample needed |
| Higher outcome variability | Larger sample needed |
| Higher power target (e.g. 90%) | Larger sample needed |
| Clustered design | Larger sample — driven by number of clusters |
| A good baseline to control for | Smaller sample needed |
| Factor | Direction |
|---|---|
| Smaller expected effect | Much larger sample |
| Higher outcome variability | Larger sample |
| Higher power target | Larger sample |
| Clustering | Larger sample |
| Expected attrition | Larger sample |
| A baseline measurement | Smaller sample needed |
| Power against sample size | Behaves |
|---|---|
| At small n | Rises steeply |
| Approaching 80% | Still rising usefully |
| Clustering costs power because | Measured by |
|---|---|
| People within a cluster are alike | Intra-cluster correlation |
| Each adds less new information | The design effect |
| The effective sample is smaller | n divided by the design effect |
| Power calculation step | What people skip |
|---|---|
| State the MDE | Assuming an implausibly large effect |
| Estimate variability | Using no prior data |
| Account for clustering | Ignoring it entirely |
| Allow for attrition | Planning to the minimum |
| Allow for partial take-up | ITT dilution |
| Type I | Type II | |
|---|---|---|
| Error | Claiming an effect that is not there | Missing one that is |
| Controlled by | Significance level, usually 5% | Power, usually 80% |
| Cost | Scaling a failure | Abandoning something that works |
| Attention received | A great deal | Very little |
| Decide before the study | Because afterwards |
|---|---|
| The primary outcome | You can pick the one that moved |
| How it is defined | Definitions can be adjusted |
| The analysis specification | Specifications can be searched |
| Subgroups of interest | Subgroups can be mined |
| Property | Failing indicator |
|---|---|
| Valid | Measures something adjacent |
| Reliable | Different reading on retest |
| Sensitive | Cannot move within the study period |
| Feasible | Requires a survey you cannot fund |
| A baseline gives you | Which is worth |
|---|---|
| Balance verification | Confidence in the design |
| Variance reduction | A smaller required sample |
| Heterogeneity by baseline status | Who it worked for |
| A fallback if randomisation fails | A DiD option |
| Survey data | Administrative data | |
|---|---|---|
| Fit to the question | Exact | Approximate |
| Cost | High | Low |
| Coverage | Your sample | Everyone in the system |
| Error type | Recall and reporting | System errors, definitions |
| Risk | Courtesy bias | Differential recording by treatment |
| Bias | Guard |
|---|---|
| Social desirability | Neutral wording; private setting |
| Recall error | Short reference periods; anchors |
| Surveyor effects | Training; rotation; matched characteristics |
| Differential measurement | Blind enumerators to treatment status |
| Instrument discipline | Why |
|---|---|
| Pilot every instrument | Comprehension failures are invisible otherwise |
| Neutral wording | Leading questions produce the expected answer |
| Blind enumerators where possible | Expectation shapes recording |
| Same instrument at both rounds | A changed tool destroys the comparison |
| Attrition is dangerous when | Check |
|---|---|
| It is high overall | The rate |
| It differs between groups | The rate by arm |
| It relates to the outcome | Baseline characteristics of leavers |
| It is unreported | Ask |
| Measure too early | Measure too late |
|---|---|
| The effect has not appeared | It may have faded |
| A false null | Comparison areas may have caught up |
| Nutrition, learning, income all lag | Programmes end; effects decay |
| Report the effect as | Not as |
|---|---|
| Percentage points | A coefficient |
| Rupees per household per month | A standardised score alone |
| Additional days of schooling | An effect size |
| Children per thousand | A log point |
| Significant but trivial | Large but uncertain |
|---|---|
| A huge sample; a tiny effect | A small sample; a big point estimate |
| Statistically detectable | Statistically indistinguishable from zero |
| Practically irrelevant | Possibly important |
| Reported as a success | Reported as a null |
| If the interval | Then |
|---|---|
| Excludes zero and includes a useful effect | Reasonable evidence of a worthwhile impact |
| Excludes zero, all values trivial | Real but not worth it |
| Includes zero, is narrow | Probably a true small effect |
| Includes zero, is wide | Uninformative — underpowered |
| ITT | ToT | |
|---|---|---|
| Measures | Effect of being offered | Effect of actually receiving |
| Reflects | Real-world rollout | The intervention itself |
| Size | Smaller if take-up is partial | Larger |
| Use for | Policy decisions on rollout | Understanding the mechanism |
| Heterogeneity analysis | Discipline required |
|---|---|
| By sex, caste, wealth, region | Pre-specify the subgroups |
| By baseline level of the outcome | Beware regression to the mean |
| By implementation quality | Not randomised — descriptive |
| By anything, post hoc | This is searching, not testing |
| A real null | An inconclusive study |
|---|---|
| Well powered | Underpowered |
| Tight interval around zero | Wide interval |
| Implementation verified | Delivery unknown |
| Take-up adequate | Almost nobody participated |
| Read against | Ask |
|---|---|
| Identification | Did the key assumption hold? |
| Balance | Were the groups alike at baseline? |
| Attrition | Low, and balanced? |
| Outcomes | Is this the pre-specified primary one? |
| Take-up | Did anyone actually receive it? |
| Testing many things means | Handle by |
|---|---|
| Some will appear significant by chance | Pre-specifying the primary outcome |
| Subgroups multiply the problem | Limiting and declaring them |
| Specifications can be searched | Reporting robustness |
| Only the winners get reported | Multiple-comparison adjustment |
| Internal validity | External validity |
|---|---|
| Is the answer right here? | Will it hold elsewhere? |
| Established by design | Established by argument and replication |
| An RCT is strongest | An RCT is not automatically strong |
| A property of the study | A property of the claim |
| Transfer breaks because of | Which you can check |
|---|---|
| Different context | Are the preconditions present? |
| Implementation quality | Was the pilot research-grade? |
| Scale effects | Does it work when everyone has it? |
| Different population | Who was in the original sample? |
| Cost-effectiveness lets you | Requires |
|---|---|
| Compare unlike programmes | A common outcome unit |
| Rank options within a budget | Full cost data |
| Argue for scale-up | Costs at scale, not pilot costs |
| Defend a smaller effect | If it is far cheaper |
| Ethical obligation | In an evaluation specifically |
|---|---|
| Informed consent | Including the right to refuse the study, not the programme |
| Do no harm | The study itself, not only the programme |
| Privacy | Small-sample re-identification |
| Ethics review | Before any data collection |
| Pre-registration records | Preventing |
|---|---|
| The hypothesis | Retrofitting one to the result |
| The primary outcome | Selecting the one that moved |
| The analysis plan | Specification searching |
| Subgroups | Post-hoc mining |
| Published record over-represents | Because |
|---|---|
| Large effects | They are more publishable |
| Positive results | Nulls are shelved |
| Novel findings | Replications are harder to publish |
| Small studies with big estimates | Chance plus selection |
| Commissioning step | Skipped at what cost |
|---|---|
| Define the decision | An unused study |
| Write a sharp question | A vague answer |
| Bring the evaluator in early | No credible design left |
| Demand a power calculation | An underpowered null |
| Require pre-registration | Findings that cannot be trusted |
| Ask the evaluator | A weak answer |
|---|---|
| What is the counterfactual? | Vague or absent |
| What must be true to identify the effect? | Cannot state it |
| What is the MDE? | Not calculated |
| How will you handle attrition and spillover? | Not considered |
| Will you pre-register? | Reluctance |
| Impact evaluation can tell you | It cannot tell you |
|---|---|
| Whether this worked here | What to value |
| By how much | How to weigh trade-offs |
| For whom, if powered for it | Whether it holds at scale |
| Compared with a specific alternative | Whether a different design is better |
| Why evidence goes unused | Fix |
|---|---|
| Arrives after the decision | Align the timeline to the decision |
| Framed for academics | Commission a practitioner brief too |
| Nobody owns the follow-through | Name an owner at commissioning |
| Inconvenient nulls are shelved | Agree publication terms in advance |
| Gertler et al. covers | Useful because |
|---|---|
| Every design in this course | Worked, with examples |
| Power calculation, step by step | Practically usable |
| Commissioning and management | Written for practitioners |
| Free to download | No barrier |
| Source | What it offers | Note |
|---|---|---|
| J-PAL | Randomised evaluations, training, evidence reviews | MIT-based; strong South Asia office |
| 3ie | Impact evaluations & systematic reviews of the Global South | Searchable evidence portal |
| World Bank DIME | Methods, guidance, the Gertler et al. handbook | Free handbook & toolkits |
| Campbell Collaboration | Systematic reviews of social interventions | Synthesis across studies |
| AEA RCT Registry | Pre-registered trial protocols | Check what was promised |
| Source | Best for |
|---|---|
| J-PAL | Randomised evaluations; training |
| 3ie | Impact evaluations and systematic reviews of the Global South |
| Campbell Collaboration | Systematic reviews in social policy |
| AEA RCT Registry | Checking pre-registration |
| World Bank DIME | Methods and country work |
| Takeaway | The question it becomes |
|---|---|
| Compared to what? | What is the counterfactual here? |
| Randomisation balances unobservables | Was assignment genuinely random? |
| Every quasi-experiment has an assumption | What is it, and does it hold? |
| Power decides what you can see | What is the MDE? |
| Significance is not importance | What does the interval include? |