A programme director asks for an impact evaluation eighteen months after the programme has rolled out to every district it was ever going to reach. The evaluation team works through the guidance in order: rationale, scope, theory of change, budget, questions. At step six they reach the counterfactual and discover there is nobody left to compare against. The decision that ended the evaluation was taken a year and a half earlier, by someone who was not thinking about evaluation at all.
The standard guidance for designing an impact evaluation lists eight steps. They are the right eight. Drawn as a numbered sequence, though, they suggest that the work proceeds in that order, and it does not. Two of the steps are decisions about the world rather than decisions on paper, and they have to be settled before things that appear earlier in the list.
The eight steps
- Define rationale and objectives. Why evaluate at all, what knowledge gap this fills, whether the question is about one context or several.
- Specify the programme. Scope, target population, anticipated outcomes.
- Build a theory of change. Activities to short, medium and long-term outcomes, with assumptions made explicit.
- Identify timeline, resources and budget. Staff time, consultancy, data collection, and where the money comes from.
- Develop measurable evaluation questions. Specific and answerable, aligned to the theory of change.
- Define the counterfactual and data requirements. Baseline and follow-up for treatment and comparison, standardised indicators.
- Identify the methodology. Experimental, quasi-experimental, or non-experimental.
- Plan dissemination and reporting. How findings reach the people who could act on them.
Nothing on that list is wrong. What follows is about the order.
The counterfactual is decided long before you reach it
A counterfactual is a claim about what would have happened to the same people, in the same period, without the programme. You cannot observe it. You can only approximate it with a group whose experience stands in for that unobserved case, and whether such a group exists is decided by how the programme is rolled out.
That decision is usually made by operations, on operational grounds, long before an evaluator is in the room. Once a programme has saturated its target population, the honest set of remaining options is narrow: compare across places that differ for reasons you cannot fully name, compare across time and hope nothing else changed, or accept that you are measuring implementation rather than impact.
So the practical question at step one is not only "why are we evaluating". It is "what will the comparison be, and does the rollout still allow it". Ask that in the first meeting and it will change the rollout plan, which is the point. Ask it at step six and you are documenting a constraint rather than choosing one.
What that changes about the earlier steps
Rationale and objectives
Two very different questions hide under "does this work". One is whether the intervention produces the outcome, which needs a counterfactual. The other is whether the intervention was delivered as designed, which needs monitoring data and no comparison group at all. Programmes frequently commission the first and need the second. If nobody has yet established that the training happened, that the transfers arrived, or that the staff were in post, an impact evaluation will produce a null result you cannot interpret: you will not know whether the theory failed or the delivery did.
Specifying the programme
The unit at which a programme varies is the unit at which you can evaluate it. A scheme delivered through the panchayat varies at panchayat level, whatever the individual-level data you collect, and your effective sample is the number of panchayats rather than the number of households. This is where sample size calculations quietly fall apart, and it is a specification question, settled at step two, not a statistics question settled later.
Theory of change
A theory of change written for a funder tends to list what the programme intends. A theory of change written for an evaluation has to include the steps that could fail and the assumptions that could turn out false, because those are what the evaluation is for. A diagram in which every arrow is a step the programme controls has no room for the evaluation to find anything.
Timeline and budget
Baseline data has to be collected before the programme reaches anyone. That single requirement moves evaluation budgeting from the second year of a project to the first, and it is the most common reason a well-designed evaluation is not affordable by the time someone asks for it.
It is worth checking what already exists before assuming a survey. India is unusually well served here. NSS rounds, NFHS, ASER's annual district-level learning data, and the administrative systems behind MGNREGA and the PDS all provide repeated measurement that predates most projects. Existing data rarely matches the outcome you would have chosen, and that is a real cost. It also rarely costs a crore.
Methodology follows the counterfactual
Step seven is often where the conversation starts, usually as an argument about whether a randomised trial is appropriate. The order is backwards. A method is a way of constructing a comparison; which methods are available depends on the comparison the world has left you.
Three situations, and what each permits:
- The programme cannot reach everyone at once. Almost always true, and the most useful fact in evaluation design. If the order of the rollout can be randomised, everyone is served and the early group can be compared with the not-yet group. Nobody is denied anything they were otherwise going to receive, which is usually the ethical objection to randomisation and is answered by the design rather than argued with.
- Eligibility is set by a threshold. Below a certain landholding, income, or score, you qualify; above it you do not. Households just either side of the line resemble each other closely, and the discontinuity can carry the comparison. This applies to a great deal of Indian social protection, where thresholds are how targeting is done.
- The programme is a universal entitlement. MGNREGA guarantees work to any rural household that asks. There is no group without access, so "does MGNREGA work" is not answerable in this frame. What is answerable is which delivery systems make the entitlement real, and that is a different and often more useful question.
The third case deserves emphasis, because treating it as a failure wastes the most policy-relevant evidence available. When access is universal, variation in implementation is the thing to study, and the studies that have most changed Indian delivery have done exactly that.
Dissemination has to include the null result
Step eight is usually written on the assumption that the finding will be positive. Presentations, policy briefs, a workshop.
Plan for the other outcome while the design is still being agreed, because that is the only moment when everyone can be honest about it. Two questions are worth putting in writing. What finding would cause this programme to change, and who has the authority to make that change? If no answer exists to the second, the evaluation is being commissioned for a reason other than learning, and it is better to know that at the start.
A null result is also worth publishing on its own terms. The evidence base for South Asian development is skewed by the fact that disappointing findings tend to stay in draft, which means the next team designing a similar programme sees only the successes.
A short checklist
Before the theory of change workshop, not after:
- Who will not receive this, and when? If the answer is "nobody", the design conversation is already over.
- At what unit does the programme vary, and how many of those units are there?
- What measurement already exists for this population, and how far back does it go?
- Can the rollout order be randomised without delaying anyone who would otherwise have been served sooner?
- What finding would change the programme, and who decides?
The eight steps are sound as a list of what an evaluation design has to cover. The sequence is where teams lose evaluations they could have run, and the counterfactual is the step that has to move earliest.