The p-value answers a question you probably aren’t asking

Every programme evaluation eventually produces a line like this one: the intervention raised handwashing-with-soap at critical times by 4.1 percentage points (p = 0.03). The room hears "p is under 0.05", writes down "it worked", and moves to the next slide.

That reading skips the question the meeting actually came to answer. A p-value is a statement about how surprising the data would be in a world where the programme did nothing. It is silent on how big the effect is, and a decision about whether to scale a programme is almost entirely a question about how big the effect is.

Two questions, two answers The p-value asks how surprising the data would be if there were no effect, and returns a single probability: p equals 0.03. The confidence interval asks which effect sizes the data is consistent with, and returns a range: a 4.1 percentage point gain with a 95 per cent interval running from 0.3 to 7.9 points. THE P-VALUE If there were no effect, how surprising would this result be? p = 0.03 Data this extreme would turn up 3% of the time in a world where the programme did nothing. It does not say there is a 97% chance the effect is real. Tells you whether to take the result seriously. THE CONFIDENCE INTERVAL Which effect sizes is this data actually consistent with? +0.3 pp +4.1 pp +7.9 pp Ninety-five per cent of intervals built this way would contain the true effect. It is a claim about the method, not a probability about this one interval. Tells you what the result is worth doing about.
The same analysis, read two ways. Only one of them tells a programme manager whether the effect is large enough to be worth the money.

The two questions, precisely

A p-value asks: if the null hypothesis were true, if the programme changed nothing at all, how often would we see data at least this extreme? At p = 0.03, the answer is about three times in a hundred. That is a statement about data, conditional on an assumption. It says nothing about the probability that the assumption is true.

A confidence interval asks a different question: given this data, which values of the effect are compatible with it? A 95% interval of 0.3 to 7.9 percentage points says the data would not be very surprising under any true effect in that range. The "95%" describes the procedure: repeat the study many times and about 95% of the intervals you construct will contain the true value. It is not the probability that this particular interval contains it.

Why the distinction changes the decision

Take the handwashing result above, and suppose the programme costs ₹340 per household reached and is worth scaling only if it raises soap use by at least 2 percentage points. Two different evaluations both report p = 0.03, and they call for opposite decisions.

Evaluation A: +1.0 pp, interval 0.1 to 1.9. The estimate is precise, and the whole interval sits below the 2-point threshold. The programme works a little, and by too little to justify the cost. Do not scale it as designed.
Evaluation B: +4.1 pp, interval 0.3 to 7.9. Same p-value as A, a point estimate four times larger, and far more uncertainty. The data is consistent with an effect large enough to be a flagship result and with one so small the programme is worse value than the alternatives. The honest answer is that this study cannot tell you which, and the next step is a better-powered evaluation rather than a scale-up.

Reported as "p = 0.03" alone, A and B look the same. The intervals show that one is a precise small effect and the other an unresolved question.

The interval in the diagram is Evaluation B's. It is the more common situation in field evaluations in South Asia: cluster designs, modest sample sizes, and intracluster correlation that widens intervals considerably. A result can clear the significance threshold and still be almost uninformative about magnitude.

Four misreadings worth naming

These misreadings are documented. Greenland and colleagues list twenty-five of them in a 2016 paper, written because misinterpretation of these tools remains rampant in the literature.

"p = 0.03 means a 97% chance the effect is real"
The p-value is computed assuming there is no effect. It cannot also be the probability that there is one. Getting from a p-value to a probability about the hypothesis requires a prior, which is what Bayesian analysis supplies and frequentist testing does not.
"p > 0.05 means no effect"
It means the data is compatible with no effect. It is usually also compatible with a substantial effect. An underpowered study will often return a non-significant result even when the programme works, so "no significant difference" and "no difference" mean different things.
"The 95% interval has a 95% chance of containing the truth"
The true effect is a fixed number; it is either inside this interval or it is not. The 95% is a property of the procedure across repetitions, which is why the term of art is confidence rather than probability.
"Two studies disagree because one was significant and one was not"
Two studies with heavily overlapping intervals can easily land on opposite sides of 0.05. Comparing significance verdicts is not the same as comparing effect estimates, and treating it as such manufactures contradictions that the underlying numbers do not support.

What to do instead

A p-value remains useful as one part of a result, and it is insufficient on its own. Three changes to how results are reported help most.

Lead with the interval. Write "+4.1 pp (95% CI: 0.3 to 7.9)" rather than "+4.1 pp, p = 0.03". The interval carries the point estimate, the uncertainty, and the significance verdict together: if it excludes zero, p is below 0.05.
Say what would change the decision before you see the result. If a 2 percentage point gain would not justify the cost, decide that during design. A pre-specified threshold turns an interval into an answer, and stops the effect size being reverse-engineered from whatever the study found.
Describe non-significant results by their interval, not their verdict. "We can rule out effects larger than 3 points" is informative. "No significant effect" is not, and is routinely read as evidence of absence.

The habit worth building is small: whenever a p-value appears without an interval beside it, ask for the interval. In most reports the number is already in the regression output and was simply not carried into the summary.

Further reading

Ronald Wasserstein and Nicole Lazar, "The ASA Statement on p-Values: Context, Process, and Purpose" (The American Statistician, 2016). The professional body's own statement of what a p-value is and is not. Six principles in five journal pages.
Sander Greenland and colleagues, "Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations" (European Journal of Epidemiology, 2016). Twenty-five numbered misinterpretations, each with a correction. The single most useful reference on this page.
Valentin Amrhein, Sander Greenland and Blake McShane, "Scientists rise up against statistical significance" (Nature, 2019). The argument for retiring the significant/non-significant dichotomy, endorsed by more than 800 signatories.