The scene is a composite, but every part of it is ordinary. A programme manager has a spreadsheet open with eleven rows in it, one for each thing her organisation does in the districts she covers. Two columns matter this week: what each line of work cost last year, and how many people it reached. She has been told to propose two closures by the end of the month, and someone senior has suggested, not unreasonably, that the ratio of those two columns should decide which two.
She is not against the ratio. She has spent enough years in this sector to have watched money disappear into work that helped nobody, defended in review meetings by colleagues who could talk fluently about pathways and assumptions and never once about what happened to anyone. Refusing to count is not a principled position, in her experience; more often it is a way of not arriving at the moment when somebody asks what the money bought.
What troubles her is narrower than that. Two of the eleven rows do not behave. The legal aid work reaches around four hundred people a year, which makes its cost per person look indefensible beside the nutrition programme, and it is also the only thing on the list that has ever changed how a government department handles a category of case. The literacy circles are cheap and score well, and when participants are asked afterwards what mattered about them, a striking number describe something that has nothing to do with literacy. Neither of those facts has a column.
This essay is about that gap: what it is, why it is not going to close, and what a person actually making these decisions should do about it. It works through the strongest version of the argument that the arithmetic should decide, sets out where that argument fails, and then proposes a set of working principles that keep the counting while withdrawing its authority to conclude. Those principles are not one organisation's method. They are assembled from several places — health economics, the econometrics of programme evaluation, the capability literature, humanitarian accountability standards, and the practical experience of organisations currently rebuilding how they decide. Where a principle has a good live example, it gets named.
Where the argument starts
The case for letting the arithmetic decide has a canonical statement, and it is worth taking in its strong form rather than the version that circulates as a punchline.
Peter Singer published "Famine, Affluence, and Morality" in Philosophy & Public Affairs in 1972, in the aftermath of the Bengal famine. The argument turns on a case: you pass a shallow pond in which a small child is drowning, and you could wade in and lift her out at the cost of ruining your clothes. Nobody thinks the clothes are a reason not to. If you walk on, you have done something seriously wrong, and this is not a controversial judgement — it is about as firm as moral judgements get.
Singer then asks what separates that child from a child who will die of a preventable illness three thousand kilometres away, whom you could also save at some cost to yourself that is comparably trivial next to a life. Physical distance is not a moral property; nobody defends the view that harm matters less when it is further away. The fact that other people could also help, and are not helping, does not seem to reduce what you could do. That you can see one child and not the other is a fact about the position of your body.
Fifty years of philosophical attention have not produced a comfortable escape. Most people who believe they have rejected the argument have in fact declined to engage it, which is a different thing and, if we are honest about it, a worse one. Two conclusions follow that the development sector should have absorbed long ago and mostly has not.
The first is that indifference to effectiveness is not neutral. If two pieces of work cost the same and one reaches twice as many people with the same benefit, choosing the second is a choice with a cost, and somebody bears it. A sector that treats good intentions as a sufficient account of itself has decided that its own comfort outranks the interests of the people it exists to serve, and anyone who has written or read a donor report knows how much of the prose in one is arranged to prevent that comparison from being made.
The second is that "how much good per rupee" is a legitimate question. Before Singer's argument found an institutional form, asking it out loud in a room of development professionals marked you as someone who had not understood that the work is complicated. The work is complicated. That is not an answer.
What the argument became
Effective altruism turned that ethical claim into a procedure. Express benefit in a common unit, usually the quality-adjusted or disability-adjusted life year. Estimate how many units a given intervention produces. Divide by what it costs. Rank the results and fund downward until the money is gone.
The best public examples of this are GiveWell's cost-effectiveness models, and it is worth saying plainly that they are serious documents. Assumptions are written down. Parameters can be changed by anyone who thinks they are wrong. Uncertainty is stated rather than implied. Very little else in this sector is that willing to show its own reasoning, and organisations that sneer at the models have generally not published anything a critic could argue with.
The procedure has a shape, though, and the shape does more work than any particular number inside it. Four features are worth naming, because each of them later turns into a specific failure:
- The question is agent-centred. It asks what I should do with my money, and the moral interest of the situation belongs to the person spending.
- The comparison is universal. Your rupee is measured against every other rupee that could be spent anywhere on anything.
- Uncertainty is handled by multiplication. Where you do not know something, you estimate it, attach a probability, and carry the product forward.
- The unit of analysis is the intervention: an effect size attributable to a discrete thing that was done, ideally isolated by design.
One objection that does not work
The most common criticism is that you cannot put a number on a human life, and it is the weakest thing said against this framework.
You can, and every institution that allocates health spending already does. When a health ministry sets a threshold above which a treatment will not be reimbursed, it has priced a life-year. When an actuary sets a premium, when a labour ministry sets a compensation schedule for industrial death, when a manager decides which of two districts gets the mobile clinic — a number has been assigned, whether or not anyone writes it down. The real choice is between doing this explicitly, where the figure can be contested, and doing it implicitly, where it cannot. On that point the effective altruists are right, and much of the discomfort their critics express is discomfort at seeing something said out loud that was already happening quietly.
The serious objections are elsewhere, and there are five of them.
Five objections that do work
1. Adding welfare across people destroys the information you needed
Summing benefit across a population treats humanity as a single reservoir of wellbeing to be filled as full as possible. Nobody, however, experiences the sum. If a hundred people each receive a modest benefit and one person is spared a catastrophe, the addition may favour the hundred, and whether it should is a question the addition cannot answer, because answering it requires a fact that the addition has already thrown away: that these are separate lives, and the catastrophe lands entirely inside one of them.
This is Rawls's objection to utilitarianism in A Theory of Justice (Harvard University Press, 1971), where he argues that the utilitarian extension of one person's rational choice to society as a whole "does not take seriously the distinction between persons". Robert Nozick presses the same point in Anarchy, State, and Utopia (Basic Books, 1974), and Bernard Williams develops a related one about integrity in his half of Utilitarianism: For and Against (Cambridge University Press, 1973): a decision procedure that requires you to treat your own deepest commitments as one more input to be traded off has, in a real sense, alienated you from your own actions.
None of this is a technicality about welfare economics. It is the reason that almost nobody, asked directly, will agree that a sufficiently large number of mild headaches relieved outweighs one person losing a leg. Sum-maximisation says it does. Most people conclude that the arithmetic has gone wrong rather than that their judgement has, and they are entitled to that conclusion until someone explains where the missing information went.
2. Poverty is not only a shortfall to be topped up
Singer's framing is one of rescue. There is a deficit; you are in a position to reduce it; the question is how efficiently you can do so. Thomas Pogge's argument in World Poverty and Human Rights (Polity, 2002) is that this misdescribes the situation of the affluent world, which is not a bystander at the pond but a participant in the institutional arrangements — trade rules, debt, resource and borrowing privileges recognised for whoever holds power in a country — that produce and sustain the deprivation. If that is right, the relevant duty is not beneficence but the negative duty not to impose an unjust order, which is a much stronger claim and has a different shape.
Iris Marion Young's Responsibility for Justice (Oxford University Press, 2011) makes a complementary move. She distinguishes the liability model of responsibility, which looks for the party who caused a harm, from what she calls the social connection model, in which many actors are responsible for a structural process none of them individually authored. A great deal of persistent deprivation is structural in exactly Young's sense: it is about who holds land, whose labour is cheap because of the caste they were born into, which settlements the municipal pipe reaches, whose consent is needed before a mine opens. These are not inefficiencies awaiting an optimiser. They are arrangements, and they persist because they suit somebody.
A cost-effectiveness model has a field for cost and a field for benefit and no field for this is a wrong. It can price a remedy. It cannot register an injustice, which in practice means it treats the two as the same kind of thing.
3. The evidence base is shaped by what is easy to measure
This is the failure that shows up most reliably in actual programme decisions, and it is worth being precise about the mechanism.
Cost-effectiveness rankings do not compare all interventions. They compare interventions that have been measured in a form the model can consume. Mass deworming has a widely cited estimate, originating in Edward Miguel and Michael Kremer's Kenya study in Econometrica (2004). Community paralegal work generally has nothing comparable. The difference is not that deworming matters more. It is that deworming produces a countable output on a short timescale in a population you can randomise, and paralegal work produces a change in what people are able to demand from institutions, over years, through mechanisms that resist attribution.
Rank by measurability and you will defund what is hard to measure, and then cite the ranking as evidence that the defunded work was less valuable. James Scott's Seeing Like a State (Yale University Press, 1998) describes the general form of this: institutions make the world legible in the ways their instruments can read, and then act as though the legible part were the whole. The version with a spreadsheet is more dangerous than the version with a cadastral map, because it arrives presenting itself as rigour.
The deworming literature is also, usefully, a caution about the measured cases. In 2015 the International Journal of Epidemiology published a pure replication and a statistical re-analysis of the Kenya data by Alexander Aiken, Calum Davey and colleagues (dyv127), a reply from Joan Hicks, Kremer and Miguel (dyv129), and a long public argument that the sector came to call the worm wars. The Cochrane review of public health deworming programmes (Taylor-Robinson et al., 2019) concludes that mass treatment shows little or no effect on weight, haemoglobin, cognition or school attendance. Reasonable people still disagree about what all this establishes. What it does establish is that a single number carried into a ranking table can be doing far less work than the table implies.
Angus Deaton made the structural version of this point in "Instruments, Randomization, and Learning about Development" (Journal of Economic Literature, 2010), and again with Nancy Cartwright in "Understanding and misunderstanding randomized controlled trials" (Social Science & Medicine, 2018). Their argument is not that trials are bad. It is that a trial establishes an average effect in one population under one implementation, and that carrying it elsewhere requires a theory of why it worked, which the trial itself does not supply. Cartwright and Jeremy Hardie develop this at book length in Evidence-Based Policy: A Practical Guide to Doing It Better (Oxford University Press, 2012), where the operative question is put as: it worked there, will it work here, and what would have to be true of here for that to follow?
There is a concrete way to feel the size of this problem. Consider what a cost-per-outcome model would have said, in 2013, about funding a petition asking the Supreme Court of India to enforce a twenty-year-old statute against manual scavenging. The cost is substantial, the direct beneficiaries in year one number zero, and attribution is hopeless. The following year, in Safai Karamchari Andolan v. Union of India (2014), the Court directed that entering a sewer without protective equipment be treated as a criminal offence even in an emergency, and fixed compensation at ten lakh rupees for each sewer death since 1993. Every family that has since claimed against that figure was claiming under something no model would have funded.
4. Whoever defines the outcome has already made the decision
Someone must decide what counts as an outcome, how to weight a year of life lived in different states of health, whether income gains matter more or less than mortality reductions, and what discount rate applies to a child's future. In the effective altruist procedure these decisions are taken by the analyst, in advance, at a distance, and then presented as the frame within which the beneficiary's interests will be represented. The beneficiary is thoroughly considered and never consulted; their welfare is an input to somebody else's optimisation.
Amartya Sen's objection to welfarism, developed across "Equality of What?" (1979) and Development as Freedom (Oxford University Press, 1999), is that utility is the wrong informational base for judging how someone is doing, partly because preferences adapt: people in long deprivation learn to want less, and a metric built on satisfaction will read that adaptation as wellbeing. What matters instead is capability — what a person is actually able to do and be. Martha Nussbaum's Creating Capabilities (Harvard University Press, 2011) makes the further point that capabilities are plural and not reducible to a common scale without losing what distinguishes them.
Robert Chambers made the practical version of this argument in Whose Reality Counts? Putting the First Last (Intermediate Technology Publications, 1997), which is now nearly thirty years old and still describes most monitoring systems accurately: the professionals' categories decide what is seen, and the people whose lives are being described are asked to confirm the categories rather than to set them.
You can defend the analyst's authority here. Preferences are inconsistent, people are not always the best judges of their own long-run interests, and a well-built metric may serve someone better than their stated preference would. Each of those defences has also been used to justify a great deal of damage done by people entirely certain that they knew better. The problem is not that expert judgement is illegitimate. It is that a procedure which structurally excludes the affected person from defining the good being maximised has made a political decision and is presenting it as a technical one.
5. Expected value is a tool for people who can absorb the variance
Expected-value reasoning is the correct tool when you hold many independent bets and can survive the spread. An insurer can. A research funder running a hundred studies can. A donor allocating across forty grants can.
A household cannot, and neither can a village. If an intervention carries a seventy per cent chance of substantial benefit and a thirty per cent chance of real harm, the expected value may be firmly positive and the thirty per cent is still somebody's actual year, borne by someone who did not choose to join your portfolio.
The same structure scales. Nick Bostrom's "Pascal's mugging" (Analysis, 2009) shows how expected-value maximisation can be driven to arbitrary conclusions by attaching tiny probabilities to enormous payoffs, and the practical form of the problem is a decision procedure that instructs you to be suspicious of intuitions resisting the calculation. Given enough time and confidence, the calculation will eventually favour something disastrous, and the intuition that might have stopped it has already been reclassified as a bias.
Whether this is what happened to effective altruism's most prominent institutional collapse is genuinely contested and this essay does not need to settle it. The facts are that FTX, whose founder was among the movement's largest funders, failed in November 2022, and that he was convicted on seven counts of fraud and conspiracy in 2023. The argument that followed inside the movement was substantially about whether its own habits of reasoning — high tolerance for variance, deference to explicit calculation over unease — had made the association more likely. Something in that structure was missing a way of saying: stop, this is not your bet to place.
What survives the objections
Five objections is enough to reject the claim that the arithmetic decides. It is not enough to reject the arithmetic, and the sector's habitual response — reciting the objections and then carrying on without counting anything — is worse than the thing it objects to, because it keeps the failure and loses the discipline.
What follows is an attempt at the harder thing: a set of principles that keep the counting and withdraw its authority to conclude. They are not novel, and no single organisation owns them. Some are codified in the humanitarian sector's Core Humanitarian Standard on Quality and Accountability, whose commitments on participation and complaints are the closest thing the field has to an agreed floor. Some come out of the state capability literature, particularly Matt Andrews, Lant Pritchett and Michael Woolcock's Building State Capability (Oxford University Press, 2017) and the problem-driven iterative adaptation approach it describes. Some are drawn from the evaluation methodologists cited above.
And one useful live example is worth naming, because it shows what the principles look like when an organisation with real scale actually rewires itself around them. In August 2026 CARE published an account of a new evidence strategy, written by Emily Janoch. Among its commitments, five matter here: benchmarking programme design against a library of more than eight thousand studies with partners including J-PAL, CEGA and IDinsight; building what it describes as the first live cost-per-outcome dashboard of any major international NGO; treating community feedback as a design input rather than a reporting category; tracking individual participant trajectories rather than headline totals across the more than a hundred countries it works in; and investing in new experimental research only where government or research partners see a credible path to adoption at scale.
Note what that is and is not. It is one organisation's operating decision, published as a strategy statement rather than as evidence that the strategy works, and the honest thing to say about it in 2026 is that it is a promising design whose results do not exist yet. It is useful here because it gives several of the principles below a concrete shape, and because its fourth commitment — cost per outcome, tracked live — is the same arithmetic operation as dollars per QALY. Anyone who finds one of these approaches congenial and the other troubling owes themselves an account of the difference that is better than "one counts and the other does not". Both count. The differences are structural, and they are what the principles are about.
Eight working principles
1. Count, and publish what you assumed
The first principle is the one the effective altruists got right, and it should be conceded without qualification. Produce the number. Write down the assumptions that generated it — the effect size you borrowed, where you borrowed it from, the population you are applying it to, the discount rate, the overhead treatment. Publish them in a form somebody outside your organisation could argue with.
The reason is not that the number will be correct. It is that an explicit assumption can be contested and an implicit one cannot, and almost every allocation decision in this sector rests on implicit ones. GiveWell's models remain the standard to beat here, and the fair criticism of most large NGOs is not that their arithmetic is wrong but that there is nothing published to check.
How it fails: published assumptions create an illusion of settledness. A model with forty documented parameters looks more trustworthy than a judgement with none, even when the forty parameters are each a guess and the judgement was made by someone with twenty years in the district.
2. Use arithmetic to eliminate, not to select
Cost-effectiveness analysis is a strong argument against your worst option and a weak argument for your best. If one line of work costs four times as much per person reached as another doing something comparable, and nobody in the room can say why, that is a finding and you should act on it. If two options sit within a factor of two of each other, the model is not telling you anything: the gap is inside the error bars of your own assumptions, and choosing on it is false precision.
This follows directly from Deaton and Cartwright's argument about external validity. An effect size transported from another context carries an uncertainty that is usually larger than the differences you are trying to adjudicate, and which is almost never propagated into the ranking.
How it fails: "within a factor of two" is a judgement call, and a manager under pressure to justify a decision will find the factor conveniently large or small.
3. Gate rather than multiply when the variance falls on someone else
This is the structural heart of the matter. Decide first whether you are holding a portfolio or making a single bet, because the correct epistemics differ and the difference is not a matter of taste.
A donor allocating across two dozen grants genuinely holds a portfolio: variance averages out, and expected value is the right tool. An implementing team working in four blocks of one district is not holding a portfolio. It is making one bet with somebody else's year, on people who did not choose to be in it. There, uncertainty should be handled by sequential gates — each one has to clear before the next is attempted — rather than by conversion into a coefficient. If the community says it is not working, you do not offset that against a strong external benchmark and proceed; you have failed a gate, and the process stops.
Staged designs of this kind are what the five commitments summarised above amount to in practice, and they are the same logic as the iterative, problem-driven approach in Building State Capability, where the unit of progress is a small authorised experiment that either survives contact with the problem or is abandoned before it is scaled.
How it fails: gates can be made nominal. A gate that has never once stopped anything is not a gate, and the only way to know is to count how often each one has been failed.
4. Put the affected person's judgement inside the measure, not after it
In most monitoring systems, beneficiary feedback is collected downstream of impact measurement and reported separately, under a heading like satisfaction or accountability. It functions as reputational data. It does not change what the impact number says.
Moving it upstream — so that whether the work is judged to have worked is partly constituted by the people it was done to — is a change in sequencing rather than in budget, and it is probably the highest-value item on this list. It is what the Core Humanitarian Standard's participation and feedback commitments are for, and it is third in the list above, placed deliberately before the cost calculation rather than after it.
This does not resolve the problem identified in objection four. Whoever designs the instrument still chooses the questions and still decides what counts as an adequately positive answer. What it does is move the affected person from being an object the metric describes to being one of the parties whose judgement the metric incorporates, which is a structural answer rather than a rhetorical one.
How it fails: feedback collected by the person delivering the programme, from someone who depends on it, is not feedback. This is the oldest problem in participatory practice and Chambers described it in 1997; the mitigations are independent collection, anonymity, and asking questions whose answers could actually embarrass somebody.
5. Use a counterfactual you can act on
The effective altruist counterfactual is global: this rupee against every other possible rupee, anywhere, for any purpose. For an individual donor with an unrestricted wallet that comparison is real. For almost any organisation it is a fiction. A field team in Bangladesh cannot be redeployed to distribute bed nets in Chad; the staff have contracts, languages, relationships, and fifteen years of knowing which union council chairman returns calls. Treating that as fungible capital misdescribes what it is.
The workable substitute is a question about durability: does this outlive us? Will a government department, a market actor or a peer organisation carry it once the grant closes? That is a real counterfactual, and an organisation can act on the answer.
How it fails: badly enough that it needs its own section below.
6. Make the unit a trajectory as well as an intervention
The effective altruist unit of analysis is the intervention: an effect attributable to a discrete thing that was done, ideally isolated by design, because isolation is what makes the effect comparable across contexts.
Most development outcomes worth having are not produced that way. A girl stays in school because of a scholarship, a road, a functioning toilet, a teacher who noticed, and a mother who held a position in an argument at home; attributing the outcome to any one of these is not merely difficult, it may be the wrong question. Following the person over time — which is what CARE means by moving from headline statistics to trajectories, and what Sen's capability framing implies about the right object of measurement — gives up clean attribution and gains a description of what actually happened.
How it fails: you lose precisely the property that let you compare your work to anyone else's. Trajectory data is also far more revealing about individuals than aggregate reporting, which is the subject of a later section.
7. Keep a protected line for work with no adoption pathway
Principle five, applied without exception, becomes a rule that you will do what powerful institutions already want done and not do what they resist. Some deliberate exception is required, and it should be budgeted rather than left to conscience.
The Indian case law makes the argument better than an abstraction can. Manual scavenging remained widespread for two decades after the 1993 prohibition precisely because no state government wished to enforce it, which is why litigation became the instrument. The mid-day meal became a universal entitlement because the Supreme Court directed it in PUCL v. Union of India in November 2001, over the objections of states that did not want the fiscal exposure. In Samatha v. State of Andhra Pradesh (1997) the Court held that leasing land in Scheduled Areas to non-tribal mining companies was void, and no ministry was waiting to adopt that finding. In Orissa Mining Corporation (2013) the Court held that the Gram Sabha, rather than a ministry, would decide whether bauxite mining proceeded at Niyamgiri. Each of these is a case where the absence of an adoption pathway was the reason to act. All four are in the Development Law Docket, with the sources.
How it fails: a protected line with no accountability becomes the place where favoured projects go to avoid scrutiny. It needs its own standard of evidence, just a different one — did this change what an institution is obliged to do — rather than exemption from evidence.
8. Ask what would have to be true for this to be wrong
Neither the multiplying nor the gating structure asks the question that protects you from your own model. Expected-value reasoning cannot ask it, having already priced its uncertainty and moved on. A gated process asks a version of it at each gate, which is better, but the gates were specified in advance and can only catch failures somebody anticipated.
This is different from sensitivity analysis. Sensitivity analysis varies the parameters inside your model; this asks what would have to be true outside it, in the part of the world the model does not represent at all. Asked of a nutrition programme, it produces answers of the form: if the reason these children are undernourished is that their mothers have no bargaining power over household food, then our supplement will show an effect that disappears the month we leave. No cost-per-outcome figure contains that.
Gary Klein's premortem, set out in Harvard Business Review in September 2007, is the cheapest practical version: before starting, assume the project has failed completely, and have the team write down why. It takes an hour and it surfaces objections that hierarchy would otherwise suppress.
How it fails: it becomes a ritual with a template, filed and never read. The test of whether it is real is whether a premortem has ever changed a design.
Four ways this apparatus still goes wrong
An essay that stopped at the principles would be doing the reader a disservice, because they have failure modes of their own and at least three are serious.
The dashboard becomes the target
Goodhart's law is not an aphorism; it is an observed regularity in every system that has tried this. Marilyn Strathern's formulation, from a 1997 paper in European Review on audit in British universities ("'Improving ratings'"), is the one usually quoted: when a measure becomes a target, it ceases to be a good measure.
The moment cost per outcome is live and visible, every manager in the system has an interest in the numerator being large and the denominator small, and the fastest route there is not better programming. It is outcome selection — a quiet drift toward outcomes that are cheap to produce and easy to record. Nobody needs to falsify anything; the system simply becomes better at producing what the dashboard rewards, and the annual report describes this as improved efficiency. The essays collected in Rosalind Eyben, Irene Guijt, Chris Roche and Cathy Shutt's The Politics of Evidence and Results in International Development (Practical Action Publishing, 2015) document the pattern across a decade of results-based management in aid.
The defences are to change what is counted often enough that nobody can optimise against it, and to keep at least one qualitative channel that is never scored. Both are expensive and both are the first things cut.
A live participant-level dashboard is also a database
A unified global system tracking individuals across a hundred countries is, from a slightly different angle, a detailed record of very poor people's circumstances, held by a foreign organisation, in jurisdictions with uneven data-protection law, for purposes those people did not specify and mostly cannot revoke.
This is not a hypothetical harm. In June 2021 Human Rights Watch reported that UNHCR had collected biometric data from Rohingya refugees in Bangladesh and shared it with the Bangladeshi government, which passed it to Myanmar — the state they had fled — without their informed consent. Beneficiary databases have also been demanded by governments as a condition of registration, and they have been breached. The person whose nutrition status, household income and location sit on the dashboard is the person least able to object and with the most to lose.
None of this makes such systems wrong. It does mean that consent, retention limits, minimisation, and a written plan for what happens when a hostile authority asks for the data are not compliance paperwork appended at the end. They are part of whether the approach is ethical at all, and a strategy that specifies its measurement architecture in detail while leaving data governance to an annexe has not finished thinking.
"Credible adoption pathway" is conservatism wearing a metric
This is the one worth pressing hardest, because it is the principle that sounds most obviously sensible.
If the test for scaling is whether a government, a market or a peer will adopt the work, then over time you will systematically do what powerful institutions already want and systematically not do what they resist. That is a coherent theory of change — it is more or less the theory behind most successful public-service delivery reform — and it is also a description of how an organisation stops being useful for anything except delivery.
The four judgments above are the counter-case, and the pattern in them is consistent: the absence of an adoption pathway was the reason the work was necessary, not a signal to stand down. The failure here is particularly quiet, because work that fails an adoption gate never generates a number that anyone has to explain. It simply does not happen, and nothing in the reporting records that it did not.
Some goods do not survive division
Cost per outcome requires outcomes to be countable and separable, and some are not.
When a woman speaks in a village meeting for the first time, the outcome is not the speech. It is a change in what she and everyone in the room now understand to be possible. You can proxy it — attendance, a self-efficacy scale, whether she spoke again three months later — and each proxy is a real measurement of something adjacent to the thing. None is the thing.
Nussbaum's argument about the plurality of capabilities is the formal version of this: goods of this kind are constitutive rather than instrumental. They are part of what a good life is, not a means to it, and dividing them by money yields a number whose units do not mean anything. This is not mysticism about the immeasurable. It is a claim about what division does to certain quantities, and it implies a practical rule: when you catch yourself computing a cost per unit of dignity, the error is in the operation, not in the estimate.
Tuesday morning
Back to the spreadsheet, and to what actually changes on the Tuesday after reading an essay like this one.
Run the ratio on all eleven rows and publish the assumptions internally. Not to rank them. To find the row that is four times the cost of a comparable row for reasons nobody can articulate, because that row is a real finding and it is the only thing in the exercise the arithmetic is competent to tell you.
Before reading what the ratio says, work out what it excludes. Every cost-per-outcome figure has a denominator somebody defined and a numerator somebody bounded. In this case the legal aid programme's four hundred people are not the outcome; the changed departmental practice is, and it is not in either column. Write down what is missing before you look at the ordering, because afterwards you will rationalise.
Decide explicitly whether this is a portfolio decision or a single bet. If the organisation is genuinely spreading risk across many independent pieces of work, expected value is legitimate. If two districts are carrying the whole consequence, it is not, and the structure should be gates.
Move the community feedback upstream of the closure decision, not after it. If it arrives after the number is computed it is public relations; if it arrives before, it is measurement. This costs a fortnight and no money.
Run a premortem on the two closures. An hour, the whole team, the assumption that in three years the decision was obviously wrong, and everyone writing down why.
Protect one line that no adoption pathway supports. Small. But if every gate in the system rewards work that powerful institutions already want, the drift is invisible from inside, because everything you do will keep passing.
What to keep from each
Singer is right that distance is not a moral property, that ignoring effectiveness has victims, and that refusing to count is usually avoidance presented as depth. Those are permanent contributions and the sector is better for having been made to hear them.
He is wrong that the calculation, once performed, settles the matter — because the calculation has already made choices about what counts, about who decides what counts, and about whether a wrong is a different kind of thing from a cost. Those choices are the substance of the ethical question rather than preliminaries to it, and burying them inside a method does not make them go away.
The principles in the middle of this essay answer several of the objections structurally: they gate instead of multiplying, they place the affected person's judgement inside the measure, and they replace an unusable global counterfactual with one an organisation can act on. They also introduce a conservatism at the durability test that would have excluded most of what development law in India actually achieved, and any live participant-level measurement system inherits a set of data-governance obligations that are part of its ethics rather than an appendix to them.
The position worth holding is narrower than a synthesis and firmer than a compromise: keep the arithmetic, refuse its authority. Count everything you can count. Publish the assumptions. Let the numbers eliminate what deserves to be eliminated. And when the number and your sense that something is wrong disagree, do not assume by default that the number wins — because the number was built by somebody who made choices, and one of those choices may have been to leave out the thing you are noticing.
The manager in the composite already knows this. The two columns are real and she should use them. So is the third thing she knows, and her job is not to force it into a column. It is to carry all three into the room and be able to say out loud which of them is deciding, and why.
Sources
Everything cited above, in the order it appears. Links go to the source where a stable one exists; books are given with publisher and year.
- Peter Singer, "Famine, Affluence, and Morality", Philosophy & Public Affairs 1(3), 1972, pp. 229–243.
- GiveWell, cost-effectiveness models and selection criteria.
- John Rawls, A Theory of Justice, Harvard University Press, 1971 (the separateness of persons argument is at §5–6).
- Robert Nozick, Anarchy, State, and Utopia, Basic Books, 1974.
- J. J. C. Smart and Bernard Williams, Utilitarianism: For and Against, Cambridge University Press, 1973 (Williams on integrity).
- Thomas Pogge, World Poverty and Human Rights, Polity, 2002.
- Iris Marion Young, Responsibility for Justice, Oxford University Press, 2011.
- Edward Miguel and Michael Kremer, "Worms: Identifying Impacts on Education and Health in the Presence of Treatment Externalities", Econometrica 72(1), 2004. doi:10.1111/j.1468-0262.2004.00481.x
- Alexander Aiken, Calum Davey, James Hargreaves and Richard Hayes, "Re-analysis of health and educational impacts of a school-based deworming programme in western Kenya: a pure replication", International Journal of Epidemiology 44(5), 2015. doi:10.1093/ije/dyv127
- Joan Hamory Hicks, Michael Kremer and Edward Miguel, "Commentary: Deworming externalities and schooling impacts in Kenya", International Journal of Epidemiology 44(5), 2015. doi:10.1093/ije/dyv129
- David Taylor-Robinson, Nicola Maayan, Sarah Donegan, Marty Chaplin and Paul Garner, "Public health deworming programmes for soil-transmitted helminths in children living in endemic areas", Cochrane Database of Systematic Reviews, 2019. doi:10.1002/14651858.CD000371.pub7
- James C. Scott, Seeing Like a State, Yale University Press, 1998.
- Angus Deaton, "Instruments, Randomization, and Learning about Development", Journal of Economic Literature 48(2), 2010. doi:10.1257/jel.48.2.424
- Angus Deaton and Nancy Cartwright, "Understanding and misunderstanding randomized controlled trials", Social Science & Medicine 210, 2018. doi:10.1016/j.socscimed.2017.12.005
- Nancy Cartwright and Jeremy Hardie, Evidence-Based Policy: A Practical Guide to Doing It Better, Oxford University Press, 2012.
- Amartya Sen, "Equality of What?" (Tanner Lecture, 1979) and Development as Freedom, Oxford University Press, 1999.
- Martha Nussbaum, Creating Capabilities: The Human Development Approach, Harvard University Press, 2011.
- Robert Chambers, Whose Reality Counts? Putting the First Last, Intermediate Technology Publications, 1997.
- Nick Bostrom, "Pascal's mugging", Analysis 69(3), 2009. doi:10.1093/analys/anp062
- United States Attorney's Office, Southern District of New York, press release on the FTX verdict, 2023; and the EA Forum's own collected discussion of the collapse.
- Core Humanitarian Standard on Quality and Accountability.
- Matt Andrews, Lant Pritchett and Michael Woolcock, Building State Capability, Oxford University Press, 2017; and the problem-driven iterative adaptation materials at Harvard's Center for International Development.
- Emily Janoch, "Built on proof: How CARE is using data to maximize impact per dollar", CARE, 24 August 2026.
- Gary Klein, "Performing a Project Premortem", Harvard Business Review, September 2007.
- Marilyn Strathern, "'Improving ratings': audit in the British University system", European Review 5(3), 1997. doi:10.1002/(SICI)1234-981X(199707)5:3<305::AID-EURO184>3.0.CO;2-4
- Rosalind Eyben, Irene Guijt, Chris Roche and Cathy Shutt (eds), The Politics of Evidence and Results in International Development, Practical Action Publishing, 2015.
- Human Rights Watch, "UN Shared Rohingya Data Without Informed Consent", 15 June 2021.
- Judgments cited — Samatha v. State of Andhra Pradesh (1997), PUCL v. Union of India WP(C) 196/2001 (mid-day meal directions, 28 November 2001), Orissa Mining Corporation v. MoEF (2013) and Safai Karamchari Andolan v. Union of India (2014) — with sources and status in the ImpactMojo Development Law Docket.