Dividing Good by Money: Effective Altruism, Cost per Outcome, and What Arithmetic Cannot Settle

A programme officer in Dhaka has a spreadsheet open. One column is what each of her eleven interventions cost last year. The next is how many people it reached. She has been asked to cut two. Somewhere in the building, someone has suggested that the arithmetic should decide.

She is not opposed to arithmetic. She has watched money go into programmes that did nothing for anyone, defended in meetings by people who spoke about theories of change and never about results. She knows that refusing to count is not a moral position. It is usually just a way of avoiding the moment when someone asks what the money bought.

But she also knows something about the two columns she cannot fit into them. The legal-aid programme reaches four hundred people a year at a cost per person that looks indefensible next to the nutrition work. It is also the only thing the organisation does that has ever changed what a government department does. The literacy circles look cheap and perform well. They also happen to be the thing that participants, when asked what mattered most, keep mentioning for reasons that have nothing to do with literacy.

This essay is about the gap between those two kinds of knowledge, and about two frameworks that try to close it. One is Peter Singer's, and the movement that grew from it. The other is the strategic framework CARE has been describing, built on what it calls cost per outcome. They are usually treated as opposites. They are not. Understanding why they are not is the whole point.

The argument that will not go away

Start with Singer, and start by taking him seriously, because the version of him that circulates in development circles is usually a caricature that is easy to dismiss and therefore useless.

In 1972, in Famine, Affluence and Morality, Singer asked you to imagine walking past a shallow pond in which a child is drowning. Wading in will ruin your clothes. Nobody thinks the clothes are a reason. You are obliged to save the child, and if you walk on you have done something wrong.

Then he asked what the morally relevant difference is between that child and a child dying of a preventable disease three thousand kilometres away, whom you could also save at comparable cost to yourself. Distance is not a moral property. Whether other people could also help but do not does not diminish what you could do. That you cannot see the second child is a fact about your eyes, not about the child.

The argument is uncomfortable precisely because the obvious escapes do not work. It has survived fifty years of attempts. Most people who say they reject it have not rejected it; they have declined to think about it, which is a different thing and a worse one.

Two things follow, and both are correct.

The first is that indifference to effectiveness has victims. If two programmes cost the same and one helps twice as many people, choosing the other is not neutral. Somebody bears that choice. A sector that treats "we meant well" as a sufficient account of itself is a sector that has decided its own comfort matters more than the people it is funded to serve. Anyone who has sat through a donor report knows how much of the writing is engineered to prevent exactly this comparison.

The second is that the question "how much good per rupee" is a legitimate question. Before the effective altruists made it respectable, asking it in a room full of development professionals was faintly vulgar, the mark of someone who did not understand that this work is complicated. It is complicated. That is not a reason to stop asking.

So: the framework earns its place. What it does not earn is the authority it claims next.

What the framework actually asks you to do

Effective altruism turned Singer's argument into a procedure. Express benefit in a common unit — typically the quality-adjusted or disability-adjusted life year. Estimate how many units an intervention produces. Divide by cost. Rank. Fund from the top of the list until the money runs out.

GiveWell's published cost-effectiveness models are the most careful public example of this, and they are genuinely impressive documents: assumptions exposed, parameters adjustable, uncertainty stated. Anyone who has tried to build one knows how much work that transparency costs. The sector has produced very little else that is so willing to show its own reasoning.

The procedure also has a specific shape, and the shape matters more than any individual number in it.

  • It is agent-centred. The question is what I should do with my money. The moral drama belongs to the donor.
  • It is comparative across everything. Your rupee competes with every other possible rupee, everywhere.
  • It resolves uncertainty by multiplying through it. Where you do not know, you estimate, attach a probability, and take the expected value.
  • It treats the intervention as the unit. What is measured is an effect size attributable to a thing that was done.

Each of those four is a choice. Each has consequences that show up in real programmes.

What is actually wrong with it

Not everything people say is wrong with it is wrong with it. Let me clear one objection out of the way first, because it is the most common and the weakest.

"You cannot put a number on a human life." You can, and everyone does. A health ministry setting a treatment threshold is doing it. An insurer is doing it. A programme officer choosing between two districts is doing it. The choice is not between valuing life numerically and refusing to; it is between doing it explicitly, where the assumption can be argued with, and doing it implicitly, where it cannot. On this point the effective altruists are simply right, and the discomfort their critics feel is mostly discomfort at seeing something stated that was already happening.

Here is what is actually wrong.

1. Aggregation erases the separateness of persons

Adding welfare across people treats humanity as one large container of wellbeing to be filled. But there is no one who experiences the sum. If a hundred people each get a small benefit and one person is spared a catastrophe, the arithmetic may favour the hundred. Whether it should is a question the arithmetic cannot answer, because the answer depends on a fact the addition destroys: that these are separate lives, and the catastrophe falls entirely on one of them.

Rawls made this objection to utilitarianism generally, and it is not a technicality. It is why most people, asked directly, will not agree that any number of mild headaches relieved outweighs one person losing a leg. Sum-maximisation says otherwise. Most of us think the arithmetic has gone wrong, not our intuition.

2. The measurement gradient decides more than the ethics does

This is the failure that shows up most reliably in practice.

Cost-effectiveness rankings do not compare all interventions. They compare all interventions that have been measured in a form the model can eat. Deworming has a cost-effectiveness estimate. Community paralegal work generally does not. That is not because deworming is more valuable. It is because deworming produces a countable output on a short timescale in a population you can randomise, and paralegal work produces a change in what people can demand from institutions, over years, through mechanisms that resist attribution.

Rank by measurability and you will systematically defund the things that are hard to measure, and then you will point at the ranking as evidence that those things were less valuable. The evidence base becomes self-confirming. This is the streetlight effect with a spreadsheet attached, and it is more dangerous than the ordinary kind because it arrives looking like rigour.

Consider what a cost-per-outcome model would have said, in 2013, about a petition asking the Supreme Court of India to enforce a twenty-year-old law against manual scavenging. Cost: substantial. Direct beneficiaries in year one: none. Attribution: hopeless. The following year the Court held in Safai Karamchari Andolan that entering a sewer without safety equipment should be a crime even in an emergency, and set compensation at ten lakh rupees for each death. Every family that has since invoked that number invoked something no model would have funded.

3. The person holding the ruler is not the person living the outcome

Someone has to decide what counts as an outcome, how to weight a year of life at different levels of health, whether income gains matter more than mortality reductions, what discount rate applies to a child's future.

In effective altruism these decisions are made by the analyst, in advance, at a distance, and are then presented as the framework within which the beneficiary's interests will be represented. The beneficiary is thoroughly considered and never consulted. Their welfare is an input to someone else's optimisation.

You can defend this. Preferences are inconsistent, people are not always the best judges of their own long-run interests, and a well-constructed metric may serve someone better than their stated preference would. Every one of those defences has also been used to justify a great deal of harm done by people certain they knew better. The point is not that expert judgement is illegitimate. It is that a framework which structurally excludes the affected person from defining the good being maximised has made a political choice and is presenting it as a technical one.

4. Politics gets laundered into technique

Cost-effectiveness analysis treats deprivation as a resource-allocation problem: there is not enough good in the world, so distribute the available good optimally.

A great deal of persistent deprivation is not that. It is a political fact. It is about who holds land, whose labour is cheap because of what caste they were born into, which settlements the municipal pipe reaches, whose consent is required before a mine opens. These are not inefficiencies. They are arrangements, maintained because they benefit somebody.

A model has no variable for this is unjust. It can price a remedy but not a wrong. So when a community in Niyamgiri refused a bauxite project on land they hold sacred, and the Supreme Court in Orissa Mining Corporation held that the Gram Sabha, not a ministry, would decide — there was no cost-effectiveness case for any of it. The cost per outcome of establishing that a village council can veto a mine is not a computable quantity. It is also one of the most consequential things that happened to Indian development law this century.

The framework does not deny that justice matters. It has nowhere to put it, which in practice amounts to the same thing.

5. Multiplying through uncertainty eventually multiplies through a catastrophe

Expected-value reasoning is the right tool when you hold many independent bets and can absorb the variance. A hedge fund can. An insurer can. A researcher running a hundred experiments can.

A person cannot. A village cannot. If your intervention has a seventy per cent chance of a large benefit and a thirty per cent chance of serious harm, the expected value may be positive and the thirty per cent is still someone's actual year.

The same structure scales up. A framework that says "take the action with the highest expected value, and be suspicious of intuitions that resist the calculation" will, given enough time and enough confidence, license something disastrous — because the calculation will eventually favour it and the intuition that would have stopped it has already been classified as a bias. Effective altruism's most public institutional failure was not a coincidence of personnel. It was the epistemics working as designed, at scale, with nothing in the structure to say stop, this bet is not yours to place.

Hold on to that phrase. It is the hinge of everything that follows.

What CARE is actually proposing

Now the other framework, and the first thing to notice is that it is not the opposite.

CARE describes a five-step process for deciding what to do and what to keep doing. Benchmark the programme against global research. Test how it performs in this specific context. Measure community feedback on whether it is working. Track cost per outcome continuously. Scale only where there is a credible pathway to adoption — by a government, a market, or a peer organisation. They are building what they describe as a live cost-per-outcome dashboard, tracking participants across more than a hundred countries, and they are explicit that the aim is to see underperformance quickly rather than in a retrospective three years later.

Read step four again. Cost per outcome. That is the same operation as dollars-per-QALY. Good divided by money. If your objection to effective altruism was that it reduces human welfare to a ratio, this framework does that too, and does it live, on a dashboard, at individual-participant resolution.

So anyone who finds one of these frameworks valuable and the other troubling owes themselves an account of the difference. "One counts and the other doesn't" is not it. Here is what I think the difference actually is. There are four parts, and they are structural rather than rhetorical.

Multiplying through uncertainty versus passing through gates Two decision structures. Above, four uncertain estimates are multiplied into a single expected value that authorises action. Below, five sequential gates each of which must be cleared before the next is attempted, so failure at any gate stops the process rather than being averaged away. Multiply through an uncertain estimate is a number you can carry forward effect sizereach p(success)cost expected value one number authorises the action Pass through gates an unanswered question stops the process instead of being averaged away benchmarkvs research test inthis context communitysays it works cost peroutcome adoptionpathway stopstop stopstop
The same ingredients, arranged two ways. Multiplying converts every uncertainty into a number that keeps moving forward. Gating lets any single unanswered question halt the programme.

Difference one: gates instead of multiplication

Effective altruism handles what it does not know by estimating it and carrying the estimate forward. Uncertainty becomes a probability, the probability becomes a coefficient, and the coefficient gets multiplied into a result that then authorises action. Nothing in the structure can stop the process, because every gap has been converted into a number.

CARE's five steps are gates. Each has to clear before the next is attempted. If the community feedback says it is not working, you do not offset that against a strong benchmark and proceed. You have failed a gate.

This is not a softer version of the same epistemics. It is a different one, appropriate to a different risk profile. Expected value is correct when variance is survivable and you hold enough independent bets for the average to be real. Gates are correct when you are operating on people who cannot absorb your variance and who did not choose to be in your portfolio.

That is the whole distinction, and it is not a matter of taste. A donor allocating across two dozen grants genuinely is holding a portfolio. An implementing organisation in one district is not. It is making a single bet with somebody else's year.

Difference two: the feedback is inside the metric

Step three — community feedback on whether it is working — sits before the cost-per-outcome calculation, not after it.

That placement is the substantive claim. In most monitoring systems, beneficiary feedback is collected downstream of impact measurement and reported separately, under headings like satisfaction or accountability. It functions as reputational data. It does not change what the impact number says.

Putting it upstream means the answer to "did this work" is partly constituted by the people it was done to. That does not resolve the ruler problem — CARE still designs the instrument, still chooses the questions, still decides what counts as adequately positive. But it moves the affected person from being an object the metric describes to being one of the parties whose judgement the metric incorporates. On the specific failure this essay identified as the third problem with effective altruism, that is a real structural answer rather than a rhetorical one.

Difference three: a counterfactual you can act on

Effective altruism's counterfactual is global. This rupee against every other possible rupee, anywhere, for any purpose.

For an individual donor with an unrestricted wallet, that comparison is real. For almost every organisation, it is a fantasy. CARE cannot redeploy a Bangladesh field team to distribute bed nets in Chad. The staff have contracts, relationships, languages, and fifteen years of knowing which union council chairman actually returns calls. Treating that as fungible capital misdescribes what it is.

CARE's fifth gate substitutes a counterfactual that an organisation can actually act on: does this outlive us? Scale only where a government, a market, or a peer will carry it. The question is not whether this is the best use of a rupee in the abstract. It is whether the thing will still be running when the grant closes.

This is a better question for an implementer, and a much harder one. It also, as we will see, comes with a cost that the framework does not advertise.

Difference four: the unit is a trajectory, not an intervention

CARE describes tracking "what changed in people's lives, how it happened, and whether it lasted." The unit of analysis is a person over time.

Effective altruism's unit is the intervention: an effect size attributable to a discrete thing that was done, ideally isolable by design. That is what makes it comparable across contexts, and the comparability is the whole point.

But most development outcomes that matter are not produced by one intervention. They are produced by a sequence — a girl stays in school because of a scholarship, a road, a toilet, a teacher who noticed, and a mother who held a position in an argument at home. Attribution to any one of these is not just difficult; it may be the wrong question. Tracking the trajectory instead of the intervention gives up clean attribution and gains a description of what actually happened.

That trade is defensible. It should be made with open eyes: what you lose is precisely the property that let you compare your programme to somebody else's.

Four ways CARE's framework could go wrong

An essay that stopped here would have done the reader a disservice. The five-gate framework has its own failure modes, and three of them are serious.

The dashboard becomes the target

Goodhart's law is not a clever aphorism; it is an empirical regularity. A measure that drives resource allocation stops measuring and starts steering.

The moment cost per outcome is live and visible, every programme manager in the system has an interest in the numerator being large and the denominator small. The fastest route to that is not better programming. It is outcome selection — quietly drifting toward outcomes that are cheap to produce and easy to record. Nobody has to falsify anything. The system will simply become better at producing what the dashboard rewards, and the annual report will describe this as improved efficiency.

The defence is to change what is counted often enough that nobody can optimise against it, and to keep at least one qualitative channel that is not scored. Both are expensive, and both are the first things cut.

A live participant-level dashboard is also a surveillance system

A unified global data system tracking individuals across a hundred countries is, viewed from a slightly different angle, a database of very poor people's circumstances, held by a foreign organisation, in jurisdictions with variable data-protection law, for purposes those people did not specify.

This is not a hypothetical harm. Beneficiary databases have been subpoenaed. They have been demanded by governments as a condition of registration. They have been breached. The person whose nutrition status, household income and location are on that dashboard is the person with the least capacity to object and the most to lose.

None of this makes the system wrong. It does mean that consent, retention limits, and a plan for what happens when a hostile authority asks for the data are not compliance paperwork. They are part of whether the framework is ethical at all. A framework that measures impact per dollar with great care and treats data governance as an annexe has not finished thinking.

"Credible adoption pathway" is conservatism wearing a metrics face

This is the one I would push hardest on, because it is the gate that sounds most obviously sensible.

If the test for scaling is whether a government, a market or a peer will adopt it, then you will systematically do the things that powerful institutions already want done, and systematically not do the things they resist. That is a coherent theory of change. It is also a description of how a sector stops being useful for anything except delivery.

Return to the cases. Manual scavenging persisted for decades after being outlawed precisely because no state government wanted to enforce the ban — there was no adoption pathway, which is why litigation was necessary. The mid-day meal became a universal entitlement because the Supreme Court ordered it in 2001, over the objections of states that did not want the fiscal burden. In Samatha, the Court held that leasing Scheduled Area land to non-tribal mining companies was void; no ministry was waiting to adopt that.

Every one of those is a case where the absence of an adoption pathway was the reason to act, not a signal to stand down. A framework that gates on institutional buy-in will get the delivery work right and lose the accountability work entirely — and it will lose it quietly, because programmes that fail gate five never generate a number that anyone has to explain.

Some goods do not survive division

Cost per outcome requires outcomes to be countable and separable. Some are not.

When a woman speaks in a village meeting for the first time, the outcome is not the speech. It is a change in what she and everyone present now understand to be possible. You can proxy it — attendance, self-efficacy scales, whether she spoke again. Each proxy is a real measurement of something adjacent. None is the thing.

Dignity works like this. So does standing, and so does the difference between being consulted and being informed. These are not immeasurable in a mystical sense. They are constitutive rather than instrumental: they are part of what a good life is, not a means to it, and dividing them by money produces a number whose units are meaningless.

Where arithmetic belongs in a decision A funnel with three stages. Arithmetic is strong at eliminating clearly worse options, weaker at ranking the remaining candidates, and does not settle the final choice, which turns on judgements about justice, standing and reversibility. EliminateNarrowDecide this option is clearlyworse than that one these three areplausibly comparable this is the onewe will do arithmetic is strong arithmetic is weak arithmetic is silent what settles it instead: — is the wrong here one we should refuse? — who gets standing to decide? — if we are wrong, can it be undone? Most of the value of cost-effectiveness analysis is spent in the first panel, and most of the trouble comes from the third.
Cost-effectiveness earns its keep by ruling options out. The damage happens when the same tool is asked to rule one in.

The question neither framework asks

Both frameworks are built to answer which option is best? Neither is built to answer the question that protects you from your own model:

What would have to be true for this to be wrong?

Expected-value reasoning cannot ask it, because it has already priced its uncertainty and moved on. A gated process asks a version of it at each gate, which is better, but the gates are fixed in advance and can only catch the failures somebody anticipated.

The question is different from sensitivity analysis. Sensitivity analysis varies the parameters inside your model. This asks what would have to be true outside it — what the model does not represent at all. Ask it about a nutrition programme and you get answers like: if the reason children are undernourished is that mothers have no bargaining power over household food, our supplement will show effects that fade the moment we leave. Nothing in a cost-per-outcome number contains that.

In practice this is three habits, and they cost almost nothing:

  • Name the mechanism, then attack it. Write down why you think this works, in one sentence, and then argue against that sentence as hard as you can.
  • Ask what your metric would look like if the programme were failing in the way you would find most embarrassing. If the answer is "about the same", your metric is decorative.
  • Ask who would tell you. If the only people positioned to notice a failure are the people whose funding depends on it not being noticed, you do not have a monitoring system. You have a reporting system.

Tuesday morning

Back to the spreadsheet in Dhaka. What actually changes?

Use arithmetic to eliminate, not to select. Cost-effectiveness is a powerful argument against your worst option and a weak argument for your best. If one programme costs four times as much per person reached and nobody can say why, that is a finding and you should act on it. If two programmes are within a factor of two of each other, the model is not telling you anything — the difference is inside the error bars of your own assumptions, and choosing on it is false precision wearing a lab coat.

Ask what the ratio excludes before you ask what it says. Every cost-per-outcome figure has a denominator that somebody defined. Find out what did not make it into the numerator. The legal-aid programme's four hundred people are not the outcome; the changed departmental practice is, and it is not in the column.

Separate the portfolio question from the person question. If you are allocating across many bets and can absorb variance, expected value is a legitimate tool. If you are making one bet on people who cannot absorb it, gates are the right structure and the phrase to keep in mind is: this bet is not yours to place.

Put the feedback upstream. If community judgement arrives after the impact number is computed, it is public relations. If it arrives before, it is measurement. This is a change in sequencing, not in budget, and it is the single highest-value thing on this list.

Keep a protected line for work with no adoption pathway. Not large. But if every gate in your system rewards work that powerful institutions already want, you will drift, and you will not notice the drift because everything you do will keep passing.

What to keep from each

Singer is right that distance is not a moral property, that effectiveness has victims when it is ignored, and that refusing to count is usually avoidance dressed as depth. Those are permanent contributions and the sector is better for having been forced to hear them.

He is wrong that the calculation, once done, settles the matter. It cannot, because the calculation has already made choices — about what counts, about who decides what counts, about whether a wrong is different from a cost — that are the substance of the ethical question rather than preliminaries to it.

CARE's framework is a serious improvement on three of those points: it gates rather than multiplies, it puts the affected person's judgement inside the measure, and it replaces an unusable global counterfactual with a real institutional one. It also introduces a conservatism at its final gate that would have excluded much of what development law in India actually achieved, and it builds an infrastructure whose ethics depend heavily on governance it has not yet described.

The position worth holding is not a synthesis of the two and not a choice between them. It is narrower than that: keep the arithmetic and refuse its authority. Count everything you can. Publish the assumptions. Let the numbers eliminate the options that deserve to be eliminated. And when the number and your sense that something is wrong disagree, do not automatically conclude that the number wins — because the number was built by someone who made choices, and one of those choices may have been to leave out the thing you are noticing.

The programme officer in Dhaka already knows this. The two columns are real and she should use them. So is the third thing she knows, the one that does not fit. Her job is not to make it fit. It is to carry all three into the room and be able to say, out loud, which of them is deciding and why.