Key Terms in Monitoring, Evaluation and Learning
A disambiguation handout — for MEL practitioners who work across donors, frameworks and disciplines, and who keep running into the same words meaning different things.
ImpactMojo · MEL & Research Track · August 2026 · companion to the MEL Rosetta Lab
How to use this
The handout is organised as pairs and clusters of terms that get confused, not as an alphabetical glossary. Each entry states the confusion, gives the competing definitions with their sources, offers a practical test for telling them apart in a live document, and where relevant names what breaks if you get it wrong. Every definition is quoted or closely reported from a named primary source with a date. Where a term has no authoritative definition, the handout says so rather than inventing one.
The reference source throughout is the OECD Development Assistance Committee glossary, second edition, published 2023. It matters that this is the second edition: the first appeared in 2002, and the 2023 revision incorporates the evaluation criteria definitions approved in 2019 along with terms from the DAC Results Community that did not exist in 2002. Anyone still working from a 2002 copy is working from superseded definitions. Where the DAC glossary has no entry for a term in daily use, that absence is itself reported — four such terms appear at the end.
Part one: the results chain, and why it is the root of most confusion
Most terminological trouble in MEL traces back to one place. The results chain is a shared idea across every framework, but the labels attached to its rows are not shared. The DAC glossary defines the results chain as the causal sequence of an intervention that stipulates the stages leading to the achievement of the desired objectives, running from inputs through activities and outputs to outcomes and impacts, and it notes that in some cases reach is included as part of the chain.
The definitions below are the DAC ones. They are the closest thing the field has to a common reference, and they are not what every donor means.
| Level | OECD-DAC definition, second edition 2023 |
| Inputs | The financial, human, material and institutional resources used for the intervention. The glossary presses users to count the resources of all involved organisations, the community and the local environment, not only those of the funder or implementer. |
| Activities | Actions taken through which inputs are mobilised to produce outputs. |
| Outputs | The products, capital goods and services that result from an intervention, including changes in knowledge, skills or abilities produced by the activities. The note is the important part: outputs are within the control of the implementing team and attributable to it. |
| Outcomes | The short-term and medium-term effects of an intervention's outputs. Often changes in institutional and behavioural capacities occurring between the completion of outputs and the achievement of impacts. |
| Impacts | The higher-level effects of an intervention's outcomes. The glossary warns directly that impacts and results are sometimes used interchangeably, which creates confusion, and that impacts should refer to higher-level results. |
| Results | The outputs, outcomes or impacts of an intervention, intended or unintended, positive or negative. A generic word for any level of the chain. |
The same words on different rows
Four donor systems in common use assign these words differently. This table is the one worth pinning above a desk.
| Level | UN RBM and OECD-DAC | Classic USAID logframe | EU and EuropeAid | DFID and FCDO |
| Highest | Impact | Goal | Overall objective | Impact |
| Medium-term | Outcome | Purpose | Specific objective | Outcome |
| Products | Output | Outputs | Expected results | Outputs |
| Actions | Activities | Inputs and activities | Activities | Inputs and activities |
| Resources | Inputs | Inputs | Means | Inputs |
Two crossings cause most of the damage. Outcome in FCDO and UN usage is purpose in the classic USAID logframe and specific objective in EU usage. Result in EU usage means output, while result in DAC usage means any level at all. An indicator copied from a EuropeAid logframe into a UN results framework under the heading it arrived with will land one row too high.
The Logical Framework itself was developed in 1969 by Leon J. Rosenberg of Fry Consultants, later Practical Concepts Incorporated, and implemented by USAID across roughly 30 country programmes in 1970 and 1971. Every later dialect is a divergence from that matrix.
Part two: terms that get confused
Attribution and contribution
The confusion. Both describe the relationship between what an intervention did and what changed. They make claims of different strength, and funders often ask for the first while the evidence can only support the second.
Attribution. The ascription of a causal link between observed, or expected to be observed, changes and a specific intervention. The DAC glossary adds a qualification that is routinely missed: the definition does not require that changes were produced solely by the intervention being evaluated. It represents the extent to which observed effects can be attributed to a specific intervention or to one or more partners, taking account of other interventions, confounding factors and external shocks. Source: OECD-DAC Glossary, second edition, 2023.
Contribution. The role or part played by an intervention, together with other interventions, in bringing about an observed or expected result. The ways an intervention helps to advance towards a goal. Source: OECD-DAC Glossary, second edition, 2023. The associated method, contribution analysis, is an approach for determining whether and how an intervention contributed to an observed result, based on verifying the underlying theory of change, developed by John Mayne.
The practical test. Ask what would have to be true for the claim to fail. An attribution claim fails if the change would have happened anyway, which is why it needs a counterfactual. A contribution claim fails if the theory of change is not borne out, which needs evidence about the causal steps and about the other factors at work, not a control group.
What breaks. Writing attribution language into a results framework for an outcome that no design can attribute. The programme then either overstates its case or reports failure against a claim it never should have made. Outputs are attributable to the implementing team by definition; outcomes and impacts usually are not.
Impact and reach
The confusion. Reach counts people. Impact describes change in their lives. Annual reports slide from one to the other, because reach numbers are large and available while impact evidence is small and expensive.
Reach. The people affected by an intervention, as a subset of the total population. The DAC glossary notes that in some cases reach is included as part of the results chain. Source: OECD-DAC Glossary, second edition, 2023. In the RE-AIM framework, reach is the first of five dimensions and is specified more tightly as the number, proportion and representativeness of those who participate. Source: Glasgow, Vogt and Boles, American Journal of Public Health 89(9), 1999, pp. 1322–1327.
Impact. As an evaluation criterion: the extent to which the intervention has generated or is expected to generate significant positive or negative, intended or unintended, higher-level effects. As a level of result, impacts are the higher-level effects of an intervention's outcomes. Source: OECD-DAC Glossary, second edition, 2023.
The practical test. Reach is a denominator question and a representativeness question. Ask who was not reached and how they differ from those who were. Impact is a change question. Ask what is different now, for whom, and compared to what.
What breaks. Reporting reach as if it were impact. Twelve thousand people trained is a reach figure and an output figure at once. It carries no information about whether anything changed, and the representativeness half of the reach construct, which is where the equity information sits, usually goes unreported.
Impact the criterion and impacts the level of result
The confusion. The DAC glossary carries two separate entries for the same word, one singular and one plural, and warns explicitly about the confusion.
Impact, singular, as a criterion. One of the six evaluation criteria. It asks about the ultimate significance of the intervention, including social, environmental and economic effects that are longer term or broader in scope than those captured under effectiveness, and looks at changes in systems or norms and effects on well-being, human rights, gender equality and the environment.
Impacts, plural, as a level of result. A row in the results chain: the higher-level effects of an intervention's outcomes. The glossary states that impacts and results are sometimes used interchangeably, which creates confusion, and that impacts should be used for higher-level results. Source for both: OECD-DAC Glossary, second edition, 2023.
The practical test. If the sentence could be replaced with "how much did this ultimately matter", it is the criterion. If it names a row in a logframe or results framework, it is the level.
What breaks. Commissioning an impact evaluation when what is wanted is a judgement against the impact criterion, or the reverse. These need different designs and different budgets.
Effectiveness, efficiency and efficacy
The confusion. The first two are DAC criteria and are regularly swapped. The third is not a DAC term at all and arrives from clinical research, where it means something specific.
Effectiveness. The extent to which the intervention achieved, or is expected to achieve, its objectives and its results, including any differential results across groups. Analysis involves taking account of the relative importance of the objectives or results. Source: OECD-DAC Glossary, second edition, 2023. The clause about differential results across groups was added in the 2019 revision and is the hook for equity analysis.
Efficiency. The extent to which the intervention delivers, or is likely to deliver, results in an economic and timely way. Economic means converting inputs into outputs, outcomes and impacts in the most cost-effective way possible compared to feasible alternatives in the context. Timely means within the intended timeframe, or one reasonably adjusted to an evolving context. Source: OECD-DAC Glossary, second edition, 2023.
Efficacy. No entry in the DAC glossary. The term comes from the distinction between explanatory and pragmatic trials drawn by Schwartz and Lellouch, Journal of Chronic Diseases 20(8), 1967, pp. 637–648. Explanatory designs test whether something works under controlled or ideal conditions, which is efficacy. Pragmatic designs test whether it works in real-world conditions, which is effectiveness. The RE-AIM framework is a live illustration of the drift: the E stood for efficacy in the 1999 paper and is now generally rendered as effectiveness.
The practical test. Effectiveness asks whether objectives were met. Efficiency asks what it cost to meet them, in money and in time. Efficacy asks whether the thing can work at all when conditions are favourable, which is a question about the intervention rather than about this programme.
What breaks. Claiming effectiveness on the strength of an efficacy trial. A pilot run with intensive supervision in three well-resourced sites says little about what happens at scale, and the two words hide that gap.
Monitoring and evaluation
The confusion. Joined by an ampersand so often that the difference in function gets lost. They answer different questions, run on different cycles, and are commissioned differently.
Monitoring. A continuing process involving the systematic collection or collation of data on specified indicators or other information, which provides management and other stakeholders with indications of the extent of implementation progress, achievement of intended results, occurrence of unintended results, use of allocated funds, and other intervention and context information. Source: OECD-DAC Glossary, second edition, 2023.
Evaluation. The systematic and objective assessment of a planned, ongoing or completed intervention, its design, implementation and results, aiming to determine relevance, coherence, effectiveness, efficiency, impact and sustainability. It also refers to the process of determining the worth or significance of an intervention. The glossary adds that not all evaluations will cover all criteria to the same degree or at all. Source: OECD-DAC Glossary, second edition, 2023.
The practical test. Monitoring tells you what is happening. Evaluation tells you what it is worth. The word doing the work in the evaluation definition is judgement: worth or significance. A study that reports numbers without reaching a judgement is monitoring with a longer report.
What breaks. Nothing structural, but budgets suffer. Monitoring data collected without an evaluative question in mind rarely answers one later, and evaluations commissioned without monitoring data behind them spend their budget reconstructing a baseline.
Indicator, performance indicator, target and milestone
The confusion. Four things at different levels of specificity, often used as if interchangeable in the same logframe column.
Indicator. A quantitative or qualitative factor or variable of interest, related to the intervention and its results or to the context in which it takes place. The note deserves attention: an indicator is always approximate only, not an exact measure, and requires interpretation and explanation even when assessed accurately. Source: OECD-DAC Glossary, second edition, 2023.
Performance indicator. A narrower term: a factor or variable providing a simple, verifiable and reliable means to measure the performance of an actor, generally in terms of the process of implementation. Source: OECD-DAC Glossary, second edition, 2023.
Target. An objective, usually quantitative, defined as a value on an established indicator, generally set at the beginning of an intervention and expected to be achieved by a specific point in time with available resources. Source: OECD-DAC Glossary, second edition, 2023.
Milestone. No entry in the DAC glossary. In FCDO logframe practice it is an interim marker on the way to a target. In payment-by-results contracting it is the payment trigger itself. The two usages carry very different consequences for missing one.
The practical test. An indicator is what you measure. A target is a value on that indicator plus a date. A milestone is an interim value, or a contractual trigger, depending on whose document you are in. If a cell contains a number with no indicator definition behind it, the indicator has gone missing and the number is doing work it cannot support.
What breaks. The glossary's insistence that every indicator is approximate and needs interpretation gets dropped when the indicator becomes a target, and the target then gets read as an exact measure of the thing itself.
Assumption and risk
The confusion. The same underlying fact, written in opposite directions in two different documents, which is why a logframe and a risk register can look inconsistent when they are not.
Assumptions. A set of untested factors and beliefs that form the basis of the intervention logic, and factors or risks which affect its relevance, progress or success. Assumptions are the conditions necessary for the cause-and-effect relationships between the different levels of results, moving from activities to outputs, outputs to outcomes, and outcomes to impacts. The glossary adds a second sense used in design and evaluation: hypothesised conditions affecting the validity of the exercise itself, such as assumptions about the size or characteristics of a population when designing a sampling procedure. Source: OECD-DAC Glossary, second edition, 2023.
Risk. Treated in the glossary as part of the same entry rather than separately. In practice, a risk register states negatively what a logframe assumption column states positively. "Rainfall is adequate" and "drought reduces yields" are one fact, written twice.
The practical test. Read the assumption column and the risk register side by side. If a condition appears in one and not the other, one of the two documents is incomplete. Note also the second sense: sampling assumptions belong to the validity of your study, not to the intervention logic, and confusing the two puts a methodological caveat in a programme document where nobody will act on it.
Theory of change, logframe, results framework and results chain
The confusion. Four distinct objects, treated as synonyms in a great many terms of reference.
Theory of change. The way the intervention is expected to achieve or achieves change. It represents how people understand change to occur in a given context, including explicit or implicit assumptions about the causal links between inputs, activities and results, and often includes evidence and risks. Source: OECD-DAC Glossary, second edition, 2023.
Logical framework. A management tool used to improve the design of interventions, most often at project level, identifying strategic elements and their causal relationships along with indicators and the assumptions or risks that may influence success and failure. Source: OECD-DAC Glossary, second edition, 2023.
Results framework. Explicit articulation, typically graphical or tabular, of how a strategy or intervention will achieve its objectives, including causal relationships and underlying assumptions and risks. Generally includes indicators with baseline, data source and means of verification for the full chain. Source: OECD-DAC Glossary, second edition, 2023.
Results chain. The causal sequence itself, from inputs to impacts. Source: OECD-DAC Glossary, second edition, 2023.
The practical test. A theory of change explains why change is expected, and can hold complexity, feedback and competing pathways. A logframe is a management matrix that compresses that explanation into columns and drops most of it. A results framework sits above the project, usually at strategy or portfolio level. The results chain is the underlying sequence all three refer to.
What breaks. Asking for a theory of change and accepting a logframe. What comes back is a linear matrix with the causal reasoning stripped out, and the assumptions that were the whole point of the exercise reduced to one column of short phrases.
Goal, objective and purpose
The confusion. Three words for higher-level intent, each anchored to a different row depending on the system.
Goal. The higher-order objective to which an intervention is intended to contribute. The glossary gives the Sustainable Development Goals as the example. Source: OECD-DAC Glossary, second edition, 2023.
Objective. Intended positive impacts contributing to physical, financial, institutional, social, well-being, environmental or other benefits to a society, community or group of people. Source: OECD-DAC Glossary, second edition, 2023.
Purpose. Not a DAC term at this level. In the classic USAID logframe, purpose is the second row, equivalent to outcome in UN and FCDO usage. The DAC glossary does define intervention purpose in the sense of overall purpose, under the entry for intervention objective.
The practical test. The word alone tells you nothing. Ask which system the document belongs to before reading any of the three. In a USAID-lineage matrix, goal is the top row and purpose is the second. In UN RBM, the equivalents are impact and outcome.
Relevance and coherence
The confusion. Coherence was added to the criteria in 2019, and much of what it covers used to be reported under relevance, so older evaluations and newer ones are not directly comparable on either.
Relevance. The extent to which the intervention objectives and design respond to beneficiaries, global, country and partner or institution needs, policies and priorities, and continue to do so if circumstances change. The glossary explains that respond to means the objectives and design are sensitive to the economic, environmental, equity, social, political economy and capacity conditions in which the intervention takes place, and that assessing relevance involves looking at differences and trade-offs between priorities. Source: OECD-DAC Glossary, second edition, 2023.
Coherence. The compatibility of the intervention with other interventions in a country, sector or institution. It splits in two. Internal coherence covers synergies and interlinkages with other interventions by the same institution or government, and consistency with the international norms and standards that institution adheres to. External coherence covers consistency with other actors' interventions in the same context. Source: OECD-DAC Glossary, second edition, 2023; criterion added in the 2019 revision.
The practical test. Relevance looks outward at needs. Coherence looks sideways at other interventions. A programme can be highly relevant to a real need and incoherent with three other programmes doing the same thing in the same district.
What breaks. Comparing a 2015 evaluation with a 2023 one on relevance. The earlier one is likely carrying coherence questions inside its relevance section, because there was nowhere else to put them.
Counterfactual, control group and comparison group
The confusion. The first is a condition, the second and third are how you estimate it. They are not the same kind of thing, and the DAC glossary defines them separately.
Counterfactual. The situation or condition that hypothetically may prevail for individuals, organisations or groups were there no intervention, that is, the status quo. The glossary notes it can be estimated by creating a control group, a comparison group or a hypothetical counterfactual. Source: OECD-DAC Glossary, second edition, 2023.
Control group. The sample or group that does not receive the intervention and against which other samples or groups that do receive it are compared in order to assess performance and results. Source: OECD-DAC Glossary, second edition, 2023.
Comparison group. Used where allocation was not randomised. The distinction matters because a comparison group carries selection concerns a randomised control group is designed to remove.
The practical test. The counterfactual is what you are trying to estimate. The control or comparison group is one method of estimating it. Saying a study has a counterfactual says nothing about how credible the estimate is.
Baseline, baseline study and validity
The confusion. Two related DAC entries plus the quality concept that determines whether either is worth anything.
Baseline, or reference value. The conditions existing prior to an intervention or at the beginning of the period, against which changes can be measured, monitored and evaluated. Source: OECD-DAC Glossary, second edition, 2023.
Baseline study. An analysis describing the situation prior to an intervention, against which changes can be assessed or comparisons made. Source: OECD-DAC Glossary, second edition, 2023.
Validity. The extent to which an evaluation is logically and factually sound. Source: OECD-DAC Glossary, second edition, 2023.
Triangulation. The use of three or more theories, sources or types of information or analysis to verify and substantiate an assessment, seeking to overcome the bias that comes from single informants, single methods, single observers or single theory studies. Source: OECD-DAC Glossary, second edition, 2023. Note the number: three or more, not two.
The practical test. A baseline is a value. A baseline study is the exercise that produces values across many indicators. A recalled or reconstructed baseline, common where a study starts after the programme did, is not the same object as a measured one and should not be reported as if it were.
Beneficiaries, target group and reach
The confusion. Three overlapping ways of describing people, with an equity caution built into the first.
Beneficiaries, or people who benefit. The individuals, groups or organisations, whether targeted or not, that benefit directly or indirectly from the intervention. The glossary notes that other terms are used depending on context, including rights holders, duty bearers and affected people, and cautions that one should not assume all people have equal access to or benefit equally, or at all, from an intervention. Source: OECD-DAC Glossary, second edition, 2023.
Target group. The specific individuals, communities or organisations that the intervention is intended to reach. Source: OECD-DAC Glossary, second edition, 2023.
Reach. The people actually affected, as a subset of the total population. Source: OECD-DAC Glossary, second edition, 2023.
The practical test. Target group is intent. Reach is what happened. Beneficiaries is who gained, which may include people never targeted and exclude people who were. The gap between the three is usually where the equity story is.
Part three: terms in daily use that the DAC glossary does not define
Searching the full text of the second edition returns no entry for any of the following. This is worth knowing, because each is used as though it had an agreed meaning, and disputes about them cannot be settled by pointing at the glossary.
| Term | Where its meaning actually comes from | The practical consequence |
| Milestone | FCDO logframe practice, where it is an interim marker; payment-by-results contracting, where it is a disbursement trigger. | Missing a milestone means a report footnote in one setting and a lost payment in the other. Confirm which applies before agreeing to any. |
| Coverage | Humanitarian standards work and public health, where it usually means the proportion of a population in need that was served. | Often used loosely as a synonym for reach, losing the denominator of people in need that gives it meaning. |
| Fidelity | Implementation science, meaning the degree to which delivery matched the intended design. | Without it, a null result cannot be read: you cannot tell whether the intervention failed or was never delivered as designed. |
| Efficacy | Clinical trial methodology, from the explanatory and pragmatic distinction in Schwartz and Lellouch, 1967. | Slides into effectiveness in reporting, converting a claim about ideal conditions into a claim about real ones. |
Quick reference — for pinning up
The full definitions are in part two.
| If you are asking | The right term is | Not |
| Did this intervention cause the change, ruling out that it would have happened anyway | Attribution | Contribution |
| Did this intervention play a part, alongside others, in the change | Contribution | Attribution |
| How many people did we touch, and were they representative | Reach | Impact |
| What is different in people's lives now | Impact | Reach or outputs |
| Did we meet our objectives | Effectiveness | Efficiency |
| What did meeting them cost, in money and time | Efficiency | Effectiveness |
| Can this work under favourable conditions | Efficacy | Effectiveness |
| What is happening right now in delivery | Monitoring | Evaluation |
| What is this worth, and should we keep doing it | Evaluation | Monitoring |
| What do we measure | Indicator | Target |
| What value by when | Target | Indicator |
| Does this fit the need | Relevance | Coherence |
| Does this fit alongside everything else being done | Coherence | Relevance |
| What would have happened without us | Counterfactual | Control group |
| How are we estimating that | Control or comparison group | Counterfactual |
| Why do we think change happens this way | Theory of change | Logframe |
Sources
OECD (2023), Glossary of Key Terms in Evaluation and Results-Based Management for Sustainable Development, second edition, OECD Publishing, Paris, DOI 10.1787/632da462-en-fr-es. The source for every definition marked as DAC in this handout. First edition published 2002; the second edition incorporates the evaluation criteria definitions approved in 2019 and terms from the DAC Results Community.
OECD/DAC (1991), Principles for Evaluation of Development Assistance. The original five evaluation criteria.
OECD/DAC Network on Development Evaluation (2019), Better Criteria for Better Evaluation: Revised Evaluation Criteria Definitions and Principles for Use, adopted 10 December 2019. Coherence added as a sixth criterion; the other five redefined.
Glasgow, R.E., Vogt, T.M. and Boles, S.M. (1999), Evaluating the public health impact of health promotion interventions: the RE-AIM framework, American Journal of Public Health 89(9), pp. 1322–1327. Reach as a construct with a representativeness requirement.
Schwartz, D. and Lellouch, J. (1967), Explanatory and pragmatic attitudes in therapeutical trials, Journal of Chronic Diseases 20(8), pp. 637–648. The origin of the efficacy and effectiveness distinction.
Roche, C. (1999), Impact Assessment for Development Agencies: Learning to Value Change, Oxfam and Novib. SPICED indicators, and the definition of impact as lasting or significant change in people's lives.
Doran, G.T. (1981), There's a S.M.A.R.T. Way to Write Management's Goals and Objectives, Management Review 70(11), pp. 35–36.
Kusek, J.Z. and Rist, R.C. (2004), Ten Steps to a Results-Based Monitoring and Evaluation System, World Bank, which attributes CREAM to Schiavo-Campo, S. (1999), Performance in the Public Sector, Asian Journal of Political Science 7(2).
UNDP (2009), Handbook on Planning, Monitoring and Evaluating for Development Results.
UNEG (2016), Norms and Standards for Evaluation, adopted April 2016.
A note on verification. Definitions here are reported from the primary sources named. Where this handout says a term has no DAC entry, that was checked against the full text of the second edition rather than assumed. Anyone quoting a definition verbatim in a publication should open the cited source, because several widely repeated definitions in the secondary literature have drifted from the originals.