Eight goals with deadlines· 2 of 4

What a target does to a programme

A wall chart of indicators with handwritten annotations
Annotated in the marginIndicator charts acquire handwriting where the definition and the available data disagree.

A number changes what a programme is

Setting a measurable target does not merely describe an ambition; it restructures the activity around the ambition. Administrators, statisticians, field officers and politicians all behave differently once a number is attached to a goal, and the restructuring is not always in the direction the target-setter intended. This is not a peripheral concern in development measurement — it is the core tension that ran through every phase of the MDG era, through the Paris effectiveness agenda, and through the debt-relief architecture that accompanied both.

The mechanism is simple. An unmeasured commitment can be honoured in many ways and judged at leisure. A measured commitment produces a number at a fixed date, and that number carries consequences: for the credibility of the institution that set the goal, for the budget allocations of governments trying to meet it, and for the statistical offices that must produce the evidence. The number disciplines the system. That discipline is the argument for targets. The same discipline is the argument against them, because what gets disciplined is not always what the target-setter had in mind.

The MDG architecture as a case in measurement

The eight Millennium Development Goals adopted in 2000 are the clearest large-scale instance of what a target does to an apparatus. Each goal decomposed into targets, each target into indicators, and each indicator into a data series produced by a statistical office or modelled by an international agency. By fixing a 2015 deadline, the UN system created a global reporting obligation that had not previously existed in that form. National statistical offices that had published intermittently were now pulled into an annual cycle, and UNICEF, the World Bank and the IMF each produced global monitoring compilations tracking progress against the agreed metrics.

The selection of what got measured was itself a policy act. Targets written in terms of halving rates — halving the proportion of people living in extreme poverty, halving the proportion without access to safe water — invited measurement of proportions rather than absolute numbers, which matters because a growing population can produce a falling proportion alongside a rising count of the deprived. The target, as written, could in principle be met by economic growth concentrated in large middle-income countries, while smaller low-income countries recorded no progress at all. Whether that constitutes failure depends entirely on what the target was understood to mean, which the text left ambiguous.

Statisticians call this construct validity: the question of whether the thing being measured corresponds to the thing the policy intends to affect. Where construct validity is poor, a programme can hit the target and miss the point, or miss the target and hit the point. The MDG poverty goal measured consumption against a dollar-a-day threshold, and the progress recorded on that metric by 2015 was substantial — the World Bank's PovcalNet database tracked the global headcount ratio falling sharply. But the threshold was contested from the start, and rebasing national accounts — the periodic revision of base years used in GDP and price calculations — moved countries across it as a statistical artefact, not as a consequence of any change in living conditions.

A printed goal framework poster on an office wall
Eight goals, printedA framework on a wall is also a reporting obligation: every line needs a number filed against it.

What the Paris indicators measured, and why that matters

The 2005 Paris Declaration introduced a different kind of target: not a development outcome target, but a process target. Its twelve indicators measured donor behaviour — whether aid was aligned to national budget systems, whether donors harmonised their reporting requirements, whether countries produced their own plans rather than parallel donor-managed ones. Each indicator came with a numerical benchmark to be achieved by 2010, later extended to Accra in 2008 and Busan in 2011.

The effect on behaviour was immediate and legible. Donors began keeping records of how much aid passed through partner-country systems — the alignment indicator — because the OECD's monitoring survey would ask. Ministries of finance in recipient governments began coding aid flows to national budget lines because doing so improved their performance on the same survey. This is the constructive version of what a target does: it makes something that was informal into something countable, and counting it creates pressure to improve it.

The problematic version appeared alongside it. Some aid was routed through national budget systems precisely to improve the alignment score, even where the national system was not yet capable of managing it without additional transaction costs. The number went up; the underlying administrative capacity did not necessarily follow. The Busan outcome document in 2011 shifted the language away from the Paris numerical commitments toward broader principles — a change widely interpreted as acknowledging that the indicator-driven approach had reached its limits, and that hitting the numbers had partly substituted for achieving the underlying goals.

A national statistics office of desks, filing cabinets and printouts
Where the estimate is assembledMost of the correction happens between the completed form and the published table, in a room like this one.

The goodhart problem, and why it is not a reason to abandon measurement

Goodhart's Law — formulated by the British economist Charles Goodhart in the context of monetary policy and applied extensively since — holds that when a measure becomes a target, it ceases to be a good measure. The MDG and Paris experiences both illustrate this, and both tempt an overcorrection: that measurement itself is suspect, or that targets cause more distortion than they prevent.

The evidence does not support that conclusion. What unmeasured programmes produce is not less distortion but invisible distortion. Debt relief under HIPC — the Heavily Indebted Poor Countries initiative — required countries to reach a completion point, which in turn required demonstrated commitment to poverty reduction strategies. The conditionality was real, the completion point was a genuine gate, and tracking what happened to fiscal space after relief required statistical offices to produce figures that had not previously existed in those countries at that level of granularity. Whatever the limitations of those figures, they produced more accountability than the pre-HIPC system of untracked arrears.

The useful question is not whether to measure but what the indicator's construct validity is, how far down the causal chain it sits from the outcome of interest, and what secondary behaviours the target will incentivise. An indicator close to the outcome — a household survey measuring consumption directly — has different properties from an indicator that is one step removed, like public spending on health. An indicator two steps removed — a donor's disbursement record — tells you less still about whether any child is better fed.

A spreadsheet printout with columns of country rows
The monitoring returnCountry rows filled in by the governments and agencies being measured.

All of this implies that a single target is almost always insufficient. It measures one slice of a programme's effect, and the programme will be managed along that slice while the unmeasured slices drift. The MDG framework addressed this partly through multiplicity — eighteen targets and forty-eight indicators across the eight goals — but multiplicity creates its own problems: indicator fatigue, inconsistent data quality across series, and the political temptation to report progress on the indicators where progress is recorded and to stay quiet on those where it is not. Choosing which numbers to lead with is itself a policy act, and the statistical offices that run the household surveys and produce the underlying figures that feed those numbers sit at the intersection of technical work and political consequence in a way that is rarely acknowledged in the published headline.

Read next