Haute Lumière
Commerce · II.10 · MMXXVI · daylight
Volume II — Foundations: The Paradigm and the Science
There is a particular kind of quiet that falls over a meeting when somebody produces a number. The argument stops. Whatever was being contested a moment ago becomes, suddenly, a matter of fact, and the room reorganises itself around the figure on the page. You have felt this. You may have used it.
This chapter is about what happens in the half-second before that quiet — the moment nobody in the room can see, in which somebody long ago decided what the number would count and what it would leave out. That decision is invisible by design. It happened in a committee, it was written into a manual, and by the time the figure reaches your meeting it has been stripped of every trace of having been chosen.
The claim of this chapter is stronger than measures are imperfect, which is true and useless. The claim is that a measure is an intervention. It does not sit outside the system reporting on it. It enters the system, changes what people do, and then reports on the changed system as though the change had nothing to do with it. Two men worked this out independently in the mid-1970s, in different disciplines, and between them they wrote the most important pair of sentences in applied social science.
You will also find here the man who built the national accounts saying, in the document that introduced them to the United States Senate in 1934, that they must not be read as a measure of national welfare — and then watching the world read them that way for ninety years.
None of this is a case against measurement. It is the opposite. This chapter ends with an instrument, a coupon, and a number that decides it, because the answer to a measure that changes behaviour is not to stop measuring. It is to measure in a way that can tell you how much of the change was real.
— The Editors
Begin where measurement is at its best, because it is better than its reputation and the best cases are not obscure.
The instrument shipped with its own warning. Simon Kuznets built the first official estimates of United States national income and delivered them to the Senate in 1934. In that same document, unprompted, he wrote the sentence that every subsequent critic has been re-deriving ever since: "The welfare of a nation can, therefore, scarcely be inferred from a measurement of national income as defined above." He went further in the same pages, noting that economic welfare cannot be judged without knowing the personal distribution of income, and that no income measurement undertakes to estimate the reverse side of income — "the intensity and unpleasantness of effort going into the earning of income." Twenty-eight years later he was still at it, writing that goals for more growth should specify "more growth of what and for what."
Read that as an achievement rather than a lament. The founding document of the national accounts states its own boundary in plain language on the page where the number appears. That is the standard. Almost nothing published since has met it.
The United Kingdom put a price on the work nobody paid for. The Office for National Statistics compiles a household satellite account, and in 2016 it valued unpaid household service work — cooking, childcare, laundry, transport, adult care — at £1,239.0 billion. Set that beside a UK gross domestic product of £1,943.0 billion in the same year and the ratio is 63.8 percent. Not a rounding error, not a footnote: an economy roughly two-thirds the size of the measured one, running in the same houses, done mostly by women, entirely absent from the headline. The ONS did not fold it into GDP. It published it beside GDP, in physical hours and in money, with its method open. That is the right move and it is the model for everything in the Design movement below.
Ninety statistical systems now keep environmental accounts. The System of Environmental-Economic Accounting became an international statistical standard in 2012, its ecosystem-accounting volume followed in 2021, and the United Nations' global assessment counts 89 national statistical systems compiling accounts under it. This matters more than any index. SEEA does not replace GDP with a better single number; it builds physical and monetary accounts for water, energy, timber, fish, land and ecosystem services that reconcile line by line with the national accounts — so a finance ministry can read both without translating. Botswana's water accounts, built under the World Bank's WAVES partnership, showed where water actually went and changed how it was allocated. Costa Rica's water and forest accounts did the same work in a different climate. In each case the accounting came first and the policy followed, which is the order that holds.
New Zealand put a measurement framework inside a budget process. The Treasury's Living Standards Framework organises four capitals — natural, human, social and financial-physical — and from 2019 the budget required spending bids to be argued against them. Whatever one concludes about the allocations, the mechanism is the thing to notice: the measure was not published as advocacy. It was wired into the document that decides money, which is the only place a measure acquires force.
Bhutan built the screen before the index. Gross National Happiness is usually reported as a survey, and it is one — nine domains, thirty-three indicators, a national sample. The more interesting half is the GNH policy screening tool, which scores proposed policies against a fixed list of variables before they proceed. A measure that can stop a proposal is doing something no dashboard does.
Five cases, one pattern. In every one, the measurement published its boundary, sat beside the existing accounts rather than replacing them, and was attached to a decision. Where any of those three is missing, what you have is a publication. Where all three are present, you have an instrument.
First, the two laws, in their real words.
Charles Goodhart, writing in 1975 for a Reserve Bank of Australia conference on the United Kingdom's monetary experience, put it this way:
"Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."
That is narrower and sharper than the version everyone quotes. Goodhart is not saying targets are bad. He is saying that the evidence that justified the target was gathered under conditions the target destroys. The familiar phrasing — when a measure becomes a target, it ceases to be a good measure — is Marilyn Strathern's, from a 1997 paper on audit in British universities, and it is a fair gloss, but it loses the mechanism.
Here is the mechanism with numbers on it. Let a measured proxy M stand beside the thing you actually want, T, correlated at ρ. Before anybody targets anything, a one standard-deviation rise in M predicts ρ² standard deviations of real gain in T. Now impose the target. The proxy can be raised by moving T, which is expensive, or by moving the part of M that has nothing to do with T, which is cheap. Suppose one fifth of the observed response is real.
ρ = 0.9 -> predicted real gain 0.81 s.d.
actual real gain 0.20 s.d.
overstatement 4.05 x
A tight correlation makes the trap worse, not better, because the tighter the historical relationship the more confidently the number is read after the target arrives. At ρ = 0.7 the overstatement is 2.45×; at ρ = 0.9 it is 4.05×.
Donald Campbell, writing in 1976, made the stronger claim, and it is the one to carry:
"The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor."
Goodhart says the statistic degrades. Campbell says the thing being measured degrades. That is a different and much more serious proposition: targeting hospital waiting times does not merely make waiting-time data less informative about care, it changes care. Bevan and Hood documented exactly this in the English health service, where a four-hour emergency target produced patients held in ambulances and reclassified trolleys. Wilson, Croxson and Atkinson found English secondary schools concentrating teaching on pupils sitting on the grade boundary that the published league table happened to cross. Wells Fargo set a cross-sell target and its own consent order with the Consumer Financial Protection Bureau in 2016 described accounts opened without customers asking. None of those organisations was unusually dishonest. All of them were unusually measured.
Second — and this is the cut — GDP has never failed at its job.
The usual reading is that GDP is a broken measure of welfare. It is not broken. It is an excellent measure of something else, and it was built to be. The modern form of the accounts was settled in the early 1940s under wartime conditions, in an argument Kuznets lost: he wanted government war expenditure treated as an intermediate cost — the price of defending a country, not part of its final output — and the war planners needed it counted as final product so that mobilisation would register as growth. The version that won is the version that answers how much output can this country throw at a war, and at that question it is superb. The alternatives are not competing with a failed instrument. They are competing with a precision instrument pointed at a target nobody alive chose, and that is a much harder contest.
Third, what the boundary excludes, with the magnitudes.
The production boundary is not a principle. It is a list. The accounts impute the rental value of owner-occupied housing — roughly 7 percent of United States GDP, a service nobody sold to anybody — and impute nothing at all for childcare done at home. Both are non-market services. One is in because a statistician decided it should be.
Fourth, one accounting change, one real country. Repetto and colleagues at the World Resources Institute recomputed Indonesian output for 1971 to 1984, charging depletion of petroleum, forests and soils. Conventional GDP grew at 7.1 percent a year. The adjusted measure grew at 4.0 percent — a gap of 3.1 percentage points a year. Over those thirteen years the conventional account reports an economy 46.5 percent larger than the adjusted account does. One change to one rule, three stocks, one country, and nearly half the reported expansion is a transfer from the balance sheet to the income statement.
The same arithmetic on the American household account: with household production falling from 39.0 to 25.7 percent of GDP across forty-five years, the identity implies −0.22 percentage points a year off measured growth if the household economy is counted. The BEA's own chain-weighted computation gives about −0.3 points, moving 2.9 percent to 2.6. Two routes, same sign, same order; the difference is that BEA deflates household output separately, and the figure stated here is the identity figure.
Now the honest negative, and it applies to everything this book advocates.
Every alternative index imports a value judgement into a number that then travels as though it were a measurement. Not as a flaw of execution — by construction. Name the judgement in each:
| Index | The judgement it imports | What the judgement is worth |
|---|---|---|
| GPI / ISEW | An inequality-aversion parameter ε, and cumulative charging of natural-capital loss | The same US quintile distribution is worth 0.8260 of mean income at ε = 0.5 and 0.4240 at ε = 2.0 — a 1.95× swing on a parameter no dataset contains |
| HDI | Equal weight on three dimensions, and a log transform of income | Implies a year of life expectancy is worth $68 a head in a poor country and $5,272 in a rich one: 77.5× |
| GNH | A sufficiency cutoff, and equal weight on nine domains | At the published cutoff the happy share is 43.0 percent; at a cutoff of half the indicators it is 90.9 percent — a 47.9-point swing |
| SEEA | Exchange value rather than welfare value | Excludes consumer surplus by design, deliberately, so the accounts reconcile |
| GDP | The market boundary, and government final expenditure as final output | The 1940s decision above |
Work the HDI line, because it is the one people find hardest to believe. The index is a geometric mean of a health index scaled between 20 and 85 years, an education index, and a log income index scaled between $100 and $75,000. Hold the index constant and ask what rise in income exactly offsets one lost year of life. In a country at 55 years and $1,000 per head the answer is $68. At 82 years and $50,000 it is $5,272. Nobody wrote that price down. It fell out of the weights, and Ravallion's 2012 paper on the index's troubling trade-offs is the formal version of this arithmetic.
And there is a second negative, harder than the first. Campbell's law does not exempt the replacement. A wellbeing target is gameable exactly as a growth target is. Worse, GDP has four things no alternative index has: an international standard revised by a statistical commission, a published compilation manual, a revision policy, and independent national compilers. SEEA has acquired all four. The GPI has none of them, which is why two research teams can produce different GPI series for the same country and neither can be adjudicated. Neumayer showed that the ISEW's celebrated turning point is produced mainly by one column — the cumulative charge for natural-capital loss — and that the famous finding is therefore a finding about an accounting convention. The condition a living systems measure needs, and usually lacks, is not better theory. It is plumbing.
In the economy that has taken this seriously, nothing dramatic has happened to the statistics. There is no single new number on the front page, because the attempt to replace one summary figure with another summary figure was abandoned early as a repetition of the original mistake.
What happened instead is that the headline acquired a companion. Every quarterly release of gross domestic product is published beside a depletion-adjusted net product for the same period, from the same agency, on the same page, computed under the environmental accounting standard. Neither is presented as the true number. The gap between them is itself reported, because the gap is the interesting quantity: it is the rate at which the country is converting stocks into income and calling it growth.
The household account is annual and unremarkable. It is quoted in hours before it is quoted in money, so the conversation about whose hours they are happens before the conversation about what an hour is worth. Nobody proposes folding it into GDP; the boundary is more useful stated than dissolved.
Every published index carries its parameters on its face. The inequality aversion used in a welfare-adjusted series is printed in the header with the series, the way a confidence interval is, and the series is published at three values of it so that a reader can see how much of the conclusion is the data and how much is the ethics. This costs nothing and it ends a certain kind of argument permanently, because the argument moves to where it belongs — to the parameter, in the open, where people who disagree can say what they disagree about.
Every target has an expiry date and a holdout. When an organisation adopts a performance measure it names, at the same moment, a second indicator that measures the same underlying thing and that nobody is allowed to target. The divergence between them is published quarterly. It is understood, in the way that inflation-adjustment is understood, that a reported gain in a targeted measure means very little until the holdout has been consulted.
And the conversation in the meeting room has changed shape. When somebody produces a number, the first question is no longer whether it is right. It is what does this count, what does it leave out, and who chose. Those three questions take about twenty seconds, they are not hostile, and they are asked of one's own figures as readily as anyone else's. The half-second of quiet is still there. It is simply followed by a better sentence.
Four mechanisms, in the order they should be built. None requires legislation.
One: the boundary statement. Every number your organisation publishes internally carries three lines — what it counts, what it excludes, and the decision rule that draws the line. Not in an appendix; in the header, above the figure. This is the Kuznets standard and it is the cheapest thing in this chapter. It also has an immediate effect nobody expects: a measure whose exclusions are written down stops being able to expand into territory it was never built for, because the expansion now requires somebody to edit the header, and somebody will object.
Two: the satellite, not the substitute. Where something important is outside the boundary, build a satellite account for it rather than arguing to fold it in. The satellite keeps the boundary visible, keeps the main series comparable with everybody else's, and — this is the operative advantage — can be built by one team in one quarter without anyone's permission, because it does not alter a number anyone depends on. Every successful case in the Discovery movement above took this route. Every unsuccessful campaign of the last fifty years took the other one.
Three: the holdout indicator. This is the anti-Goodhart mechanism, and it is the one genuinely new piece of engineering here. For each targeted measure, name a second indicator that moves with the same underlying reality, that is measured with equal care, and that is never targeted, never in anyone's objectives, and never in any incentive plan. Publish both. The divergence between them is your estimate of the gaming component, and it is the only such estimate available anywhere.
targeted KPI improvement 18.0 %
holdout indicator improvement 4.2 %
divergence 13.8 percentage points
real share of the reported gain 23.3 %
That last line is the number that matters, and note what it does not say. It does not say the programme failed. It says roughly a quarter of the reported gain survived contact with an untargeted measure, which is a finding you can act on, unlike a suspicion. The holdout must be genuinely protected: the moment it enters a bonus formula it stops being a holdout and becomes a second target, and you are back where you started with two corrupted series instead of one.
Four: the parameter register. Any index your organisation uses that contains a weight, a threshold, a discount rate or an aversion parameter gets a one-page register entry naming the parameter, its current value, who set it, when, and what the result looks like at two other plausible values. A parameter that has never been varied is a parameter nobody has examined. The arithmetic above — 1.95× on ε, 77.5× on the HDI's weights, 47.9 points on a sufficiency cutoff — is what a register is for. It takes an afternoon per index and it converts an argument about conclusions into an argument about assumptions, which is the only kind of argument that ever finishes.
Sequence. Boundary statements first, because they cost nothing and they train the reflex. Satellite account second, on the single largest excluded item you can identify. Holdout third, attached to whichever target is currently most consequential. Register fourth, when somebody has begun quoting an index at you.
A measurement regime sustains itself on three conditions, and fails on the absence of any one.
It is compiled by someone who does not own the result. The team that computes the number cannot be the team judged on it. This is why national statistical offices are constituted at arm's length, and it is the single structural feature most often dropped when the practice is imported into a company.
It has a revision policy written before the first revision. How and when the series changes, who signs, and how far back it is restated. Without this, the first inconvenient revision becomes a negotiation, and a series that can be negotiated is no longer evidence.
The holdout is protected by someone senior enough to refuse. The pressure to put the holdout into the incentive plan is constant, reasonable-sounding and eventually irresistible unless one named person's job is to say no.
Now the failure modes, honestly. This fails when the boundary statement becomes boilerplate — copied forward, unread, and therefore no longer a constraint on anything. It fails when the satellite account is built once for a launch and never updated, at which point it is an artefact rather than an instrument. It fails when a divergence between target and holdout is discovered and the response is to question the holdout, which is the most natural response available and the wrong one every time. And it fails, most commonly, when somebody senior decides that a composite index would communicate better than four separate series — which it does, and that is precisely the problem. Composites communicate by hiding the weights, and a weight that has been hidden will not be argued with.
There is a real pleasure in reading a number whose boundary is printed above it. It is the pleasure of being trusted. Somebody who could have handed you a clean figure handed you the figure and the seam it was cut along, and the effect is not doubt — it is the opposite. You believe it more, and you stop holding the small private reservation you had been carrying about every other number in the pack.
There is a second pleasure, quieter, in the moment a holdout indicator confirms a target. It happens more often than the sceptical framing of this chapter would suggest. The two series move together, the divergence is nearly nothing, and what you have is no longer a claim but a measurement that survived an attempt to break it. People who have never done this underestimate how good it feels.
And there is the pleasure of the question itself. What does this count, what does it leave out, and who chose. Once it becomes a reflex it makes reading the world more interesting rather than less — every league table, every ranking, every index in a newspaper becomes a small puzzle with a solvable structure. It is not cynicism. Cynicism is the belief that the numbers mean nothing. This is the much more enjoyable belief that they mean something specific, and that you can find out what.
The instrument: a sustainability-linked facility with a holdout clause.
Sustainability-linked loans and bonds are established structures. The issuer names a key performance indicator, a sustainability performance target, an observation date and a coupon step-up if the target is missed, with a second-party opinion on the calibration. The framework is the ICMA Sustainability-Linked Bond Principles, and your treasury desk will already know it.
The modification is one clause, and it is the whole product. The step-up is triggered by the divergence between the targeted KPI and a named holdout indicator, not by the KPI alone. The holdout is agreed at signing, measured by the same verifier, and contractually excluded from every management incentive plan for the life of the facility.
The mechanics.
The balance-sheet treatment. The facility is ordinary debt and is accounted for as such; the step-up is a contingent cash flow disclosed in the notes and, under the relevant standards, assessed for whether it is closely related to the host instrument. Your auditors will want that conversation early, and it is a conversation they have every year about every variable-rate instrument.
The counterparty. Internal first — treasury lending to a business unit on exactly these terms. It documents in a fortnight, it produces a track record, and it lets you discover what your own holdout does before a bank is watching. Take the structure external once two internal facilities have completed.
The number that decides it. A step-up deters gaming only if it costs more than gaming saves:
cost of real compliance − cost of gaming
step-up > ---------------------------------------------
notional × duration
Work it on a live shape. A $500,000,000 facility, 5.0 years to observation, real compliance costing $12,000,000 and reclassification costing $1,000,000:
required step-up 44 bps
market-standard step-up 25 bps
value of the standard step-up $6,250,000
cost advantage of gaming $11,000,000
deterrence shortfall $4,750,000
the standard step-up is short 1.76 x
The market convention under-deters by nearly a factor of two on this shape, and it is not close. That single inequality is the most useful line in this chapter for anyone with a treasury function, because it converts an argument about integrity into an argument about basis points, and basis points get approved.
The first ninety days.
| Day | Action | Artifact |
|---|---|---|
| 1–15 | Boundary statements on the five numbers in the standing pack | Five headers |
| 16–30 | Choose the KPI and name the holdout | One-page measurement memo |
| 31–45 | Agree and sign the baseline and the verification method | The signed baseline |
| 46–60 | Compute the required step-up; draft the term sheet | Facility memo |
| 61–75 | Internal facility executed, treasury to business unit | Signed facility |
| 76–90 | First divergence reading published | KPI, holdout, and the gap |
Discovery — what is already working
Dream — what becomes possible
Design — what we build
Destiny — how it holds
Alkire, S. and Foster, J. (2011). "Counting and Multidimensional Poverty Measurement." Journal of Public Economics, 95(7–8).
Atkinson, A. B. (1970). "On the Measurement of Inequality." Journal of Economic Theory, 2(3).
Bevan, G. and Hood, C. (2006). "What's Measured Is What Matters: Targets and Gaming in the English Public Health Care System." Public Administration, 84(3).
Bridgman, B., Dugan, A., Lal, M., Osborne, M. and Villones, S. (2012). "Accounting for Household Production in the National Accounts, 1965–2010." Survey of Current Business, 92(5). Bureau of Economic Analysis.
Campbell, D. T. (1976). Assessing the Impact of Planned Social Change. Occasional Paper Series 8. Public Affairs Center, Dartmouth College. Reprinted in Evaluation and Program Planning, 2(1), 1979.
Coyle, D. (2014). GDP: A Brief but Affectionate History. Princeton University Press.
Daly, H. E. and Cobb, J. B. (1989). For the Common Good: Redirecting the Economy Toward Community, the Environment, and a Sustainable Future. Beacon Press.
Goodhart, C. A. E. (1975). "Problems of Monetary Management: The U.K. Experience." In Papers in Monetary Economics, Volume I. Reserve Bank of Australia. Reprinted in Goodhart, C. A. E. (1984), Monetary Theory and Practice: The UK Experience. Macmillan.
International Capital Market Association (2023). Sustainability-Linked Bond Principles. ICMA, Zurich.
Kubiszewski, I., Costanza, R., Franco, C., Lawn, P., Talberth, J., Jackson, T. and Aylmer, C. (2013). "Beyond GDP: Measuring and Achieving Global Genuine Progress." Ecological Economics, 93.
Kuznets, S. (1934). National Income, 1929–1932. Senate Document 124. Seventy-Third Congress, Second Session. United States Government Printing Office.
Kuznets, S. (1962). "How To Judge Quality." The New Republic, 20 October.
Leipert, C. (1989). "National Income and Economic Growth: The Conceptual Side of Defensive Costs." Journal of Economic Issues, 23(3).
Neumayer, E. (1999). "The ISEW: Not an Index of Sustainable Economic Welfare." Social Indicators Research, 48(1).
Neumayer, E. (2003). Weak Versus Strong Sustainability: Exploring the Limits of Two Opposing Paradigms, 2nd edn. Edward Elgar.
New Zealand Treasury (2018). Our People, Our Country, Our Future: Living Standards Framework — Introducing the Dashboard. The Treasury, Wellington.
Office for National Statistics (2018). Household Satellite Account, UK: 2015 and 2016. ONS, Newport.
Ravallion, M. (2012). "Troubling Tradeoffs in the Human Development Index." Journal of Development Economics, 99(2).
Repetto, R., Magrath, W., Wells, M., Beer, C. and Rossini, F. (1989). Wasting Assets: Natural Resources in the National Income Accounts. World Resources Institute.
Stiglitz, J. E., Sen, A. and Fitoussi, J.-P. (2009). Report by the Commission on the Measurement of Economic Performance and Social Progress. Paris.
Strathern, M. (1997). "'Improving Ratings': Audit in the British University System." European Review, 5(3).
United Nations et al. (2014). System of Environmental-Economic Accounting 2012 — Central Framework. United Nations, New York.
United Nations et al. (2021). System of Environmental-Economic Accounting — Ecosystem Accounting. United Nations, New York.
United Nations Statistics Division (2020). Global Assessment of Environmental-Economic Accounting and Supporting Statistics. United Nations, New York.
United Nations Development Programme (2010 onward). Human Development Report, Technical Note 1: Calculating the Human Development Indices. UNDP, New York.
Ura, K., Alkire, S., Zangmo, T. and Wangdi, K. (2012). An Extensive Analysis of GNH Index. Centre for Bhutan Studies, Thimphu.
Wilson, D., Croxson, B. and Atkinson, A. (2006). "'What Gets Measured Gets Done': Headteachers' Responses to the English Secondary School Performance Management System." Policy Studies, 27(2).
Note on figures. Every figure in this chapter is computed in lib/verify/II_10.py with its inputs printed. The Goodhart overstatement, the Atkinson equally-distributed-equivalent incomes, the HDI's implied price of a life year, the GNH headcounts at three cutoffs, the Indonesian and American growth adjustments and the step-up inequality are all reproducible there.