Haute Lumière

Commerce · II.10 · MMXXVI · daylight

La Bourse  /  Volume II  /  Nº II.10  /  Ten concept briefs

A man in a dark suit writing in a ledger beside printed charts, a library of books behind him.
Plate II.10 · Ten concept briefsThe Weight and the Weighing.The balance does not describe the cloth. It describes the cloth in the one respect somebody decided to care about, and then the ledger forgets that a decision was made.

TEN CONCEPT BRIEFS · Chapter II.10 — Measurement and What It Does

One page each. A reader who reads only these ten pages has the chapter.


BRIEF 1 — Goodhart's Law, in Its Original Form

The idea. Charles Goodhart, writing in 1975 on the United Kingdom's monetary experience, said this:

"Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."

Note what he did not say. He did not say targets are bad, or that measuring is futile. He said something precise: the relationship you observed was observed under conditions in which nobody was trying to move it. Once you start trying, the relationship you relied on is no longer the relationship you are in.

The popular phrasing — when a measure becomes a target, it ceases to be a good measure — is Marilyn Strathern's, from a 1997 paper on university audit. It is a fair summary and it loses the mechanism, which is the part you can act on.

Worked example. Let a measured proxy M sit beside the thing you want, T, correlated at ρ. Before targeting, a one standard-deviation rise in M predicts ρ² standard deviations of real gain. After targeting, the rise can come from the part of M unrelated to T. If a fifth of the response is real:

  ρ = 0.9   predicted 0.81 s.d.   actual 0.20 s.d.   overstated 4.05 x
  ρ = 0.7   predicted 0.49 s.d.   actual 0.20 s.d.   overstated 2.45 x

The counter-intuitive part: the tighter the historical correlation, the worse the trap, because a tight correlation is exactly what persuades everyone to trust the number after the target arrives.

You already know this because you have watched a metric go green in a way that made nobody's life better, and you knew within a week that the green was real and the improvement was not.


BRIEF 2 — Campbell's Law, and Why It Is the Stronger Claim

The idea. Donald Campbell, writing in 1976:

"The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor."

The difference that matters. Goodhart says the statistic degrades. Campbell says the process degrades. Under Goodhart, the waiting-time figure stops telling you about the quality of care. Under Campbell, the quality of care itself changes shape to fit the figure, and it does so permanently, whether or not anyone ever looks at the figure again.

Worked examples, documented. Bevan and Hood studied the English health service's four-hour emergency target and found patients held in ambulances and trolleys reclassified. Wilson, Croxson and Atkinson found English secondary schools concentrating teaching effort on pupils sitting exactly on the grade boundary the published table happened to cross — pupils well above and well below it received less. Wells Fargo's own 2016 consent order with the Consumer Financial Protection Bureau described accounts opened without customers asking, under a cross-sell target.

What this is not. It is not a claim about dishonest people. Every organisation above contained mostly conscientious ones. Campbell's law describes what happens to conscientious people under a single number.

You already know this because you have taught to a test, or been taught to one, and you can still remember which parts of the subject were quietly dropped.


BRIEF 3 — The Production Boundary

The idea. The national accounts do not measure "the economy." They measure everything inside a line, and the line is a list rather than a principle.

What is inside. All market transactions. Own-account production of goods — subsistence farming counts. And one large imputed service: the rental value of owner-occupied housing, worth roughly 7.0 percent of United States GDP, a service nobody sold to anybody.

What is outside. Own-account production of services. Childcare at home, cooking, laundry, elder care, the unpaid transport of other people's children. Depletion of a forest, a fishery, an aquifer or a soil. Most of the value of volunteering.

The trap this closes. The usual defence of the boundary is that non-market activity cannot be valued. But the accounts already value one enormous non-market service — housing — using a replacement-cost method that would work equally well on childcare. The boundary is not where valuation becomes impossible. It is where a committee stopped.

Worked example. A parent leaves paid work to care for a child at home. Measured GDP falls by their salary and by the nursery fee. Measured output of childcare falls to zero. The actual quantity of childcare provided has, at minimum, not decreased.

You already know this because you have heard someone say they "don't work" while doing forty hours a week of it.


BRIEF 4 — Kuznets's Warning

The idea. The man who built the United States national income estimates shipped the warning in the same document as the instrument.

In National Income, 1929–1932, presented to the Senate in 1934, Simon Kuznets wrote:

"The welfare of a nation can, therefore, scarcely be inferred from a measurement of national income as defined above."

And in the same pages: economic welfare cannot be judged without knowing the personal distribution of income, and no income measurement estimates the reverse side of income — "the intensity and unpleasantness of effort going into the earning of income." In 1962 he added that goals for more growth should specify "more growth of what and for what."

The part usually left out. Kuznets also lost an argument. In the wartime recasting of the accounts he wanted government war expenditure treated as an intermediate cost rather than final output. The war planners needed it as final output, so that mobilisation would show up as growth. The version we use is the version that won that argument, and it answers the question it was built for — how much output can this country mobilise — extremely well.

Why it matters. GDP is not a failed welfare measure. It is a successful measure of something else. That reframing changes what an alternative has to do.

You already know this because you have used a tool designed for one job to do another, and blamed the tool.


BRIEF 5 — Household Production

The idea. The largest excluded item is not exotic. It is housework.

The numbers. The Bureau of Economic Analysis satellite account values non-market household production at 39.0 percent of GDP in 1965 and 25.7 percent in 2010 — in 2010, $3,853.0 billion against a GDP of $14,992.1 billion. The United Kingdom's Office for National Statistics, on a broader definition, valued unpaid household service work at £1,239.0 billion in 2016 against a GDP of £1,943.0 billion — 63.8 percent.

What the fall means, and does not mean. The share fell because household work moved into the market: restaurant meals, laundry services, paid childcare. Measured GDP grew partly by reclassifying work that was already being done.

Worked example — the growth effect. If household production falls from 39.0 to 25.7 percent of GDP over forty-five years, the identity implies about −0.22 percentage points a year off measured growth once it is counted. BEA's own chain-weighted computation gives about −0.3 points, moving 2.9 percent to 2.6. Two routes, same sign, same order; the difference is the deflator.

You already know this because you have noticed that the same dinner counts for nothing when you cook it and counts for something when you buy it.


BRIEF 6 — Depletion, and the Adjusted Product

The idea. The accounts charge depreciation on a lathe and nothing on an oil field. Sell the lathe and it is a capital transaction; sell the oil and it is income.

Worked example, and the chapter's headline arithmetic. Repetto and colleagues at the World Resources Institute recomputed Indonesian output for 1971 to 1984, charging depletion of petroleum, forests and soils:

  conventional GDP growth              7.1 % per year
  depletion-adjusted growth            4.0 % per year
  gap                                  3.1 percentage points per year
  over 13 years, conventional ends    46.5 % larger than adjusted

One change, three stocks, one country, and nearly half the reported expansion turns out to be a transfer from the balance sheet to the income statement.

What already exists to fix it. The System of Environmental-Economic Accounting became an international statistical standard in 2012, with ecosystem accounting following in 2021, and 89 national statistical systems now compile accounts under it. SEEA does not replace GDP. It builds physical and monetary accounts that reconcile line by line with the national accounts, which is why finance ministries can actually read it.

You already know this because no business you respect would report the sale of its buildings as revenue.


BRIEF 7 — Defensive Expenditure

The idea. Some spending restores a condition rather than improving one. It counts in GDP exactly like spending that makes life better.

Commuting costs. Cleaning up a spill. Security. Litigation. The health costs of pollution. Each is a real transaction with real employment attached, and each is a cost of maintaining a position rather than advancing one.

The numbers. Leipert estimated defensive expenditure in West Germany at 5.0 percent of GDP in 1970 and 10.0 percent by 1988. Netting it out would remove about −0.30 percentage points a year from measured growth over that period. For scale in another economy: United States health expenditure was $4,867.0 billion in 2023 against a GDP of $27,720.7 billion — 17.6 percent.

The honest caveat, which is the actual lesson. Health spending is only partly defensive; a hip replacement is not a repair to a damage somebody caused. Where the line falls is a judgement, and no dataset settles it. That is precisely why the figure should be published as a satellite account with its boundary stated rather than folded into a headline.

You already know this because you can tell the difference between a bill you were glad to pay and a bill that put you back where you started.


BRIEF 8 — Composite Indices and the Weight You Cannot See

The idea. A composite index communicates well because it hides its weights, and a weight that is hidden will not be argued with.

Worked example one — the Human Development Index. The HDI is a geometric mean of a health index scaled between 20 and 85 years, an education index, and a log income index scaled between $100 and $75,000. Hold the index constant and ask what income rise exactly offsets one lost year of life expectancy:

  poor country   LE 55 and GNI $1,000     one life year is worth      $68 a head
  rich country   LE 82 and GNI $50,000    one life year is worth   $5,272 a head
  ratio                                                            77.5 x

Nobody wrote that price down. It fell out of equal weighting and a log transform. Ravallion's 2012 paper makes this argument formally.

Worked example two — the GPI's inequality adjustment. The Genuine Progress Indicator weights consumption by an Atkinson inequality index with an aversion parameter ε. On United States quintile shares, the equally-distributed equivalent is 0.8260 of mean income at ε = 0.5 and 0.4240 at ε = 2.0 — a 1.95× swing on a parameter no dataset contains.

Why it matters. Neumayer showed the ISEW's celebrated turning point is produced mainly by its cumulative natural-capital column. A finding produced by a convention is a finding about the convention.

You already know this because you have seen two league tables of the same schools disagree, and understood immediately that the disagreement was in the formula.


BRIEF 9 — The Sufficiency Cutoff

The idea. Any index that counts how many people have enough contains a threshold, and the threshold decides the headline.

Bhutan's Gross National Happiness Index uses the Alkire–Foster method: nine domains, thirty-three indicators, sufficiency thresholds on each, and a person counted as happy when they are sufficient in a stated share of the weighted indicators.

Worked example. Using the published distribution:

  cutoff at 77% of indicators      headcount    8.4 %
  cutoff at 66% — the published    headcount   43.0 %
  cutoff at 50%                    headcount   90.9 %

Moving one threshold takes the national happiness rate from 43.0 percent to 90.9 percent — a 47.9-point swing — without a single person's circumstances changing. Equal weight on nine domains is also a choice, and it is the one nobody argues about because it looks like arithmetic.

What this does not mean. It does not mean the index is worthless. Bhutan publishes the cutoff, the domains and the method, which puts it ahead of most national statistics. And the GNH policy screening tool — which scores proposals before they proceed — does something no dashboard does.

You already know this because you have seen a pass mark moved and watched the pass rate move with it.


BRIEF 10 — The Holdout Indicator

The idea. For every targeted measure, name a second indicator that tracks the same underlying reality, measure it with equal care, and never target it. The divergence between the two is your estimate of the gaming component.

Why it is new. Goodhart's law and Campbell's law are eighty years old between them and have generated an enormous literature of complaint and almost no instrumentation. A holdout produces a number for the thing everyone suspects.

Worked example.

  targeted KPI improvement           18.0 % of the base
  holdout indicator improvement       4.2 % of the base
  divergence                         13.8 percentage points
  real share of the reported gain    23.3 %

Read the last line carefully. It does not say the programme failed. It says roughly a quarter of the reported gain survived contact with an untargeted measure — a finding you can act on, unlike a suspicion.

The one rule that keeps it alive. The holdout must be contractually and culturally excluded from every incentive plan. The moment it enters a bonus formula it becomes a second target, and you have two corrupted series where you had one.

You already know this because you have kept a private measure of how something was really going, alongside the one you reported.


All figures in these briefs are computed in lib/verify/II_10.py, with their inputs printed, and sourced in the chapter's Works Cited.