Haute Lumière

Commerce · II.10 · MMXXVI · daylight

La Bourse  /  Volume II  /  Nº II.10  /  Quiz, reflection, essays

A man in a dark suit writing in a ledger beside printed charts, a library of books behind him.
Plate II.10 · Quiz, reflection, essaysThe Weight and the Weighing.The balance does not describe the cloth. It describes the cloth in the one respect somebody decided to care about, and then the ledger forgets that a decision was made.

ASSESSMENT · Chapter II.10 — Measurement and What It Does

Three instruments: a ten-point quiz, eight reflection questions, five essay prompts. The quiz checks comprehension rather than recall. The reflections are private and first-person. The essays are arguable from more than one side.


THE QUIZ — ten points

Four on recall.

1. State Goodhart's law in its original 1975 form, and say how it differs from the popular phrasing.

"Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." The popular version — when a measure becomes a target it ceases to be a good measure — is Marilyn Strathern's 1997 gloss. One mark for the original wording, one for the difference: Goodhart names the mechanism, which is that the evidence justifying the target was gathered under conditions the target destroys.

2. State Campbell's law and explain why it is the stronger claim.

"The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor." Goodhart says the statistic degrades; Campbell says the measured process itself degrades — the care changes, not merely the waiting-time data about it.

3. What did Kuznets write in 1934 about national income and welfare, and what argument did he later lose?

"The welfare of a nation can, therefore, scarcely be inferred from a measurement of national income as defined above." He lost the wartime argument over government war expenditure, which he wanted treated as an intermediate cost rather than final output. Credit any answer that also names the distribution of income or the unpleasantness of effort, both of which he raised in the same document.

4. Name three things the production boundary excludes, and one non-market service it includes.

Excludes: own-account household services (childcare, cooking, laundry, elder care); depletion of natural stocks; most volunteering. Includes: the imputed rental value of owner-occupied housing, roughly 7.0 percent of United States GDP. The point of the question is the last part — the boundary is a list, not a principle.

Four on application.

5. A colleague says the exclusions do not matter because non-market activity cannot be valued. Answer them.

The accounts already value an enormous non-market service — owner-occupied housing — by a replacement-cost method that would work equally well on childcare. The BEA satellite account values household production at 25.7 percent of GDP in 2010; the ONS values unpaid household service work at 63.8 percent of UK GDP in 2016. Full marks require naming a published satellite account, not merely asserting that valuation is possible.

6. Your division adopts a customer-satisfaction target. Design the holdout indicator and say what would destroy it.

A second indicator tracking the same reality, independently sourced where possible — repeat purchase rate, unprompted referral, complaint escalation — measured with equal care and never targeted, never in an incentive plan. What destroys it: putting it into a bonus formula, at which point it becomes a second target and both series are corrupted.

7. A team reports an 18.0 percent improvement in a targeted KPI while the holdout indicator moved 4.2 percent. What do you conclude, and what do you not conclude?

Divergence of 13.8 percentage points; roughly 23.3 percent of the reported gain survived contact with an untargeted measure. Conclude that about a quarter of the gain is supported. Do not conclude that the programme failed or that anyone was dishonest. The commonest wrong answer here is to question the holdout, which is the natural response and wrong every time.

8. Why does the chapter argue for satellite accounts rather than a better single index?

A single replacement index repeats the original error — it hides its weights and invites Goodhart. A satellite keeps the boundary visible, keeps the main series internationally comparable, and can be built by one team in a quarter without altering a number anyone depends on. Credit any answer naming SEEA, now a standard in 89 national statistical systems, as the working case.

Two that require the arithmetic to be done.

9. Using the HDI formula — health index (LE − 20) / 65, income index (ln GNI − ln 100) / (ln 75,000 − ln 100), combined as a geometric mean — compute what rise in GNI per head exactly offsets one lost year of life expectancy for a country at 60 years and $4,000. Compare it with the rich-country case in the chapter. Show your working.

Inside a geometric mean, holding the index constant requires the proportional loss in the health index to equal the proportional gain in the income index. I_health = 40/65 = 0.6154; one year is 1/65 = 0.015385, a proportional loss of 0.025. I_income = ln(40)/ln(750) = 0.5572, so the income index must rise by 0.5572 × 0.025 = 0.013931, i.e. Δln GNI = 0.013931 × 6.6201 = 0.09222, i.e. GNI must rise by e^0.09222 − 1 = 9.66% — about $386 a head. The rich case in the chapter is $5,272, a ratio of 13.6×. The mark is for the method. The finding is that the index contains a price list nobody wrote down.

10. A borrower issues a $300,000,000 sustainability-linked facility with a 4.0-year observation period. Real compliance costs $9,000,000; reclassifying to hit the KPI costs $500,000. What step-up deters gaming, and how does it compare with the market-standard 25 basis points?

Required step-up = (9,000,000 − 500,000) / (300,000,000 × 4.0) = 0.00708 = 71 basis points. The market standard of 25 bps is worth 0.0025 × 300,000,000 × 4.0 = $3,000,000 against a gaming advantage of $8,500,000 — short by 2.8 times. The stronger answer states the conclusion in the right register: the convention under-deters, and the fix is a number on a term sheet, not an appeal to integrity.


REFLECTION — eight questions, for one person and a pen

These are not for a room. Write the answers by hand if you can; the slowness is the point.

  1. What is the number you are currently measured on? Write down what it counts, what it leaves out, and — as far as you can tell — who chose.
  1. Recall a time you moved a metric without improving the thing the metric was for. What made that the rational choice at the time?
  1. What have you stopped doing because it did not show up anywhere? Was the decision a good one?
  1. Where do you keep a private measure of how something is really going, alongside the one you report? What does the private one measure that the reported one cannot?
  1. Think of a number you trust completely. What is it about how it is compiled that earns the trust — and have you ever checked that those conditions still hold?
  1. Which of your own figures could you publish the boundary of tomorrow, and what is the honest reason you have not?
  1. When somebody last produced a number that ended an argument you were having, what would you have asked if you had had the three questions ready?
  1. What in your life is worth roughly $21,840 a year at replacement cost and appears in no account anywhere? Name it precisely, then decide whether you want it counted.

ESSAY PROMPTS — five

Each is arguable from more than one side. Each requires at least one source the chapter cites and at least one it does not.

1. Is Campbell's law a law? Campbell's claim is universal in form and was offered without formal proof. Argue either that it is a robust empirical regularity with a mechanism, or that it is an observation about badly designed incentive systems that well-designed ones escape. Engage Campbell directly and Bevan and Hood on the English health service, plus one empirical study of a target regime that the chapter does not cite.

2. The war that built GDP. The chapter argues that GDP is not a failed welfare measure but a successful measure of mobilisable output, settled under wartime conditions in an argument Kuznets lost. Argue either that this history explains the accounts' present shape and constrains reform, or that the shape has been re-legitimated so many times since that its origin is now irrelevant. Use Kuznets (1934) and Coyle, and one history of national accounting the chapter does not cite.

3. Which judgement would you import? Every alternative index imports a value judgement: the GPI an inequality aversion parameter worth a 1.95× swing, the HDI a set of weights implying a 77.5× difference in the price of a life year, GNH a cutoff worth 47.9 points of headcount, SEEA exchange value in place of welfare value. Choose one and argue that its imported judgement is defensible — not invisible, but correct — then state precisely what evidence would change your mind. Use Ravallion or Neumayer, and one source the chapter does not cite.

4. Satellite or substitute. The chapter argues for satellite accounts over replacement indices, on the grounds that SEEA has the plumbing — a standard, a manual, a revision policy, independent compilers — and the GPI does not. Argue the counter-case: that incremental satellite accounts are absorbed without changing any decision, and that only a rival headline number can displace a headline number. Use the SEEA framework documents and Kubiszewski et al., plus one account of an actual policy change driven by a measurement the chapter does not cite.

5. Can a holdout survive? The holdout indicator is proposed here as the anti-Goodhart mechanism, on the condition that it is never targeted and never enters an incentive plan. Argue either that this condition is institutionally achievable — naming the governance that would hold it — or that any indicator a senior decision-maker watches becomes a target regardless of what the contract says. Use Goodhart and Strathern, and one source from the literature on audit, gaming or performance management that the chapter does not cite.