Haute Lumière

Commerce · II.10 · MMXXVI · daylight

La Bourse  /  Volume II  /  Nº II.10  /  Workbook — the student

A man in a dark suit writing in a ledger beside printed charts, a library of books behind him.
Plate II.10 · Workbook — the studentThe Weight and the Weighing.The balance does not describe the cloth. It describes the cloth in the one respect somebody decided to care about, and then the ledger forgets that a decision was made.

WORKBOOK — THE STUDENT

Chapter II.10 · Measurement and What It Does

For the person studying this alone, or in a seminar, with no organisation to change yet. You are surrounded by measures of yourself — grades, credits, rankings, follower counts, step counts — and almost none of them were designed by anyone who would meet you. This is the term you learn to read them.


WHY THIS WORKBOOK IS DIFFERENT

The chapter was written for someone who can commission a satellite account and sign a term sheet. You can do neither yet.

What you can do is far more valuable at your stage, and almost nobody does it: you can acquire the reflex before the habits set. Three questions, asked of every number that crosses your desk for one term — what does this count, what does it leave out, who chose — will change how you read for the next forty years. The instrument is free, it takes twenty seconds, and the people who have it are visibly different in a room from the people who do not.

You are also, right now, living inside the single cleanest natural experiment in Campbell's law available to anyone: a grading system. You have already watched yourself and everyone around you optimise against it. You have the data. What you lack is the vocabulary, and that is what this term supplies.


PART ONE — DISCOVERY

Days 1–30: find the measures already around you

Exercise 1.1 — The measurement inventory (90 minutes)

List every number by which you are currently assessed, ranked, sorted or compared. Be thorough and unflattering. Grade point average. Credits. Word counts. Attendance. Reading lists completed. Hours billed at a part-time job. Followers. Streaks. Times at a sport. Sleep scores.

For each, write three lines and nothing else:

  1. Counts — what the number actually registers.
  2. Excludes — what it cannot see, however good you get at it.
  3. Chose — who set it, when, and for what purpose.

The third column is the hard one and it is the point. Most of these numbers were designed by someone solving an administrative problem, not by someone trying to describe you. A measure designed for administration will describe you administratively. That is not a conspiracy; it is a boundary.

Exercise 1.2 — The appreciative measurement interview (45 minutes)

Find someone whose judgement you trust — a supervisor, a tutor, a practitioner — and ask exactly this:

"Tell me about a measure in your field that you genuinely trust. Not one you put up with — one you believe. What is it about how it is compiled that earns that?"

Then stay quiet and take notes on the compilation, not the number. You are collecting the conditions under which measurement works: independent compilers, published method, a revision policy, a stated boundary, a long series. Do this five times over the term and you will have an education in the difference between a number and an instrument.

Exercise 1.3 — Find a boundary statement in the wild (60 minutes)

Go to a national statistical office — the ONS, BEA, Eurostat, your own country's — and find a published methodology note for a headline figure. Read the section that says what is excluded. Write down one exclusion you did not expect.

Most people have never once read one of these, and they are written in plain language by people who wanted to be understood. This exercise mostly teaches you that the honesty was always there and nobody went to look.


PART TWO — THE ARITHMETIC

Days 31–50: compute before you argue

Exercise 2.1 — Reproduce the chapter's figures (2–3 hours)

Open lib/verify/II_10.py, read it, and then compute these independently, by hand or in a spreadsheet:

  1. The Goodhart overstatement. At ρ = 0.9, a one standard-deviation rise in the proxy predicts ρ² = 0.81 standard deviations of real gain. If a fifth of the response is real, the overstatement factor is 4.05. Confirm it, then do it at ρ = 0.7 and at ρ = 0.8.
  2. Indonesia. Conventional growth 7.1 percent a year, depletion-adjusted 4.0 percent, thirteen years. Confirm that the conventional account ends 46.5 percent larger than the adjusted one.
  3. The HDI's price of a life year. Health index (LE − 20)/65, income index (ln GNI − ln 100)/(ln 75,000 − ln 100), geometric mean. Confirm $68 a head at 55 years and $1,000, and $5,272 at 82 years and $50,000. Ratio 77.5×.
  4. The Atkinson swing. United States quintile shares, equally-distributed equivalent 0.8260 of the mean at ε = 0.5 and 0.4240 at ε = 2.0. Ratio 1.95×.

Then do the thing that matters most: take one figure in this chapter and check it against a source the chapter does not cite. If you find a discrepancy, write it down and bring it to your seminar. This edition wants to be checked.

Exercise 2.2 — Value your own household production (60 minutes)

For one week, log the hours you spend on unpaid work that somebody could be paid to do: cooking, cleaning, laundry, shopping, care of another person, unpaid transport, unpaid administration for a family member.

Then apply a replacement wage — what it would cost to hire someone locally. The chapter's illustrative figures are 28.0 hours a week at $15.00 an hour, which is $21,840 a year.

  hours per week  ×  replacement wage  ×  52  =  annual replacement value

Write down your number. Then write one sentence on what changes when you can say it out loud.

Exercise 2.3 — Build your own composite, then break it (90 minutes)

Construct a three-component index of something you care about — your own week, your degree, a city you might move to. Pick three indicators, normalise each to a zero-to-one scale, and combine them.

Now do the part the exercise exists for. Change the weights and watch the ranking change. Try equal weights, then 0.5/0.3/0.2, then a geometric mean instead of an arithmetic one. Record how many of your rankings survive.

Then answer in writing: if I published only the final score, what would a reader be unable to argue with?


PART THREE — DESIGN

Days 51–70: build the apparatus

Exercise 3.1 — Write three boundary statements (45 minutes)

Take the three most important numbers from your inventory in Exercise 1.1 and write a proper boundary statement for each: what it counts, what it excludes, the decision rule that draws the line, and who chose.

Put them somewhere you will see them. This is the Kuznets standard, applied at personal scale, and it is the cheapest thing in this workbook.

Exercise 3.2 — Run a holdout on yourself (four weeks)

Choose one thing you are currently optimising — a grade, a word count, a running time, a revision schedule.

Now name a holdout: a second indicator that tracks the same underlying reality and that you will never target. If you are optimising word count, the holdout might be the number of paragraphs you would still defend a month later. If you are optimising a grade, the holdout might be how much of the material you can explain to someone else without notes.

Measure both weekly for four weeks. Target only the first. Then compute:

  divergence  =  targeted gain  −  holdout gain
  real share  =  holdout gain  /  targeted gain

The chapter's worked case has a targeted gain of 18.0 percent, a holdout gain of 4.2 percent, a divergence of 13.8 percentage points and a real share of 23.3 percent. Yours will be different. Whatever it is, you now have a number for something most people only suspect.

Exercise 3.3 — The parameter register (45 minutes)

Pick any published index that gets quoted at you — a university ranking, a liveability index, a sustainability score, a credit score if you can find its method. Write a one-page register entry: the parameters, their values, who set them, and what the result looks like at two other plausible values.

Where the method is not published, that is your finding, and it is a finding worth stating plainly.


PART FOUR — DESTINY AND DELIGHT

Days 71–90: make it a reflex

Exercise 4.1 — The twenty-second habit (ongoing)

For the rest of the term, every time a number is used to end an argument in your presence, ask the three questions internally: what does this count, what does it leave out, who chose. Once a week, ask them out loud.

Keep a tally of how often the person quoting the number knows the answer. The tally itself becomes one of the more interesting artifacts of your term.

Exercise 4.2 — Ask them of your own work first (ongoing)

The practice collapses into cynicism unless it is aimed inward at least as often as outward. Before you present any figure — in an essay, a presentation, a job application — write its boundary statement first.

You will find that roughly one figure in five does not survive the exercise. It is better that you find this than that a marker does.

Exercise 4.3 — Delight, on purpose (30 minutes)

Find one index in the world that you think is beautifully made — genuinely well designed, honest about its boundary, useful. Write three hundred words on why.

This is not a soft exercise. The whole chapter risks curdling into the belief that all numbers lie, which is both false and useless. Cynicism is the belief that numbers mean nothing. The useful belief is that they mean something specific, and that you can find out what. One admired example is what keeps you on the right side of that line.


A NOTE ON READING BADLY MADE MEASURES CHARITABLY

One thing will happen to you this term and it is worth naming in advance. Once the three questions become a reflex, every number starts to look flimsy, and there is a stage — usually around week six — where the honest conclusion seems to be that nothing can be measured and everyone is pretending.

That stage is a symptom of having learned half the chapter. The other half is this: almost every measure you will take apart was made by someone with less time, less data and less freedom than you have in the exercise. The administrator who set the attendance rule was solving a real problem under a real constraint. The statistician who drew the production boundary in 1953 was building something that had never existed, in a hurry, for a world that needed it.

So the discipline is to write, for every measure you criticise, one sentence naming what it does well and one naming the constraint it was built under. It costs two lines. It keeps the practice from turning into the thing it was meant to cure — a confident judgement made from outside, with the boundary unstated.


THE TERM PROJECT

One measure, followed all the way down

Choose a single published measure that affects people you can talk to — a university ranking, a hospital league table, a school performance table, a sustainability score, a neighbourhood deprivation index — and take it apart.

Deliverables.

  1. The boundary memo (seven hundred words). What it counts, what it excludes, the decision rule, and the documentary evidence for each. Cite the methodology note.
  2. The parameter register (1 page). Every weight, threshold and transform, with the result recomputed at two alternative values. Show your working.
  3. The Campbell search (seven hundred words). Find and document at least one behavioural response to the measure among the people it measures. Interview two of them, appreciatively: when has this measure helped you do something you were proud of? — and then what did you stop doing because of it?
  4. The holdout proposal (1 page). Name an indicator that tracks the same reality and could plausibly be protected from targeting. Say who would have to protect it.
  5. The reflection (six hundred words). What you expected, what you found, and one thing you now believe less strongly than you did.

How it is assessed. Not on whether you exposed a scandal. On whether your recomputation is reproducible, whether your interviews were appreciative rather than prosecutorial, and whether you stated at least one thing the measure does well.

A project that concludes the measure is broadly sound, with the arithmetic shown, is a first-class piece of work. A project that concludes it is corrupt and cannot be reproduced is not.


SELF-ASSESSMENT

Not yetBeginningSolidFluent
I can state Goodhart's law in its original form and say what it means
I can say why Campbell's claim is stronger, and give a documented case
I ask what a number counts and excludes before I use it
I ask it of my own figures first
I can recompute a composite index at different weights
I can design a holdout indicator and say what would destroy it
I can name a measure I admire and say precisely why
I can read a methodology note without being told to

The two that matter most are the fourth and the seventh. The fourth is what keeps the practice honest. The seventh is what keeps it from becoming corrosive.


CARRYING IT FORWARD

By the end of the term you will have:

That last one is worth more than it sounds. The person in the room who can say "that ranking is mostly the weight on research income, and here it is at a different weight" changes what the room is able to decide — and you can be that person with no authority at all, which is the fastest route to being given some.


APPRECIATIVE QUESTIONS FOR YOUR SEMINAR

  1. Which measure in our field do we genuinely trust, and what is it about how it is compiled that earns that trust?
  2. When has one of us been measured well — in a way that recognised something real? What made that measurement work?
  3. What would this course look like if its assessment had a holdout indicator nobody was allowed to teach to?
  4. If each of us wrote a boundary statement for the one number we are most judged on, what would we learn about this institution?