Haute Lumière

Commerce · I.05 · MMXXVI · daylight

La Bourse  /  Volume I  /  Nº I.05  /  Workbook — the student

A woman standing alone in an empty room before a wide window of trees, sepia light running across the floor toward her.
Plate I.05 · Workbook — the studentTwo Rooms, One Question.A pilot is not a small version of the plan. It is a comparison you built on purpose, in a place where the only difference is the one you made.

WORKBOOK — THE STUDENT

Chapter I.05 · The Pilot That Pays

For the person studying this alone, or in a seminar, with no organisation to change yet. You are learning the one skill in this volume that transfers completely intact from a dissertation to a boardroom: the ability to design a comparison that a hostile reader cannot dismantle.


WHY THIS WORKBOOK IS DIFFERENT

The chapter was written for someone with a site, a budget and a finance director. You may have none of those. What you do have is the thing the chapter spends its whole Arithmetic movement on, and you have it in abundance: decisions you make repeatedly, in conditions that vary, with no idea whether what you do makes any difference.

That is a pilot problem. It is exactly the same problem, at a scale you control completely, with no one to persuade and nothing to lose. And the skill you build on it — power, confounders, pre-registration, the difference between an effect and an artefact — is not a management technique that becomes useful in fifteen years. It is the thing that makes you the most useful person in any room you enter, starting immediately.

There is one more reason to do this now. Almost nobody can do a power calculation on request. It takes four lines and one constant. The person who can say "we would need sixty-three per arm and we have eleven, so this design cannot answer that question" changes what a room is able to decide, and can do it before they have any authority at all.


PART ONE — DISCOVERY

Days 1–20: find the comparisons you already have

Exercise 1.1 — The natural experiments in your own life (60 minutes)

The chapter lists five places a firm already has a comparison it has never used. Here they are translated.

In a firmIn your life
Multiple sites doing the same workModules, weeks, training sessions, shifts — anything repeated under varying conditions
A queueAnything you have more demand for than supply: study hours, energy, attention
A staggered rolloutAny change you will make to several things in some order — the order can be randomised
Thirty-six months of historyYour own logs: steps, spending, sleep, submissions, grades, hours
A near missSomething you intended to change and did not, for an unrelated reason

Write five. Against each, write in one line what the comparison would be — which units are treated, which are not, and what would make them otherwise alike.

Exercise 1.2 — The appreciative interview, aimed at method (45 minutes)

Find someone who evaluates things for a living — a researcher, a clinician, an analyst, an engineer — and ask exactly this:

"Tell me about a time you were convinced something worked and the measurement disagreed. What was the measurement doing that your judgement was not?"

Take notes on the mechanism, not the story. You are collecting confounders, and people who have been caught by one remember it precisely.

Exercise 1.3 — Read one trial properly (2 hours)

Take Finkelstein et al. (2012) on the Oregon Health Insurance Experiment. Do not read it for the findings. Read it for the design: why the lottery existed, what it made possible, what the authors said they could not conclude, and how they described the null results on the clinical markers.

Write 300 words on one question: what did the design make unarguable, and what did it leave open? That distinction is the whole of this chapter.


PART TWO — THE ARITHMETIC

Days 21–45: compute before you argue

Exercise 2.1 — Reproduce every figure in the chapter (2 hours)

Open lib/verify/I_05.py, read it, then compute these independently — by hand, in a spreadsheet, or in whatever language you use. Do not take them on trust.

  1. (z_α/2 + z_β) at α = 0.05 two-sided and 80 percent power. Confirm 2.8016.
  2. K = 2(z_α/2 + z_β)². Confirm 15.70.
  3. The power of a 3σ effect at m = k = 1. Confirm 56.4 percent.
  4. The detection threshold at m = 12, k = 3. Confirm 1.81σ, and confirm the ratio to the m = k = 1 case: 2.19 on the effect, 4.80 on sample size.
  5. The AR(1) inflation at k = 12, ρ = 0.5. Confirm 2.667.
  6. The fund's regeneration rate and doubling time at φ = 0.65, S = 96,000, C = 120,000. Confirm 0.52 and 1.66 years.

Then do the thing that matters most: find one figure in this chapter you can check against an outside source, and check it. A statistics textbook, a statistical package, an online power calculator. If you find a discrepancy, write it down and bring it to your seminar. This edition wants to be checked.

Exercise 2.2 — Your own power calculation (60 minutes)

Take one of the five comparisons from Exercise 1.1 and compute its MDE.

  1. Gather the history. You need at least twelve periods of the outcome under ordinary conditions.
  2. Compute σ, the period-to-period standard deviation.
  3. Compute the lag-1 correlation ρ of the series. Inflate σ accordingly.
  4. Decide m and k — pre-periods and post-periods you can actually run.
  5. Compute detectable effect = 2.8016 · σ · √(1/m + 1/k).

Now answer one question in writing: do I believe in an effect that large?

If the answer is no, you have just saved yourself a term of work and learned the most useful thing in this workbook. Change the design — more history, less noisy outcome, more periods — and compute again.

Exercise 2.3 — Manufacture a false positive on purpose (90 minutes)

This is the exercise people remember.

Generate 200 rows of pure random noise in a spreadsheet — two columns, both from the same distribution, no effect whatsoever. Now try to find a significant difference between them. You are allowed to: choose which subset to analyse, drop outliers by a rule you invent afterwards, try several transformations, and stop looking when you find something.

Keep a tally of how many things you tried.

You will find one. Everybody does. Then read Simmons, Nelson and Simonsohn (2011) and discover that this has a name and a measured magnitude.

Write one paragraph on what you noticed about your own honesty while doing it. The finding is not that analysts are dishonest. It is that each individual choice felt entirely reasonable at the time, which is exactly why the constraint has to be structural.

Exercise 2.4 — Find the honest negative (30 minutes)

The chapter's honest negative is the pilot that succeeds and does not scale — the step function, and the voltage drop. In your own words, write why a saving of 990 hours can be completely real and worth nothing to a budget.

Then practise the move: take an improvement you personally believe in and write the strongest honest negative against it. Not a straw version — the one that troubles you.


PART THREE — DREAM AND DESIGN

Days 46–70: build the apparatus

Exercise 3.1 — The pre-registration (45 minutes)

Build the artifact that does the most work for the least effort in this whole volume. One page:

  1. The question, in one sentence.
  2. The primary outcome — exactly one, defined so that another person would compute it identically.
  3. The secondary outcomes, listed now, labelled secondary.
  4. The comparison: which units treated, which not, how assigned.
  5. The periods: m before, k after, with dates.
  6. The exclusion rules, written before any exclusion is tempting.
  7. The failure condition: what result would make you say this did not work.
  8. Date it and give it to one named person.

Item seven is the one that earns the room. Write it last and write it honestly.

Exercise 3.2 — Randomise something this week (30 minutes)

Take any change you were going to make to several things in sequence — which modules to try a new revision method on, which weeks to change a routine, which of your regular tasks to reorganise.

Randomise the order. Use a spreadsheet function or a coin. Write down the assignment before you start.

You have just converted a plan into evidence at a cost of thirty minutes, and this is the single most transferable habit in the book.

Exercise 3.3 — The present-tense description (60 minutes)

Write 400 words describing, in the present tense, how you make decisions eight years from now.

Constraints:

The last constraint defeats most people. A description of your future in which you are never shown to be wrong is a description of somebody who has stopped measuring.


PART FOUR — DESTINY AND DELIGHT

Days 71–90: make it hold, and enjoy it

Exercise 4.1 — The baseline register (45 minutes, then ongoing)

Start one. A single document, one page per metric, each with: the definition, the period, the method, the verifier, and the date.

Three metrics is enough to begin. Keep it for a year and you will have something almost nobody your age has: a comparable series, which is the raw material of every claim you will ever make.

Exercise 4.2 — The second reader (this week)

Find one person who will read your analysis looking for the fault, and ask them explicitly to try to break it. Give them the pre-registration first.

The chapter's governance rule applies at every scale: the person who did the work is the one person who cannot see the flaw in it. This is structural, not a comment on your ability. Build the habit now, when it costs nothing, rather than when a result matters.

Exercise 4.3 — Delight, on purpose (ongoing)

Delight is the adoption mechanism, not a reward. Choose one element of this practice and make it genuinely pleasant: the spreadsheet you like the look of, the hour, the place, the particular satisfaction of a chart that goes back far enough to be interesting.

Write one sentence: the part of this I look forward to is ___. If you cannot complete it, redesign the practice until you can.


THE TERM PROJECT

One comparison, designed and run

Choose one system you genuinely control — your own study, a society you run, a team you play in, a household, a job — and run a full designed comparison on it.

Deliverables.

  1. The comparison memo (700 words). What you are comparing, why those units are otherwise alike, and which of the five natural-experiment structures you are using. Name the confounder that worries you most.
  2. The power calculation (1 page plus workings). σ from at least twelve periods of history, ρ estimated and applied, m and k stated, the MDE computed. State plainly whether you believe in an effect that large.
  3. The pre-registration (1 page, dated, signed, given to a named person before you start). All eight items from Exercise 3.1.
  4. The intervention (run, not described). With a log of the date, the co-interventions, and anything unusual.
  5. The analysis (600 words). Exactly the analysis you pre-registered. Any additional analysis appears in a clearly labelled second section titled Exploratory.
  6. The reflection (600 words). What you expected, what happened, where you were tempted to deviate from the pre-registration and what you did about it.

How it is assessed. Not on whether the intervention worked. On whether the design could have detected it, and on whether the analysis matches the pre-registration.

A null result from a well-powered, pre-registered design is a first-class piece of work. A positive result from an under-powered design chosen after the fact is not a result at all.


SELF-ASSESSMENT

Score yourself honestly. This is for you.

Not yetBeginningSolidFluent
I can compute a minimum detectable effect from a series of history
I know what power my design has before I run it
I can explain regression to the mean to someone who has never heard of it
I check for serial correlation before trusting a σ
I know why randomising sites is not the same as randomising people
I pre-register the analysis and I do not deviate from it
I can tell a measured saving from a banked one
I ask someone else to try to break my result

The two that matter most are the second and the sixth. Everything else can be looked up. Those two are habits, and habits take a term.


CARRYING IT FORWARD

What you will have at the end of ninety days, if you do this properly:

That last one is worth more than it sounds. An analyst who has written down in advance what would change their mind is believed about everything else — and you can be that person long before anybody gives you anything to decide.


APPRECIATIVE QUESTIONS FOR YOUR SEMINAR

  1. When has one of us designed a comparison before acting, and what did it let us say afterwards that we could not otherwise have said?
  2. Which series available to this group goes back the furthest, and what could we now ask of it?
  3. What is already working in how this seminar decides what to believe — and what makes it work?
  4. If every one of us randomised one decision order this term, what would this group collectively know by the summer?