Haute Lumière
Commerce · I.05 · MMXXVI · daylight
For the person studying this alone, or in a seminar, with no organisation to change yet. You are learning the one skill in this volume that transfers completely intact from a dissertation to a boardroom: the ability to design a comparison that a hostile reader cannot dismantle.
The chapter was written for someone with a site, a budget and a finance director. You may have none of those. What you do have is the thing the chapter spends its whole Arithmetic movement on, and you have it in abundance: decisions you make repeatedly, in conditions that vary, with no idea whether what you do makes any difference.
That is a pilot problem. It is exactly the same problem, at a scale you control completely, with no one to persuade and nothing to lose. And the skill you build on it — power, confounders, pre-registration, the difference between an effect and an artefact — is not a management technique that becomes useful in fifteen years. It is the thing that makes you the most useful person in any room you enter, starting immediately.
There is one more reason to do this now. Almost nobody can do a power calculation on request. It takes four lines and one constant. The person who can say "we would need sixty-three per arm and we have eleven, so this design cannot answer that question" changes what a room is able to decide, and can do it before they have any authority at all.
Exercise 1.1 — The natural experiments in your own life (60 minutes)
The chapter lists five places a firm already has a comparison it has never used. Here they are translated.
| In a firm | In your life |
|---|---|
| Multiple sites doing the same work | Modules, weeks, training sessions, shifts — anything repeated under varying conditions |
| A queue | Anything you have more demand for than supply: study hours, energy, attention |
| A staggered rollout | Any change you will make to several things in some order — the order can be randomised |
| Thirty-six months of history | Your own logs: steps, spending, sleep, submissions, grades, hours |
| A near miss | Something you intended to change and did not, for an unrelated reason |
Write five. Against each, write in one line what the comparison would be — which units are treated, which are not, and what would make them otherwise alike.
Exercise 1.2 — The appreciative interview, aimed at method (45 minutes)
Find someone who evaluates things for a living — a researcher, a clinician, an analyst, an engineer — and ask exactly this:
"Tell me about a time you were convinced something worked and the measurement disagreed. What was the measurement doing that your judgement was not?"
Take notes on the mechanism, not the story. You are collecting confounders, and people who have been caught by one remember it precisely.
Exercise 1.3 — Read one trial properly (2 hours)
Take Finkelstein et al. (2012) on the Oregon Health Insurance Experiment. Do not read it for the findings. Read it for the design: why the lottery existed, what it made possible, what the authors said they could not conclude, and how they described the null results on the clinical markers.
Write 300 words on one question: what did the design make unarguable, and what did it leave open? That distinction is the whole of this chapter.
Exercise 2.1 — Reproduce every figure in the chapter (2 hours)
Open lib/verify/I_05.py, read it, then compute these independently — by hand, in a spreadsheet, or in whatever language you use. Do not take them on trust.
(z_α/2 + z_β) at α = 0.05 two-sided and 80 percent power. Confirm 2.8016.K = 2(z_α/2 + z_β)². Confirm 15.70.m = k = 1. Confirm 56.4 percent.m = 12, k = 3. Confirm 1.81σ, and confirm the ratio to the m = k = 1 case: 2.19 on the effect, 4.80 on sample size.k = 12, ρ = 0.5. Confirm 2.667.φ = 0.65, S = 96,000, C = 120,000. Confirm 0.52 and 1.66 years.Then do the thing that matters most: find one figure in this chapter you can check against an outside source, and check it. A statistics textbook, a statistical package, an online power calculator. If you find a discrepancy, write it down and bring it to your seminar. This edition wants to be checked.
Exercise 2.2 — Your own power calculation (60 minutes)
Take one of the five comparisons from Exercise 1.1 and compute its MDE.
σ, the period-to-period standard deviation.ρ of the series. Inflate σ accordingly.m and k — pre-periods and post-periods you can actually run.detectable effect = 2.8016 · σ · √(1/m + 1/k).Now answer one question in writing: do I believe in an effect that large?
If the answer is no, you have just saved yourself a term of work and learned the most useful thing in this workbook. Change the design — more history, less noisy outcome, more periods — and compute again.
Exercise 2.3 — Manufacture a false positive on purpose (90 minutes)
This is the exercise people remember.
Generate 200 rows of pure random noise in a spreadsheet — two columns, both from the same distribution, no effect whatsoever. Now try to find a significant difference between them. You are allowed to: choose which subset to analyse, drop outliers by a rule you invent afterwards, try several transformations, and stop looking when you find something.
Keep a tally of how many things you tried.
You will find one. Everybody does. Then read Simmons, Nelson and Simonsohn (2011) and discover that this has a name and a measured magnitude.
Write one paragraph on what you noticed about your own honesty while doing it. The finding is not that analysts are dishonest. It is that each individual choice felt entirely reasonable at the time, which is exactly why the constraint has to be structural.
Exercise 2.4 — Find the honest negative (30 minutes)
The chapter's honest negative is the pilot that succeeds and does not scale — the step function, and the voltage drop. In your own words, write why a saving of 990 hours can be completely real and worth nothing to a budget.
Then practise the move: take an improvement you personally believe in and write the strongest honest negative against it. Not a straw version — the one that troubles you.
Exercise 3.1 — The pre-registration (45 minutes)
Build the artifact that does the most work for the least effort in this whole volume. One page:
m before, k after, with dates.Item seven is the one that earns the room. Write it last and write it honestly.
Exercise 3.2 — Randomise something this week (30 minutes)
Take any change you were going to make to several things in sequence — which modules to try a new revision method on, which weeks to change a routine, which of your regular tasks to reorganise.
Randomise the order. Use a spreadsheet function or a coin. Write down the assignment before you start.
You have just converted a plan into evidence at a cost of thirty minutes, and this is the single most transferable habit in the book.
Exercise 3.3 — The present-tense description (60 minutes)
Write 400 words describing, in the present tense, how you make decisions eight years from now.
Constraints:
The last constraint defeats most people. A description of your future in which you are never shown to be wrong is a description of somebody who has stopped measuring.
Exercise 4.1 — The baseline register (45 minutes, then ongoing)
Start one. A single document, one page per metric, each with: the definition, the period, the method, the verifier, and the date.
Three metrics is enough to begin. Keep it for a year and you will have something almost nobody your age has: a comparable series, which is the raw material of every claim you will ever make.
Exercise 4.2 — The second reader (this week)
Find one person who will read your analysis looking for the fault, and ask them explicitly to try to break it. Give them the pre-registration first.
The chapter's governance rule applies at every scale: the person who did the work is the one person who cannot see the flaw in it. This is structural, not a comment on your ability. Build the habit now, when it costs nothing, rather than when a result matters.
Exercise 4.3 — Delight, on purpose (ongoing)
Delight is the adoption mechanism, not a reward. Choose one element of this practice and make it genuinely pleasant: the spreadsheet you like the look of, the hour, the place, the particular satisfaction of a chart that goes back far enough to be interesting.
Write one sentence: the part of this I look forward to is ___. If you cannot complete it, redesign the practice until you can.
Choose one system you genuinely control — your own study, a society you run, a team you play in, a household, a job — and run a full designed comparison on it.
Deliverables.
σ from at least twelve periods of history, ρ estimated and applied, m and k stated, the MDE computed. State plainly whether you believe in an effect that large.How it is assessed. Not on whether the intervention worked. On whether the design could have detected it, and on whether the analysis matches the pre-registration.
A null result from a well-powered, pre-registered design is a first-class piece of work. A positive result from an under-powered design chosen after the fact is not a result at all.
Score yourself honestly. This is for you.
| Not yet | Beginning | Solid | Fluent | |
|---|---|---|---|---|
| I can compute a minimum detectable effect from a series of history | ||||
| I know what power my design has before I run it | ||||
| I can explain regression to the mean to someone who has never heard of it | ||||
| I check for serial correlation before trusting a σ | ||||
| I know why randomising sites is not the same as randomising people | ||||
| I pre-register the analysis and I do not deviate from it | ||||
| I can tell a measured saving from a banked one | ||||
| I ask someone else to try to break my result |
The two that matter most are the second and the sixth. Everything else can be looked up. Those two are habits, and habits take a term.
What you will have at the end of ninety days, if you do this properly:
K = 15.70. Being able to produce it in a meeting is a form of authority that requires no title.That last one is worth more than it sounds. An analyst who has written down in advance what would change their mind is believed about everything else — and you can be that person long before anybody gives you anything to decide.