Haute Lumière
Commerce · I.05 · MMXXVI · daylight
One page each. A reader who reads only these ten pages has the chapter.
The idea. Power is the probability that your pilot detects an effect that is genuinely there. It is not a property of the truth. It is a property of your design, and you choose it before you start.
Four things set it: the size of the effect, the noise in the measurement, the number of observations, and how sure you insist on being. Fix any three and the fourth follows.
detectable effect = (z_α/2 + z_β) · σ · √(1/m + 1/k)
= 2.8016 · σ · √(1/m + 1/k)
m = periods before k = periods after σ = period-to-period sd
Worked example. One month before, one month after: the bar is 3.96σ. Twelve months before, three after: 1.81σ. Same intervention, same site, same cost — less than half the effect required, purely because of how the comparison was arranged.
The trap. A pilot with low power that returns "no significant effect" has told you nothing. It has not shown the intervention failed. Most pilots that are reported as failures were never capable of succeeding, and that is a design fault, not a finding.
You already know this because you have watched somebody call one good week proof and one bad week noise, and you knew that both claims came from the same person for the same reason.
The idea. Before you spend anything, compute the smallest improvement your design could see. Then ask whether you believe in an improvement that large.
The MDE turns power on its head. Instead of how likely am I to detect my hoped-for effect, it asks what is the smallest effect I could detect at all — and it is a far better question, because it can be answered in week two with data you already have.
n per arm = 2 (z_α/2 + z_β)² (σ / δ)² = 15.70 (σ / δ)²
σ / δ = 1.0 -> 16 per arm
σ / δ = 1.5 -> 36 per arm
σ / δ = 2.0 -> 63 per arm
σ / δ = 3.0 -> 142 per arm
Worked example. Monthly variation 12 percent; you want to detect a 6 percent improvement. σ/δ = 2, so you need 63 units per arm — 126 in total. If you have eleven branches, you cannot run this pilot as designed, and you have learned that for the price of an afternoon.
What to do when the MDE is too big. Reduce σ (measure better, stratify, use a longer pre-period), increase n (more sites, more weeks), or change the outcome to one with less noise in it. Halving σ is worth as much as quadrupling n and is usually cheaper.
Why it matters. The MDE is the cheapest piece of honesty in the whole method. It is the number that stops you spending £120,000 to learn nothing.
The idea. The most powerful thing you can add to a pilot costs nothing, because it already happened.
The standard error of a before-and-after comparison is σ√(1/m + 1/k). The post-periods are expensive — every one is a month of running the intervention. The pre-periods are free: they are history, sitting in a system somebody maintains.
m pre k post detect at power at a 3σ effect
------------------------------------------------------
1 1 3.96 σ 56.4 %
1 3 3.23 σ 73.8 %
3 3 2.29 σ 95.7 %
6 3 1.98 σ 98.9 %
12 3 1.81 σ 99.6 %
Worked example. Going from one pre-period to twelve moves the bar from 3.96σ to 1.81σ. That is a factor of 2.19 on the effect and 4.80 on the equivalent sample size — a pilot almost five times the size, for the cost of an email to whoever owns the reporting system.
The condition. The history must be comparable. A system upgrade, a definitional change, a reorganisation or a merger inside the pre-period breaks it. Check the series for breaks before you trust it, and say in the pre-registration which months you are using and why.
Why it matters. Most pilots are under-powered and over-budgeted at the same time. The correction is not more money. It is thirty-six months of a spreadsheet nobody thought to ask for.
The idea. If you choose the worst-performing site, it will improve whether or not you do anything. The improvement is arithmetic, not achievement.
Any measurement is part signal and part noise. Let ρ be the reliability — the share of observed variation that is real. A unit selected because it scored badly was partly unlucky, and luck does not persist.
expected 'improvement' with no intervention = (1 − ρ) × (distance from mean)
percentile chosen reliability ρ free improvement
-------------------------------------------------------
10th 0.6 0.51 σ
10th 0.8 0.26 σ
5th 0.6 0.66 σ
25th 0.6 0.27 σ
Worked example. You run the pilot at the depot with the worst cost per drop. Reliability is 0.6. The depot improves by half a standard deviation next quarter on its own. If your intervention is worth 0.4σ, you will report 0.9σ and be wrong by more than twice the truth.
The fixes, in order. Randomise which sites get it. Or select on a period you then exclude from the analysis. Or — best and cheapest — keep a control group, which absorbs the regression identically in both arms and cancels it.
You already know this because you have seen a terrible quarter followed by a good one, and a manager credited for the recovery who did nothing but arrive.
The idea. Monthly figures are not independent. A good month follows a good month, so the variation you measure understates the uncertainty you face.
Under an AR(1) process with correlation ρ, the variance of a k-period mean is inflated:
k periods ρ inflation worth this many independent periods
------------------------------------------------------------------------
12 0.0 1.00 12.0
12 0.3 1.76 6.8
12 0.5 2.67 4.5
12 0.7 4.39 2.7
Worked example. Twelve correlated months at ρ = 0.5 carry the information of four and a half independent ones. A power calculation that treats them as twelve overstates its precision by 63 percent and will produce a confident result that does not replicate.
The evidence. Bertrand, Duflo and Mullainathan (2004) showed that difference-in-differences estimates ignoring serial correlation rejected a true null far more often than the nominal 5 percent — in their simulations, at rates several times higher.
What to do. Estimate ρ from your own series — one line in a spreadsheet — and inflate σ before sizing. Or aggregate the whole pre-period and the whole post-period into single means and compare those, which is cruder and honest.
Why it matters. This is the commonest way a careful analyst gets a wrong answer while doing everything else correctly.
The idea. You randomise sites. Your sample size is the number of sites, not the number of people inside them.
People in the same site share a manager, a rota, a building and a local market. Their outcomes are correlated. The intra-cluster correlation ICC measures how much, and the design effect converts people into effective observations.
DEFF = 1 + (m − 1) · ICC effective n = n / DEFF
people per site ICC DEFF effective n
-----------------------------------------------------
30 0.01 1.29 n / 1.29
30 0.05 2.45 n / 2.45
30 0.10 3.90 n / 3.90
120 0.05 6.95 n / 6.95
Worked example. The 63-per-arm figure from Brief 2, with thirty people per site and an ICC of 0.05, becomes 154 per arm. Two and a half times the sample, from one correlation nobody measured.
The useful consequence. Many small clusters beat few large ones. Eleven branches of thirty beat three branches of a hundred and ten, by a wide margin, for the same number of people.
Why it matters. It is the difference between a trial and a demonstration. Both are legitimate; presenting the second as the first is not.
The idea. Write down the analysis before you see the data, date it, and give it to someone who is not you.
One page. The primary outcome — exactly one. The secondary outcomes, listed in advance and labelled secondary. The analysis period. The exclusion rules. And the line that earns the room: what result would count as a failure.
Why it works. Simmons, Nelson and Simonsohn (2011) showed how few free choices it takes — which outcome, which period, which exclusion — to produce a significant result from data with no effect in it at all. The choices are made in good faith. That is precisely why removing them has to be structural rather than moral.
Worked example. Two analysts, same data, same honesty. One chose the outcome after seeing which moved; the other chose it in week three and signed it. Only the second can answer the sceptic's question — did you pick this because it worked? — with a document instead of an assurance.
What it costs. Forty minutes, once, before deployment. There is no cheaper credibility available anywhere in this book.
Why it matters. It changes what the results meeting is about. With a pre-registration, the meeting is about the result. Without one, it is about the analyst.
The idea. A measured saving and a banked saving are different things. The ratio between them has a name and it should be measured, never assumed.
φ = banked saving / measured saving
measured £96,000/yr φ = 1.00 -> banked £96,000
φ = 0.80 -> banked £76,800
φ = 0.65 -> banked £62,400
φ = 0.40 -> banked £38,400
The test for whether a saving is banked. Name the budget line it falls out of, and the person whose budget falls. If you cannot name both, φ = 0 for that saving, regardless of how well it was measured.
Worked example. A team saves 900 hours of rework a year. Real, verified, uncontested. Nobody's budget falls, because the team simply does other work with the hours. That is a genuine improvement and a φ of zero, and it must be presented as capacity released rather than cost avoided — which is a different claim, made to a different person, for a different purpose.
Why it matters. Every proposal in this field is built on savings. Almost none of them state φ. A proposal that states a measured φ from a completed pilot is believed about everything else on the page.
The idea. Savings are continuous when you measure them and lumpy when you claim them. Below one whole step, nothing falls.
Costs come in units: a person, a shift, a vehicle, a lease, a contracted volume, a licence band. You can measure a 12 percent reduction. You cannot bank 12 percent of a person.
claimable = step × floor(saving / step)
saving (one shift-year = 1,800 h) claimable φ
-----------------------------------------------------------
12 % = 216 h 0 h 0.00
42 % = 756 h 0 h 0.00
55 % = 990 h 0 h 0.00
110 % = 1,980 h 1,800 h 0.91
230 % = 4,140 h 3,600 h 0.87
Worked example. A picking improvement releases 990 hours a year. It is real. It banks nothing, because a rota does not release half a person. The same improvement at twice the scale banks 1,800 hours and φ jumps from zero to 0.91 — a step, not a slope.
The design rule that follows. Size the pilot so the saving crosses at least one whole claimable step, and name the step in the proposal. If no achievable pilot crosses one: aggregate several savings that fall on the same step; label the work a feasibility study with no financial claim; or choose a saving that lands on a continuous line — energy, consumables, freight — where every unit counts.
Why it matters. This is how an excellent pilot dies. Not disputed. Absorbed.
The idea. Point Chapter I.01's scarcity ratio at the pilot programme itself. The pilot budget is a stock; the banked savings are its regeneration.
r = φ · S / C doubling time = ln 2 / ln(1 + r)
φ S C r doubles in
------------------------------------------------------
0.65 96,000 120,000 0.52/yr 1.66 years
0.65 96,000 240,000 0.26/yr 3.00 years
0.40 96,000 120,000 0.32/yr 2.50 years
0.80 150,000 120,000 1.00/yr 1.00 year
Worked example. At r = 0.52 the fund doubles in twenty months with no new allocation. The successor a first pilot can genuinely fund in one cycle is 0.52 times its size — not three times. Ask for 3× and you will be refused and the ladder stops. Ask for 0.52× four times and you are larger in three years than the refused proposal would have made you in five.
The portfolio inequality that underwrites it. A single pilot is a bet; only a portfolio is a rate of return.
p · φ · S / (C + V + A) > WACC
0.60 × 0.65 × 96,000 / (120,000 + 9,000 + 6,000) = 27.7 % vs 9 %
break-even success rate: p = 19.5 %
Fewer than one pilot in five needs to work. That sentence converts an act of faith into an underwriting question.
Why it matters. Because the ladder is the only mechanism by which small, honest, well-measured work reaches a scale that changes anything — and it is made of arithmetic rather than permission.
All figures in these briefs are computed in lib/verify/I_05.py and printed by python3 lib/verify.py I.05. Sources are in the chapter's Works Cited.