Haute Lumière
Commerce · I.05 · MMXXVI · daylight
Volume I — Transition: From Here to the Living Economy Nine movements, one ladder.
Chapter I.01 gave you a rule for sizing a pilot and asked you to memorise it: the effect must be greater than three times the period-to-period noise, and the cost must fit inside one persuaded person's discretion. That rule has got a great many people started, which is what a rule of thumb is for.
This chapter is the rule done properly, because you are about to spend real money and stake a reputation on the answer, and at that point a rule of thumb is no longer a kindness.
Here is what we are going to do. We will turn the noise rule into an actual power calculation, which takes about four lines and one constant. We will name the four confounders that will otherwise eat your result — seasonality, the site you chose because it was struggling, the correlation between one month and the next, and the enthusiasm of the person running it. We will build the measurement so that it survives the meeting where somebody who does not want it to be true examines it line by line, because that meeting is the only audience worth designing for. And then we will do the part almost nobody does: the accounting that decides whether the saving you measured can actually be claimed, or whether it is absorbed into the ordinary working of the business and disappears.
That last piece is the difference between a pilot that is admired and a pilot that pays — and a pilot that pays funds the next one, which is the only mechanism by which any of this reaches a scale that matters.
You are not trying to prove that regeneration works. Somebody has already proven that, in your own accounts, and I.02 showed you where to find it. You are trying to produce one result that a hostile reader cannot dismantle, at a price that one signature can carry, in a form that puts money back into the budget it came from.
— The Editors
People have been designing measurements that survive scepticism for a very long time, and they have left the instructions. The pilot you are about to build has excellent ancestors, and each one solved a problem you are about to have.
Progresa, Mexico, 1997. Santiago Levy designed a conditional cash transfer programme for rural households, and he designed something else at the same time: the order of the rollout. Five hundred and six communities were assigned to receive the programme immediately or eighteen months later, and the assignment was randomised. Nobody was refused anything — everybody was getting it — but the eighteen-month lag produced a control group, and the control group produced evidence. When the government changed in 2000, the programme survived under a new name, because the evidence had been built into it rather than commissioned afterwards. Levy has written openly that this was the intention: the evaluation was a survival strategy. It is the single most useful precedent in this chapter. Randomise the order, not the access.
The Oregon Health Insurance Experiment, 2008. Oregon had funds to extend Medicaid to about ten thousand people and roughly ninety thousand applicants. Faced with an allocation problem, the state ran a lottery — which, entirely incidentally, created the first randomised evidence on the effects of health insurance coverage in a developed country. Finkelstein and colleagues published the first-year results in 2012, Baicker and colleagues the two-year clinical outcomes in 2013, and the findings were mixed in exactly the way honest findings are mixed: large effects on financial strain and self-reported health, no detectable effect on measured blood pressure, cholesterol or glycated haemoglobin over two years. The scarcity was the experiment. When there is less of a thing than there are people who want it, you already have the structure of a trial; the only decision is whether to waste it.
Interface, and the loop that funded itself. Ray Anderson's waste-elimination programme began in 1995 as a set of small, measured plant-level interventions. What made it compound was not the ambition — it was that each verified saving was returned to the programme rather than absorbed into general overhead, so the second project was funded by the first. Interface's own reporting through the following decade attributes several hundred million dollars of cumulative avoided cost to that programme. The mechanism was a ledger, not a vision.
The Empire State Building, 2009. A deep retrofit was designed jointly by the building's owners, Rocky Mountain Institute, Johnson Controls and Jones Lang LaSalle, and the interesting choice was made before any work started: the measurement method, the baseline and the verification schedule were written into the contract, and the models were published. The projected reduction — around 38 percent of energy use, on the order of four million dollars a year — could therefore be argued with in public, by name, against a stated method. A number that can be argued with is a number that can be believed. Publishing the method before the result is what converts a claim into evidence.
Four cases, four continents of practice, one pattern: in every one, the comparison was designed before the intervention, and the design is what made the result unarguable afterwards. Progresa randomised the order. Oregon randomised the queue. Interface ring-fenced the saving. The Empire State Building pre-published the method. None of them is expensive. All of them are decisions made in week one that cannot be made in week twelve.
So the discovery question for your own organisation is not what should we pilot. You already know. It is this: where do you already have a natural comparison you have never used?
First, the rule of thumb, made honest.
I.01 said: size the pilot so the effect exceeds three times the period-to-period standard deviation. Let us see exactly what that buys. Take one period before and one period after, each with standard deviation σ. The standard error of the difference is σ√(1/m + 1/k) where m is the number of pre-periods and k the number of post-periods — with one each, that is σ√2 = 1.414σ. A three-sigma effect is therefore 2.12 standard errors, which clears the conventional 5 percent threshold, and gives you:
m pre k post SE (sd) z power
---------------------------------------------
1 1 1.4142 2.12 56.4 %
1 3 1.1547 2.60 73.8 %
3 3 0.8165 3.67 95.7 %
6 3 0.7071 4.24 98.9 %
12 3 0.6455 4.65 99.6 %
At one period each side, the three-sigma rule gives you a 56 percent chance of detecting an effect that is genuinely there. That is a coin flip wearing a decimal point, and it is how a real improvement gets written off as noise.
Turn it round and ask for the threshold instead of the power. To detect an effect at 5 percent significance with 80 percent power, you need:
detectable effect = (z_α/2 + z_β) · σ · √(1/m + 1/k)
= 2.8016 · σ · √(1/m + 1/k)
m pre k post detect at
-------------------------------
1 1 3.96 σ
1 3 3.23 σ
3 3 2.29 σ
6 3 1.98 σ
12 3 1.81 σ
12 12 1.14 σ
Here is the Balenciaga cut, and it is worth the whole movement: the threshold is not a property of your intervention. It is a property of how much history you bothered to load. Going from one pre-period to twelve moves the bar from 3.96σ to 1.81σ — a factor of 2.19 on the effect you need, which is 4.80 times the statistical power, equivalent to running a pilot almost five times the size. And it costs nothing, because those twelve months already happened. They are in a system, on a disk, maintained by somebody who will send them to you this afternoon.
The cheapest power on sale is the past, and almost nobody buys it.
For the two-arm case — sites getting the change against sites that are not — the sample size is one line:
n per arm = 2 (z_α/2 + z_β)² · (σ / δ)²
= 15.70 · (σ / δ)²
σ / δ n per arm
----------------------
1.0 16
1.5 36
2.0 63
3.0 142
If your monthly variation is 12 percent and you want to detect a 6 percent improvement, σ/δ = 2, and you need 63 sites, teams or weeks per arm. That number stops a great many conversations, which is exactly what it is for. It also tells you which lever to pull: halving σ by measuring better is worth the same as quadrupling the sample, and measuring better is usually cheaper.
Second, the four confounders, each with its arithmetic.
Regression to the mean. You will be tempted to run the pilot at the worst- performing site, because the upside looks largest and the site will welcome you. This is the most expensive instinct in the chapter. If site performance has reliability ρ — the share of the measured variation that is real rather than noise — a site chosen at the 10th percentile is expected to improve by (1 − ρ) × 1.28σ with no intervention at all:
percentile chosen reliability free 'improvement'
------------------------------------------------------
10th 0.6 0.51 σ
10th 0.8 0.26 σ
5th 0.6 0.66 σ
25th 0.6 0.27 σ
Half a standard deviation of counterfeit success, delivered by arithmetic. Galton described the mechanism in 1886 and it has not changed. Choose your site at random, or choose it on a criterion measured in a period you will not use in the comparison.
Serial correlation. Monthly figures are not independent draws. A good month follows a good month. Under an AR(1) process with correlation ρ, the variance of a k-period mean is inflated:
k periods ρ inflation worth this many independent periods
---------------------------------------------------------------------
12 0.0 1.00 12.0
12 0.3 1.76 6.8
12 0.5 2.67 4.5
12 0.7 4.39 2.7
Twelve correlated months at ρ = 0.5 are worth four and a half. Bertrand, Duflo and Mullainathan demonstrated in 2004 that ignoring this produces false positives at several times the nominal rate in published work. Estimate ρ from your own history — it takes one line in a spreadsheet — and inflate σ accordingly before you size anything.
Clustering. You will randomise sites, not people, so your sample is the number of sites. The design effect is 1 + (m − 1)·ICC, where m is people per site:
people/site ICC design effect effective n
----------------------------------------------------
30 0.01 1.29 n / 1.29
30 0.05 2.45 n / 2.45
30 0.10 3.90 n / 3.90
120 0.05 6.95 n / 6.95
With thirty people per site and an intra-cluster correlation of 0.05, the 63 per arm above becomes 154 per arm. That is the number that decides whether you are running a trial or a demonstration, and both are legitimate — but they are not the same document and must never be presented as though they were.
The enthusiast. The person running the pilot wants it to work. This is not dishonesty; it is the condition of getting anything done. It becomes a measurement problem the moment the analysis is chosen after the result is seen. Simmons, Nelson and Simonsohn showed in 2011 how few free choices it takes — which outcome, which period, which exclusion — to manufacture significance from nothing. The fix is free and it is in the Design movement: write down the analysis before you deploy, and date it.
Third, and this is the honest negative: the pilot that succeeds and still does not scale.
Suppose everything above goes well. The design was clean, the effect was real, the number cleared. The pilot can still fail to produce a single pound that anybody can spend, for a reason that has nothing to do with the measurement.
A saving is only banked when somebody's budget line falls. Define the realisation rate:
φ = banked saving / measured saving
And the mechanism that drives φ to zero is a step function. Savings are continuous at pilot scale and lumpy at operating scale. Take a warehouse where one shift-year is 1,800 picking hours:
measured saving hours claimable φ
---------------------------------------------------
12 % of a shift 216 0 0.00
42 % of a shift 756 0 0.00
55 % of a shift 990 0 0.00
110 % of a shift 1,980 1,800 0.91
230 % of a shift 4,140 3,600 0.87
You saved 990 hours. They are real. Nobody's budget fell by a penny, because you do not employ fractions of a shift and the manager will not release a person for half a rota. Below one whole step, φ is zero however good your measurement was. This is the commonest way an excellent pilot dies: not disputed, absorbed.
The consequence is a design rule, and it is the second thing in this chapter worth memorising:
Size the pilot so that its saving crosses at least one whole
claimable step — a person, a vehicle, a shift, a lease, a
contracted volume — and name that step in the proposal.
If no achievable pilot crosses a step, you have three honest options and none of them is to proceed quietly: aggregate several small savings that fall on the same step; run the pilot as an explicitly labelled feasibility study with no financial claim attached; or find a saving that lands on a continuous line — energy, consumables, freight — where every unit counts. Two of those three are better than the pilot you had planned, which is why the constraint is worth meeting early.
There is a second version of the same failure, and Cartwright and Hardie name it precisely: it worked there is not it will work here. John List calls the collapse a voltage drop. A single site run by the person who invented the method, with attention nobody else will receive, tells you the effect is possible — not that it is typical. The cure is in the design: more sites, chosen at random, run by ordinary staff, before you claim a rate.
In the organisation that has learned this, the pilot is not an event. It is a standing capability, and it is boring in the way that good infrastructure is boring.
There is a baseline register: one page per metric, each with a definition precise enough that two people compute it identically, a period, a method, a named verifier and two signatures. It is maintained the way a fixed-asset register is maintained, and for the same reason. When someone proposes a change, the baseline is already there. The eight weeks that used to be spent establishing what was true before are simply not spent.
The estate is understood as a natural experiment. Nobody says let us find a control group; the control group is the other eleven branches, and the query that produces it is saved. Rollouts are staggered because rollouts are always staggered, and the order is drawn at random because drawing it at random costs nothing and buys evidence. Nobody is denied anything. Everybody gets it; the sequence is what is randomised.
Analyses are pre-registered — one page, dated, lodged with finance before deployment, saying which outcome, which period, which exclusions. It takes forty minutes. Its effect is that the meeting where the result is presented is about the result, and not about the analyst.
And the money moves. Each verified saving is tagged with the budget line it falls out of and the name of the person whose budget falls. The finance function does not resist this, because it makes the forecast better: a saving with an owner is a saving that can be planned against. The share that returns to the pilot fund is written into the scheme, so the fund replenishes without an allocation, and the question at the capital committee is not may we have money for pilots but the fund is returning 27 percent, how much more of it would you like.
The people running it can tell you the fund's regeneration rate to one decimal place, and its doubling time in years. It is the same arithmetic as Chapter I.01's S = D / (R · r), pointed at the programme itself: the pilot budget is a stock, and a well-run pilot programme is a stock with a positive regeneration rate. Nobody finds this unusual. It is simply how the pilots are funded.
Twelve weeks, four decisions, one page each. Everything here is done before the intervention starts, and that is the entire design.
Week 1–2 — The comparison. Choose, in this order of preference: randomised staggered rollout across your own sites; matched controls from your own estate; long pre-period interrupted time series. Never a single site with one month before and one month after. If the only available design is that one, say so in writing and label the output a feasibility study.
Week 2–3 — The power calculation. Compute σ from history, inflate it for serial correlation, apply the design effect for clustering, and state the minimum detectable effect you can afford. If the MDE is larger than the effect you believe in, you have learned the most valuable thing available this quarter, and you have learned it before spending anything. Change the design, extend the history, add sites, or measure something with less noise in it.
Week 3–4 — The pre-registration. One page, dated, two signatures, lodged with finance and the verifier. It names: the primary outcome, exactly one; the secondary outcomes, listed in advance and labelled secondary; the analysis period; the exclusion rules; and what result would count as a failure. That last line is the one that earns you the room. An analyst who has written down in advance what would change their mind is believed about everything else.
Week 4 — The claim map. For each pound of expected saving, write the budget line it falls out of, the step it must cross to fall, and the name of the person whose budget falls. Get that person's agreement in writing before deployment. This is not bureaucracy; it is the difference between φ = 0.65 and φ = 0, and it is discovered on day thirty or never.
Weeks 5–12 — Deploy, measure, and do not touch the analysis. Log the intervention date, the co-interventions — there will be co-interventions — and anything unusual. Deming's lesson applies precisely here: reacting to common-cause variation as though it were a signal makes the system worse. Hold still.
Governance, in three roles that must not be held by two people.
Where a standard is wanted, IPMVP gives you the vocabulary — Option C for whole-facility comparison, Option B where you can meter the affected system directly — and a named standard in a proposal removes an entire category of objection.
The pilot programme sustains itself if, and only if, the money comes back. That is the whole of it, and it is a mechanical condition rather than a cultural one.
Treat the pilot fund as a stock with a regeneration rate:
r = φ · S / C doubling time = ln 2 / ln(1 + r)
φ S C r doubles in
-----------------------------------------------------
0.65 96,000 120,000 0.52/yr 1.66 years
0.65 96,000 240,000 0.26/yr 3.00 years
0.40 96,000 120,000 0.32/yr 2.50 years
0.80 150,000 120,000 1.00/yr 1.00 year
At r = 0.52 the fund doubles in twenty months with no new allocation. That is the ladder, and it is why the successor must be sized to the saving rather than to the ambition: a first pilot at these numbers funds a successor 0.52 times its size in one cycle, not three times. Ask for 3× and you will be refused and the ladder stops. Ask for 0.52× and take it four times, and you are larger in three years than the refused proposal would have made you in five.
Three failure modes, named honestly.
The fund is raided. A good year ends, a gap appears elsewhere, and the replenishment is swept into it. Prevent it the way any ring-fence is prevented: a written reversion clause, a named owner, and a line in the standing pack so that removing it requires an explanation.
The baseline register rots. Definitions drift, metrics are redefined by a system upgrade, and two years of comparison quietly become incomparable. A register with no owner is a register that is already wrong. Review it annually, and version every definition.
The verifier becomes the advocate. This is the slowest and most dangerous, because it looks like success. The verifier who has approved eleven results has a stake in the twelfth. Rotate the role, or have the verifier's own work read by somebody who did not produce it — the estate's own practice, applied to the instrument that checks the estate.
There is a specific and underrated pleasure in the meeting where the sceptic's first objection is already answered on page one, in writing, dated before anybody knew the result. You do not have to defend anything. You simply point. The room reorganises itself around the fact that the work was done properly, and the conversation moves on to what to do next — which is the conversation you wanted all along.
Then there is the quieter pleasure, the one that arrives months later: the fund replenishes and nobody has to ask. Money that went out comes back with more behind it, and the second pilot is financed by the first without a single meeting. A ladder you built has a rung on it that you did not put there.
And there is the smallest one, which is the best. Somebody in another division, whom you have never met, asks the finance team for the baseline template. It has stopped being your method. That is the point at which it becomes the way things are done here, and it happens without a mandate, an announcement or a programme — which is the only way it ever happens.
The instrument: an evergreen pilot facility with a claim ledger and a covenanted reversion.
Not a budget line. A revolving internal fund, underwritten on a portfolio return rather than on any single pilot, with a written claim on a fixed share of the savings it produces.
The mechanics.
p well below one. A fund that pretends every pilot works is a fund that will be embarrassed once and closed.The balance-sheet treatment. Pilot spend is ordinarily expensed, and should be — a pilot is an information-acquisition cost and it is genuinely period cost. Where the intervention leaves behind a long-lived improvement, capitalise that portion and depreciate it over the asset's regenerated life, as I.01 argued. The information itself has value and no line, which is precisely why the claim ledger exists: it is the asset register for knowledge the business has bought and would otherwise lose at the next reorganisation.
The counterparty. Treasury to the business unit, internally, first. Once the fund has completed two cycles with a verified portfolio return, the same structure is financeable externally — this is recognisably an energy-performance contract with a wider outcome definition, and lenders already understand the shape. Do not go outside before you have two cycles. The track record is the security.
The number that decides it. One inequality, on the front page, and it is a portfolio number because a single pilot is a bet and only a portfolio is a rate:
p · φ · S
----------------------- > WACC
C + V + A
p share of pilots producing a verified saving
φ realisation rate — the share that falls out of a budget line
S verified annual saving per successful pilot
C pilot cost V verification A administration
Worked at the figures above — p = 0.60, φ = 0.65, S = £96,000, C = £120,000, V = £9,000, A = £6,000:
0.60 × 0.65 × 96,000 = 37,440
120,000 + 9,000 + 6,000 = 135,000
37,440 / 135,000 = 27.7 % vs a 9 % WACC
-> clears by 18.7 points
break-even success rate at this WACC: p = 19.5 %
Fewer than one pilot in five needs to work. Put that sentence in the paper. It converts the proposal from an act of faith into an underwriting question, and underwriting questions get answered by people who never answer questions of faith.
The first ninety days on a page.
| Day | Action | Artifact |
|---|---|---|
| 1–10 | Find the natural comparison; pull thirty-six months of history | The comparison memo |
| 11–20 | Compute σ, inflate for ρ and ICC, state the MDE | The power calculation |
| 21–30 | Write the claim map; get the budget owner's signature | The claim map |
| 31–40 | Pre-register: outcome, period, exclusions, failure condition | The pre-registration |
| 41–45 | Name the verifier; agree the method reference | Verification memo |
| 46–75 | Deploy. Log co-interventions. Change nothing in the analysis | Intervention log |
| 76–85 | Verify against the pre-registration | The verified result |
| 86–90 | Bank it: the ledger row, the reversion, the successor at φS/C | The claim ledger |
Discovery — what is already working
Dream — what becomes possible
Design — what we build
Destiny — how it holds
Levy, S. (2006). Progress Against Poverty: Sustaining Mexico's Progresa-Oportunidades Program. Brookings Institution Press.
Skoufias, E. (2005). PROGRESA and Its Impacts on the Welfare of Rural Households in Mexico. IFPRI Research Report 139. International Food Policy Research Institute.
Finkelstein, A., Taubman, S., Wright, B., Bernstein, M., Gruber, J., Newhouse, J. P., Allen, H., Baicker, K. and the Oregon Health Study Group (2012). "The Oregon Health Insurance Experiment: Evidence from the First Year." Quarterly Journal of Economics, 127(3), 1057–1106.
Baicker, K., Taubman, S. L., Allen, H. L., Bernstein, M., Gruber, J. H., Newhouse, J. P., Schneider, E. C., Wright, B. J., Zaslavsky, A. M. and Finkelstein, A. N. (2013). "The Oregon Experiment — Effects of Medicaid on Clinical Outcomes." New England Journal of Medicine, 368(18), 1713–1722.
Anderson, R. C. and White, R. (2009). Confessions of a Radical Industrialist. St. Martin's Press.
Rocky Mountain Institute (2009). Empire State Building Retrofit: Cost-Effective Greenhouse Gas Reductions via Whole-Building Retrofits — Process, Findings and Tools. RMI, with Johnson Controls and Jones Lang LaSalle.
Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences, 2nd edn. Lawrence Erlbaum.
Bloom, H. S. (1995). "Minimum Detectable Effects: A Simple Way to Report the Statistical Power of Experimental Designs." Evaluation Review, 19(5), 547–556.
Duflo, E., Glennerster, R. and Kremer, M. (2007). "Using Randomization in Development Economics Research: A Toolkit." In Handbook of Development Economics, Vol. 4. Elsevier.
Gertler, P. J., Martinez, S., Premand, P., Rawlings, L. B. and Vermeersch, C. M. J. (2016). Impact Evaluation in Practice, 2nd edn. World Bank.
Hussey, M. A. and Hughes, J. P. (2007). "Design and Analysis of Stepped Wedge Cluster Randomized Trials." Contemporary Clinical Trials, 28(2), 182–191.
Bertrand, M., Duflo, E. and Mullainathan, S. (2004). "How Much Should We Trust Differences-in-Differences Estimates?" Quarterly Journal of Economics, 119(1), 249–275.
Galton, F. (1886). "Regression Towards Mediocrity in Hereditary Stature." Journal of the Anthropological Institute of Great Britain and Ireland, 15, 246–263.
Campbell, D. T. and Stanley, J. C. (1963). Experimental and Quasi-Experimental Designs for Research. Rand McNally.
Simmons, J. P., Nelson, L. D. and Simonsohn, U. (2011). "False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant." Psychological Science, 22(11), 1359–1366.
Ioannidis, J. P. A. (2005). "Why Most Published Research Findings Are False." PLoS Medicine, 2(8), e124.
Cartwright, N. and Hardie, J. (2012). Evidence-Based Policy: A Practical Guide to Doing It Better. Oxford University Press.
List, J. A. (2022). The Voltage Effect: How to Make Good Ideas Great and Great Ideas Scale. Currency.
Deming, W. E. (1986). Out of the Crisis. MIT Center for Advanced Engineering Study.
Shewhart, W. A. (1931). Economic Control of Quality of Manufactured Product. D. Van Nostrand.
Efficiency Valuation Organization. International Performance Measurement and Verification Protocol (IPMVP), Core Concepts. Successive editions.
Ostrom, E. (1990). Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press.
Cooperrider, D. L. and Whitney, D. (2005). Appreciative Inquiry: A Positive Revolution in Change. Berrett-Koehler.
Note on figures. Every constant in this chapter — the critical values, the sample-size constant K = 15.70, the power table, the detection thresholds, the regression-to-the-mean drift, the AR(1) inflation, the design effects, the step function, the fund's regeneration rate and the portfolio inequality — is computed in lib/verify/I_05.py and printed by python3 lib/verify.py I.05. The z values are computed at run time by bisection on the normal cumulative distribution rather than quoted, so any reader with a statistical table can check them. Where this chapter tightens I.01's three-sigma rule, the tightened figure is given and the difference stated in the text.