Haute Lumière

Commerce · I.05 · MMXXVI · daylight

La Bourse  /  Volume I  /  Nº I.05

The Pilot That Pays

Volume I — Transition: From Here to the Living Economy Nine movements, one ladder.


THE PLATE

A woman standing alone in an empty room before a wide window of trees, sepia light running across the floor toward her.
Plate I.05Two Rooms, One Question.A pilot is not a small version of the plan. It is a comparison you built on purpose, in a place where the only difference is the one you made.

THE LETTER

Chapter I.01 gave you a rule for sizing a pilot and asked you to memorise it: the effect must be greater than three times the period-to-period noise, and the cost must fit inside one persuaded person's discretion. That rule has got a great many people started, which is what a rule of thumb is for.

This chapter is the rule done properly, because you are about to spend real money and stake a reputation on the answer, and at that point a rule of thumb is no longer a kindness.

Here is what we are going to do. We will turn the noise rule into an actual power calculation, which takes about four lines and one constant. We will name the four confounders that will otherwise eat your result — seasonality, the site you chose because it was struggling, the correlation between one month and the next, and the enthusiasm of the person running it. We will build the measurement so that it survives the meeting where somebody who does not want it to be true examines it line by line, because that meeting is the only audience worth designing for. And then we will do the part almost nobody does: the accounting that decides whether the saving you measured can actually be claimed, or whether it is absorbed into the ordinary working of the business and disappears.

That last piece is the difference between a pilot that is admired and a pilot that pays — and a pilot that pays funds the next one, which is the only mechanism by which any of this reaches a scale that matters.

You are not trying to prove that regeneration works. Somebody has already proven that, in your own accounts, and I.02 showed you where to find it. You are trying to produce one result that a hostile reader cannot dismantle, at a price that one signature can carry, in a form that puts money back into the budget it came from.

— The Editors


DISCOVERY

What is already working

People have been designing measurements that survive scepticism for a very long time, and they have left the instructions. The pilot you are about to build has excellent ancestors, and each one solved a problem you are about to have.

Progresa, Mexico, 1997. Santiago Levy designed a conditional cash transfer programme for rural households, and he designed something else at the same time: the order of the rollout. Five hundred and six communities were assigned to receive the programme immediately or eighteen months later, and the assignment was randomised. Nobody was refused anything — everybody was getting it — but the eighteen-month lag produced a control group, and the control group produced evidence. When the government changed in 2000, the programme survived under a new name, because the evidence had been built into it rather than commissioned afterwards. Levy has written openly that this was the intention: the evaluation was a survival strategy. It is the single most useful precedent in this chapter. Randomise the order, not the access.

The Oregon Health Insurance Experiment, 2008. Oregon had funds to extend Medicaid to about ten thousand people and roughly ninety thousand applicants. Faced with an allocation problem, the state ran a lottery — which, entirely incidentally, created the first randomised evidence on the effects of health insurance coverage in a developed country. Finkelstein and colleagues published the first-year results in 2012, Baicker and colleagues the two-year clinical outcomes in 2013, and the findings were mixed in exactly the way honest findings are mixed: large effects on financial strain and self-reported health, no detectable effect on measured blood pressure, cholesterol or glycated haemoglobin over two years. The scarcity was the experiment. When there is less of a thing than there are people who want it, you already have the structure of a trial; the only decision is whether to waste it.

Interface, and the loop that funded itself. Ray Anderson's waste-elimination programme began in 1995 as a set of small, measured plant-level interventions. What made it compound was not the ambition — it was that each verified saving was returned to the programme rather than absorbed into general overhead, so the second project was funded by the first. Interface's own reporting through the following decade attributes several hundred million dollars of cumulative avoided cost to that programme. The mechanism was a ledger, not a vision.

The Empire State Building, 2009. A deep retrofit was designed jointly by the building's owners, Rocky Mountain Institute, Johnson Controls and Jones Lang LaSalle, and the interesting choice was made before any work started: the measurement method, the baseline and the verification schedule were written into the contract, and the models were published. The projected reduction — around 38 percent of energy use, on the order of four million dollars a year — could therefore be argued with in public, by name, against a stated method. A number that can be argued with is a number that can be believed. Publishing the method before the result is what converts a claim into evidence.

Four cases, four continents of practice, one pattern: in every one, the comparison was designed before the intervention, and the design is what made the result unarguable afterwards. Progresa randomised the order. Oregon randomised the queue. Interface ring-fenced the saving. The Empire State Building pre-published the method. None of them is expensive. All of them are decisions made in week one that cannot be made in week twelve.

So the discovery question for your own organisation is not what should we pilot. You already know. It is this: where do you already have a natural comparison you have never used?

  1. Multiple sites doing the same work. Five depots, eleven branches, four plants, nine wards. Four of them are a control group that costs a query.
  2. A queue. Anything with a waiting list is a lottery you have not run.
  3. A staggered rollout already scheduled. If the new system reaches sites in some order, that order can be randomised at no cost, this week.
  4. History. Thirty-six months of monthly data on the thing you are about to change is sitting in a system somebody maintains. It is the cheapest statistical power available to anyone, and we will do that arithmetic in a moment.
  5. A near miss. A site that was going to get the change and did not, for a reason unrelated to the change itself. That is a natural experiment, and it is often better than one you could have designed.

THE ARITHMETIC

The power calculation, the confounders, and the saving that cannot be claimed

First, the rule of thumb, made honest.

I.01 said: size the pilot so the effect exceeds three times the period-to-period standard deviation. Let us see exactly what that buys. Take one period before and one period after, each with standard deviation σ. The standard error of the difference is σ√(1/m + 1/k) where m is the number of pre-periods and k the number of post-periods — with one each, that is σ√2 = 1.414σ. A three-sigma effect is therefore 2.12 standard errors, which clears the conventional 5 percent threshold, and gives you:

  m pre   k post    SE (sd)     z       power
  ---------------------------------------------
    1       1       1.4142     2.12     56.4 %
    1       3       1.1547     2.60     73.8 %
    3       3       0.8165     3.67     95.7 %
    6       3       0.7071     4.24     98.9 %
   12       3       0.6455     4.65     99.6 %

At one period each side, the three-sigma rule gives you a 56 percent chance of detecting an effect that is genuinely there. That is a coin flip wearing a decimal point, and it is how a real improvement gets written off as noise.

Turn it round and ask for the threshold instead of the power. To detect an effect at 5 percent significance with 80 percent power, you need:

  detectable effect  =  (z_α/2 + z_β) · σ · √(1/m + 1/k)
                     =  2.8016 · σ · √(1/m + 1/k)

   m pre   k post    detect at
   -------------------------------
     1       1        3.96 σ
     1       3        3.23 σ
     3       3        2.29 σ
     6       3        1.98 σ
    12       3        1.81 σ
    12      12        1.14 σ

Here is the Balenciaga cut, and it is worth the whole movement: the threshold is not a property of your intervention. It is a property of how much history you bothered to load. Going from one pre-period to twelve moves the bar from 3.96σ to 1.81σ — a factor of 2.19 on the effect you need, which is 4.80 times the statistical power, equivalent to running a pilot almost five times the size. And it costs nothing, because those twelve months already happened. They are in a system, on a disk, maintained by somebody who will send them to you this afternoon.

The cheapest power on sale is the past, and almost nobody buys it.

For the two-arm case — sites getting the change against sites that are not — the sample size is one line:

  n per arm  =  2 (z_α/2 + z_β)² · (σ / δ)²
             =  15.70 · (σ / δ)²

   σ / δ      n per arm
   ----------------------
     1.0          16
     1.5          36
     2.0          63
     3.0         142

If your monthly variation is 12 percent and you want to detect a 6 percent improvement, σ/δ = 2, and you need 63 sites, teams or weeks per arm. That number stops a great many conversations, which is exactly what it is for. It also tells you which lever to pull: halving σ by measuring better is worth the same as quadrupling the sample, and measuring better is usually cheaper.

Second, the four confounders, each with its arithmetic.

Regression to the mean. You will be tempted to run the pilot at the worst- performing site, because the upside looks largest and the site will welcome you. This is the most expensive instinct in the chapter. If site performance has reliability ρ — the share of the measured variation that is real rather than noise — a site chosen at the 10th percentile is expected to improve by (1 − ρ) × 1.28σ with no intervention at all:

   percentile chosen   reliability   free 'improvement'
   ------------------------------------------------------
        10th              0.6            0.51 σ
        10th              0.8            0.26 σ
         5th              0.6            0.66 σ
        25th              0.6            0.27 σ

Half a standard deviation of counterfeit success, delivered by arithmetic. Galton described the mechanism in 1886 and it has not changed. Choose your site at random, or choose it on a criterion measured in a period you will not use in the comparison.

Serial correlation. Monthly figures are not independent draws. A good month follows a good month. Under an AR(1) process with correlation ρ, the variance of a k-period mean is inflated:

   k periods   ρ      inflation    worth this many independent periods
   ---------------------------------------------------------------------
      12      0.0       1.00                 12.0
      12      0.3       1.76                  6.8
      12      0.5       2.67                  4.5
      12      0.7       4.39                  2.7

Twelve correlated months at ρ = 0.5 are worth four and a half. Bertrand, Duflo and Mullainathan demonstrated in 2004 that ignoring this produces false positives at several times the nominal rate in published work. Estimate ρ from your own history — it takes one line in a spreadsheet — and inflate σ accordingly before you size anything.

Clustering. You will randomise sites, not people, so your sample is the number of sites. The design effect is 1 + (m − 1)·ICC, where m is people per site:

   people/site   ICC     design effect   effective n
   ----------------------------------------------------
        30       0.01        1.29          n / 1.29
        30       0.05        2.45          n / 2.45
        30       0.10        3.90          n / 3.90
       120       0.05        6.95          n / 6.95

With thirty people per site and an intra-cluster correlation of 0.05, the 63 per arm above becomes 154 per arm. That is the number that decides whether you are running a trial or a demonstration, and both are legitimate — but they are not the same document and must never be presented as though they were.

The enthusiast. The person running the pilot wants it to work. This is not dishonesty; it is the condition of getting anything done. It becomes a measurement problem the moment the analysis is chosen after the result is seen. Simmons, Nelson and Simonsohn showed in 2011 how few free choices it takes — which outcome, which period, which exclusion — to manufacture significance from nothing. The fix is free and it is in the Design movement: write down the analysis before you deploy, and date it.

Third, and this is the honest negative: the pilot that succeeds and still does not scale.

Suppose everything above goes well. The design was clean, the effect was real, the number cleared. The pilot can still fail to produce a single pound that anybody can spend, for a reason that has nothing to do with the measurement.

A saving is only banked when somebody's budget line falls. Define the realisation rate:

      φ  =  banked saving  /  measured saving

And the mechanism that drives φ to zero is a step function. Savings are continuous at pilot scale and lumpy at operating scale. Take a warehouse where one shift-year is 1,800 picking hours:

   measured saving        hours    claimable    φ
   ---------------------------------------------------
    12 % of a shift         216         0      0.00
    42 % of a shift         756         0      0.00
    55 % of a shift         990         0      0.00
   110 % of a shift       1,980     1,800      0.91
   230 % of a shift       4,140     3,600      0.87

You saved 990 hours. They are real. Nobody's budget fell by a penny, because you do not employ fractions of a shift and the manager will not release a person for half a rota. Below one whole step, φ is zero however good your measurement was. This is the commonest way an excellent pilot dies: not disputed, absorbed.

The consequence is a design rule, and it is the second thing in this chapter worth memorising:

  Size the pilot so that its saving crosses at least one whole
  claimable step — a person, a vehicle, a shift, a lease, a
  contracted volume — and name that step in the proposal.

If no achievable pilot crosses a step, you have three honest options and none of them is to proceed quietly: aggregate several small savings that fall on the same step; run the pilot as an explicitly labelled feasibility study with no financial claim attached; or find a saving that lands on a continuous line — energy, consumables, freight — where every unit counts. Two of those three are better than the pilot you had planned, which is why the constraint is worth meeting early.

There is a second version of the same failure, and Cartwright and Hardie name it precisely: it worked there is not it will work here. John List calls the collapse a voltage drop. A single site run by the person who invented the method, with attention nobody else will receive, tells you the effect is possible — not that it is typical. The cure is in the design: more sites, chosen at random, run by ordinary staff, before you claim a rate.


DREAM

What becomes ordinary

In the organisation that has learned this, the pilot is not an event. It is a standing capability, and it is boring in the way that good infrastructure is boring.

There is a baseline register: one page per metric, each with a definition precise enough that two people compute it identically, a period, a method, a named verifier and two signatures. It is maintained the way a fixed-asset register is maintained, and for the same reason. When someone proposes a change, the baseline is already there. The eight weeks that used to be spent establishing what was true before are simply not spent.

The estate is understood as a natural experiment. Nobody says let us find a control group; the control group is the other eleven branches, and the query that produces it is saved. Rollouts are staggered because rollouts are always staggered, and the order is drawn at random because drawing it at random costs nothing and buys evidence. Nobody is denied anything. Everybody gets it; the sequence is what is randomised.

Analyses are pre-registered — one page, dated, lodged with finance before deployment, saying which outcome, which period, which exclusions. It takes forty minutes. Its effect is that the meeting where the result is presented is about the result, and not about the analyst.

And the money moves. Each verified saving is tagged with the budget line it falls out of and the name of the person whose budget falls. The finance function does not resist this, because it makes the forecast better: a saving with an owner is a saving that can be planned against. The share that returns to the pilot fund is written into the scheme, so the fund replenishes without an allocation, and the question at the capital committee is not may we have money for pilots but the fund is returning 27 percent, how much more of it would you like.

The people running it can tell you the fund's regeneration rate to one decimal place, and its doubling time in years. It is the same arithmetic as Chapter I.01's S = D / (R · r), pointed at the programme itself: the pilot budget is a stock, and a well-run pilot programme is a stock with a positive regeneration rate. Nobody finds this unusual. It is simply how the pilots are funded.


DESIGN

The instrument, built

Twelve weeks, four decisions, one page each. Everything here is done before the intervention starts, and that is the entire design.

Week 1–2 — The comparison. Choose, in this order of preference: randomised staggered rollout across your own sites; matched controls from your own estate; long pre-period interrupted time series. Never a single site with one month before and one month after. If the only available design is that one, say so in writing and label the output a feasibility study.

Week 2–3 — The power calculation. Compute σ from history, inflate it for serial correlation, apply the design effect for clustering, and state the minimum detectable effect you can afford. If the MDE is larger than the effect you believe in, you have learned the most valuable thing available this quarter, and you have learned it before spending anything. Change the design, extend the history, add sites, or measure something with less noise in it.

Week 3–4 — The pre-registration. One page, dated, two signatures, lodged with finance and the verifier. It names: the primary outcome, exactly one; the secondary outcomes, listed in advance and labelled secondary; the analysis period; the exclusion rules; and what result would count as a failure. That last line is the one that earns you the room. An analyst who has written down in advance what would change their mind is believed about everything else.

Week 4 — The claim map. For each pound of expected saving, write the budget line it falls out of, the step it must cross to fall, and the name of the person whose budget falls. Get that person's agreement in writing before deployment. This is not bureaucracy; it is the difference between φ = 0.65 and φ = 0, and it is discovered on day thirty or never.

Weeks 5–12 — Deploy, measure, and do not touch the analysis. Log the intervention date, the co-interventions — there will be co-interventions — and anything unusual. Deming's lesson applies precisely here: reacting to common-cause variation as though it were a signal makes the system worse. Hold still.

Governance, in three roles that must not be held by two people.

Where a standard is wanted, IPMVP gives you the vocabulary — Option C for whole-facility comparison, Option B where you can meter the affected system directly — and a named standard in a proposal removes an entire category of objection.


DESTINY

How it holds when you stop pushing

The pilot programme sustains itself if, and only if, the money comes back. That is the whole of it, and it is a mechanical condition rather than a cultural one.

Treat the pilot fund as a stock with a regeneration rate:

        r  =  φ · S / C        doubling time  =  ln 2 / ln(1 + r)

   φ      S          C          r        doubles in
   -----------------------------------------------------
  0.65   96,000    120,000    0.52/yr     1.66 years
  0.65   96,000    240,000    0.26/yr     3.00 years
  0.40   96,000    120,000    0.32/yr     2.50 years
  0.80  150,000    120,000    1.00/yr     1.00 year

At r = 0.52 the fund doubles in twenty months with no new allocation. That is the ladder, and it is why the successor must be sized to the saving rather than to the ambition: a first pilot at these numbers funds a successor 0.52 times its size in one cycle, not three times. Ask for 3× and you will be refused and the ladder stops. Ask for 0.52× and take it four times, and you are larger in three years than the refused proposal would have made you in five.

Three failure modes, named honestly.

The fund is raided. A good year ends, a gap appears elsewhere, and the replenishment is swept into it. Prevent it the way any ring-fence is prevented: a written reversion clause, a named owner, and a line in the standing pack so that removing it requires an explanation.

The baseline register rots. Definitions drift, metrics are redefined by a system upgrade, and two years of comparison quietly become incomparable. A register with no owner is a register that is already wrong. Review it annually, and version every definition.

The verifier becomes the advocate. This is the slowest and most dangerous, because it looks like success. The verifier who has approved eleven results has a stake in the twelfth. Rotate the role, or have the verifier's own work read by somebody who did not produce it — the estate's own practice, applied to the instrument that checks the estate.


DELIGHT

What it feels like

There is a specific and underrated pleasure in the meeting where the sceptic's first objection is already answered on page one, in writing, dated before anybody knew the result. You do not have to defend anything. You simply point. The room reorganises itself around the fact that the work was done properly, and the conversation moves on to what to do next — which is the conversation you wanted all along.

Then there is the quieter pleasure, the one that arrives months later: the fund replenishes and nobody has to ask. Money that went out comes back with more behind it, and the second pilot is financed by the first without a single meeting. A ladder you built has a rung on it that you did not put there.

And there is the smallest one, which is the best. Somebody in another division, whom you have never met, asks the finance team for the baseline template. It has stopped being your method. That is the point at which it becomes the way things are done here, and it happens without a mandate, an announcement or a programme — which is the only way it ever happens.


OPERATIONALIZE THIS

At the level of finance

The instrument: an evergreen pilot facility with a claim ledger and a covenanted reversion.

Not a budget line. A revolving internal fund, underwritten on a portfolio return rather than on any single pilot, with a written claim on a fixed share of the savings it produces.

The mechanics.

The balance-sheet treatment. Pilot spend is ordinarily expensed, and should be — a pilot is an information-acquisition cost and it is genuinely period cost. Where the intervention leaves behind a long-lived improvement, capitalise that portion and depreciate it over the asset's regenerated life, as I.01 argued. The information itself has value and no line, which is precisely why the claim ledger exists: it is the asset register for knowledge the business has bought and would otherwise lose at the next reorganisation.

The counterparty. Treasury to the business unit, internally, first. Once the fund has completed two cycles with a verified portfolio return, the same structure is financeable externally — this is recognisably an energy-performance contract with a wider outcome definition, and lenders already understand the shape. Do not go outside before you have two cycles. The track record is the security.

The number that decides it. One inequality, on the front page, and it is a portfolio number because a single pilot is a bet and only a portfolio is a rate:

              p · φ · S
        -----------------------   >   WACC
          C  +  V  +  A

  p  share of pilots producing a verified saving
  φ  realisation rate — the share that falls out of a budget line
  S  verified annual saving per successful pilot
  C  pilot cost   V  verification   A  administration

Worked at the figures above — p = 0.60, φ = 0.65, S = £96,000, C = £120,000, V = £9,000, A = £6,000:

   0.60 × 0.65 × 96,000  =  37,440
   120,000 + 9,000 + 6,000  =  135,000
   37,440 / 135,000  =  27.7 %      vs a 9 % WACC
   -> clears by 18.7 points
   break-even success rate at this WACC:  p = 19.5 %

Fewer than one pilot in five needs to work. Put that sentence in the paper. It converts the proposal from an act of faith into an underwriting question, and underwriting questions get answered by people who never answer questions of faith.

The first ninety days on a page.

DayActionArtifact
1–10Find the natural comparison; pull thirty-six months of historyThe comparison memo
11–20Compute σ, inflate for ρ and ICC, state the MDEThe power calculation
21–30Write the claim map; get the budget owner's signatureThe claim map
31–40Pre-register: outcome, period, exclusions, failure conditionThe pre-registration
41–45Name the verifier; agree the method referenceVerification memo
46–75Deploy. Log co-interventions. Change nothing in the analysisIntervention log
76–85Verify against the pre-registrationThe verified result
86–90Bank it: the ledger row, the reversion, the successor at φS/CThe claim ledger

APPRECIATIVE QUESTIONS

Twelve, for a room

Discovery — what is already working

  1. When has this organisation measured something before it changed it, and what did that let us say afterwards that we could not otherwise have said?
  2. Where do we already have a natural comparison — several sites doing the same work, a queue, a rollout order — that we have never used as evidence?
  3. Which of our data series goes back the furthest, and who maintains it? What has their care made possible that they have never been thanked for?

Dream — what becomes possible

  1. If every rollout in this organisation were staggered in a random order as a matter of course, what would we know in three years that we do not know now?
  2. Imagine a baseline register as well kept as our fixed-asset register. What is the first question it would let us answer on the day somebody asks?
  3. If a saving always arrived with the name of the budget it fell out of, what would change about how we plan?

Design — what we build

  1. What is the one result we could produce this quarter that the most sceptical person in the building would accept — and what would it take to design it that way from the start?
  2. Who should be the verifier, and what would make that role genuinely independent rather than nominally independent?
  3. What is the smallest whole step — a shift, a vehicle, a lease, a contract — that one of our candidate savings could cross, and what size of pilot crosses it?

Destiny — how it holds

  1. What share of a verified saving should return to the fund, and what would make the operating unit glad to send it rather than resigned to it?
  2. What would have to be true for the pilot fund still to be revolving when everyone in this room has moved on?
  3. What is the first sign that our verification has quietly become advocacy, and who outside this room would notice it first?

WORKS CITED

Levy, S. (2006). Progress Against Poverty: Sustaining Mexico's Progresa-Oportunidades Program. Brookings Institution Press.

Skoufias, E. (2005). PROGRESA and Its Impacts on the Welfare of Rural Households in Mexico. IFPRI Research Report 139. International Food Policy Research Institute.

Finkelstein, A., Taubman, S., Wright, B., Bernstein, M., Gruber, J., Newhouse, J. P., Allen, H., Baicker, K. and the Oregon Health Study Group (2012). "The Oregon Health Insurance Experiment: Evidence from the First Year." Quarterly Journal of Economics, 127(3), 1057–1106.

Baicker, K., Taubman, S. L., Allen, H. L., Bernstein, M., Gruber, J. H., Newhouse, J. P., Schneider, E. C., Wright, B. J., Zaslavsky, A. M. and Finkelstein, A. N. (2013). "The Oregon Experiment — Effects of Medicaid on Clinical Outcomes." New England Journal of Medicine, 368(18), 1713–1722.

Anderson, R. C. and White, R. (2009). Confessions of a Radical Industrialist. St. Martin's Press.

Rocky Mountain Institute (2009). Empire State Building Retrofit: Cost-Effective Greenhouse Gas Reductions via Whole-Building Retrofits — Process, Findings and Tools. RMI, with Johnson Controls and Jones Lang LaSalle.

Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences, 2nd edn. Lawrence Erlbaum.

Bloom, H. S. (1995). "Minimum Detectable Effects: A Simple Way to Report the Statistical Power of Experimental Designs." Evaluation Review, 19(5), 547–556.

Duflo, E., Glennerster, R. and Kremer, M. (2007). "Using Randomization in Development Economics Research: A Toolkit." In Handbook of Development Economics, Vol. 4. Elsevier.

Gertler, P. J., Martinez, S., Premand, P., Rawlings, L. B. and Vermeersch, C. M. J. (2016). Impact Evaluation in Practice, 2nd edn. World Bank.

Hussey, M. A. and Hughes, J. P. (2007). "Design and Analysis of Stepped Wedge Cluster Randomized Trials." Contemporary Clinical Trials, 28(2), 182–191.

Bertrand, M., Duflo, E. and Mullainathan, S. (2004). "How Much Should We Trust Differences-in-Differences Estimates?" Quarterly Journal of Economics, 119(1), 249–275.

Galton, F. (1886). "Regression Towards Mediocrity in Hereditary Stature." Journal of the Anthropological Institute of Great Britain and Ireland, 15, 246–263.

Campbell, D. T. and Stanley, J. C. (1963). Experimental and Quasi-Experimental Designs for Research. Rand McNally.

Simmons, J. P., Nelson, L. D. and Simonsohn, U. (2011). "False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant." Psychological Science, 22(11), 1359–1366.

Ioannidis, J. P. A. (2005). "Why Most Published Research Findings Are False." PLoS Medicine, 2(8), e124.

Cartwright, N. and Hardie, J. (2012). Evidence-Based Policy: A Practical Guide to Doing It Better. Oxford University Press.

List, J. A. (2022). The Voltage Effect: How to Make Good Ideas Great and Great Ideas Scale. Currency.

Deming, W. E. (1986). Out of the Crisis. MIT Center for Advanced Engineering Study.

Shewhart, W. A. (1931). Economic Control of Quality of Manufactured Product. D. Van Nostrand.

Efficiency Valuation Organization. International Performance Measurement and Verification Protocol (IPMVP), Core Concepts. Successive editions.

Ostrom, E. (1990). Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press.

Cooperrider, D. L. and Whitney, D. (2005). Appreciative Inquiry: A Positive Revolution in Change. Berrett-Koehler.

Note on figures. Every constant in this chapter — the critical values, the sample-size constant K = 15.70, the power table, the detection thresholds, the regression-to-the-mean drift, the AR(1) inflation, the design effects, the step function, the fund's regeneration rate and the portfolio inequality — is computed in lib/verify/I_05.py and printed by python3 lib/verify.py I.05. The z values are computed at run time by bisection on the normal cumulative distribution rather than quoted, so any reader with a statistical table can check them. Where this chapter tightens I.01's three-sigma rule, the tightened figure is given and the difference stated in the text.