Haute Lumière
Commerce · I.05 · MMXXVI · daylight
Three instruments: a ten-point quiz, eight reflection questions, five essay prompts. The quiz checks comprehension rather than recall. The reflections are private and first-person. The essays are arguable from more than one side.
Four on recall.
1. Write the expression for the detectable effect of a before-and-after comparison, and define every term.
detectable effect = (z_α/2 + z_β) · σ · √(1/m + 1/k), wheremis the number of pre-periods,kthe number of post-periods, andσthe period-to-period standard deviation. At α = 0.05 two-sided and 80 percent power,(z_α/2 + z_β) = 2.8016. One mark for the expression, one for identifying thatmandkenter through their reciprocals — which is why adding pre-periods, which are free, moves the threshold.
2. Define the realisation rate φ, and state the test for whether a saving counts as banked.
φ = banked saving / measured saving. A saving is banked only if you can name the budget line it falls out of and the person whose budget falls. If you cannot name both,φ = 0for that saving however well it was measured.
3. Name the four confounders the chapter treats, and one sentence on each.
Regression to the mean — a site chosen because it performs badly improves on its own. Serial correlation — successive periods are not independent, so the measured
σunderstates uncertainty. Clustering — you randomise sites, so yournis sites, not people. The enthusiast — analysis chosen after the result is seen manufactures significance from nothing.
4. What are the three governance roles, and which two may not be held by the same person?
Sponsor (carries the decision and the budget), operator (runs the intervention), verifier (named before the result, reporting to finance rather than to the sponsor). The operator and the verifier may never be the same person; the verifier's independence is the whole of their value.
Four on application.
5. A colleague reports: "We ran it for a month at the Leeds depot, our worst site, and cost per drop fell eleven percent. It works." Diagnose it.
Three faults, in descending seriousness. Site selection: Leeds was chosen because it was worst, so regression to the mean predicts improvement with no intervention — at reliability 0.6 and the 10th percentile, about 0.51σ of it. Power: one period before and one after sets the detection threshold near 3.96σ, so an eleven percent move may well be inside the noise. No control: without a comparison arm, nothing separates the intervention from the season, the weather or the regression. The strongest answer notes that the fix for all three is one decision — a control group from the same estate — and that it costs a query.
6. Your pilot releases 990 hours a year of warehouse picking time, verified, uncontested. The finance director asks what it saves. Answer precisely.
Nothing yet. One shift-year is 1,800 hours;
claimable = step × floor(saving / step), and 990 hours is below one whole step, soφ = 0. The honest statement is capacity released, not cost avoided — a different claim, made to a different person, for a different purpose. Credit any answer proposing the two legitimate routes: scale the intervention until it crosses a whole step, or aggregate it with other savings falling on the same step.
7. You have eleven branches, thirty staff each, and an ICC of 0.05 on the outcome you care about. Your power calculation said 63 per arm. What do you tell the sponsor?
The design effect is
1 + (30 − 1) × 0.05 = 2.45, so the requirement is 154 per arm, and eleven branches of thirty cannot supply it. Say so before spending. The stronger answer gives the three routes: lengthen the pre-period to lower the threshold; choose an outcome with a lower ICC; or relabel the work a feasibility study with no rate claimed.
8. Why does the chapter insist on a portfolio inequality rather than a per-pilot return?
Because a single pilot is a bet and only a portfolio is a rate. Attrition is real — some pilots produce no verified saving — and a fund that prices every pilot as a success will be embarrassed once and closed. The portfolio form carries
p, the share of pilots that work, explicitly, which lets the fund be underwritten rather than believed in.
Two that require the arithmetic to be done.
9. Your monthly variation is 12 percent and you want to detect a 6 percent improvement, at 5 percent significance and 80 percent power. How many units per arm? Then repeat with thirty people per site and an ICC of 0.05. Show your working.
n per arm = 2(z_α/2 + z_β)² (σ/δ)² = 15.70 × (12/6)² = 15.70 × 4 = 62.79, so 63 per arm, 126 in total. With clustering,DEFF = 1 + 29 × 0.05 = 2.45, so62.79 × 2.45 = 153.8→ 154 per arm. Credit any method reaching 62–63 and 153–155. The point of the question is that the design effect is not a rounding — it is two and a half times the trial.
10. A pilot fund spends £180,000 a pilot. A successful pilot produces £132,000 a year of measured saving, and the measured realisation rate is 0.55. What is the fund's regeneration rate, and how long does it take to double?
r = φ S / C = 0.55 × 132,000 / 180,000 = 0.4033, i.e. 40.3 percent a year. Doubling time= ln 2 / ln(1 + r) = 0.6931 / 0.3389 = 2.05 years. The stronger answer states the consequence in the right register: the successor this pilot can fund in one cycle is 0.40 times its size, so the proposal should ask for 0.40× and be taken four times, not 3× and be refused once.
These are not for a room. Write the answers by hand if you can; the slowness is the point.
Each is arguable from more than one side. Each requires at least one source the chapter cites and at least one it does not.
1. The ethics of the staggered rollout. The chapter recommends randomising the order of a rollout rather than access to it, on the grounds that nobody is denied anything. Argue either that this fully resolves the ethical objection to experimenting on colleagues or customers, or that it merely relocates it — since being last still carries a cost, and the people who bear it did not consent. Use Levy on Progresa, and one source on research ethics or informed consent in field experiments that the chapter does not cite.
2. Is the pre-registration a cure or a costume? Pre-registration is imported here from clinical trials and from the credibility reforms in psychology and economics. Argue whether it genuinely constrains analytical freedom inside a firm — where nobody polices it and the document can be revised — or whether its real function is rhetorical, buying credibility rather than creating it. Engage Simmons, Nelson and Simonsohn (2011), and at least one published critique or empirical evaluation of pre-registration that the chapter does not cite.
3. The step function and the moral status of released capacity. The chapter argues that a saving which does not make a budget line fall has φ = 0. Take a position on whether this accounting is honest realism or a framework that systematically undervalues improvements to human working life — less rework, less strain, more slack — precisely because they are absorbed rather than banked. Use the chapter's warehouse arithmetic, and one source on slack, capacity or the economics of working time that it does not cite.
4. Scale is a different question from proof. Cartwright and Hardie argue that it worked there does not establish it will work here; List describes the collapse as a voltage drop. Argue either that rigorous small-scale evidence is the necessary foundation for scaling and that the voltage problem is an engineering matter downstream, or that the two are different epistemic problems and that optimising a pilot for internal validity actively harms its external validity. Use Cartwright and Hardie or List, and one empirical study of a scaled intervention that the chapter does not cite.
5. The measured negative. The Oregon Health Insurance Experiment found substantial effects on financial strain and self-reported health, and no detectable effect on three clinical markers over two years. Write the case that this is a model of what honest evaluation should produce — then the case that a two-year window on chronic disease markers was under-powered by design and that the null was over-read. Conclude with which you find more persuasive and why. Use Finkelstein et al. (2012) or Baicker et al. (2013), and at least one subsequent commentary or re-analysis that the chapter does not cite.