Haute Lumière
Commerce · V.09 · MMXXVI · daylight
One page each. A reader who reads only these ten pages has the chapter.
The idea. People who join a voluntary health programme are, on average, already healthier and already cheaper than people who do not. That is not a failure of the programme. It is a fact about volunteering, and it exists before the programme does. Any evaluation that compares participants with non-participants is therefore measuring two things at once — who joined, and what joining did — and reporting the sum as though it were the second.
Worked example. A self-insured employer of 5,000 people runs a programme at 150.00 USD a head, so 750,000 USD a year. Participation is 56.0 per cent, giving 2,800 participants and 2,200 others. Afterwards, participants average 5,100 USD of claims and non-participants 6,600 USD — a gap of 1,500 USD. Multiply by 2,800 and the programme appears to have saved 4,200,000 USD, a return of 5.60 : 1.
The figure to carry. In the same data, before the programme existed, the participants averaged 4,900 USD and the others 6,100 USD. 1,200 USD of that 1,500 USD gap was already there.
Why it matters. The Illinois Workplace Wellness Study randomised 12,459 employees precisely so this could be measured rather than argued about, and its confidence intervals exclude 78.0 per cent of the previously published savings estimates.
You already know this because you have noticed which of your colleagues sign up for the running club, and it is not a random sample of the office.
The idea. If two groups differ before an intervention and after it, the estimate of the effect is not the difference between them; it is the change in the difference. Subtract the before-gap from the after-gap and what remains is the part the intervention could plausibly have caused, on the assumption that the two groups would otherwise have moved in parallel.
Worked example. Continuing Brief 1: the after-gap is 1,500 USD and the before-gap is 1,200 USD, so the estimated effect is 300 USD a participant. Times 2,800 participants, that is 840,000 USD against a cost of 750,000 USD — a return of 1.12 : 1 rather than 5.60 : 1.
The figure to carry. The naive figure is 5.00 x the difference-in- differences figure. The general form is worth memorising:
naive / true = 1 + (baseline gap / causal effect)
A baseline gap two to four times the size of the effect makes the naive number 3, 4 or 5 times too large. That is where the phrase three to five times comes from — it is arithmetic, not an accusation.
What it still cannot fix. Parallel trends is an assumption, not a fact. People often join a programme in the year they have decided to change, so the groups are diverging for reasons the method cannot see. At the randomised point estimate the effect is 0.00 USD and the year's net is −750,000 USD.
You already know this because you have compared this month with last month and then remembered that last month had a bank holiday in it.
The idea. An extreme measurement is extreme partly because of a lasting property and partly because of luck. Luck does not repeat. So any group selected because it was extreme will, on its next measurement, be less extreme — with no intervention at all.
Worked example. Medical spending has a population mean of 6,000 USD, a standard deviation of 12,000 USD, and a year-to-year individual correlation of about 0.35. Select the top decile. For a standard normal, the expected value above the ninetieth percentile is 1.7550 standard deviations, so this year's selected group averages 27,060 USD. Next year, doing nothing whatever, its expected average is 6,000 + 0.35 × (27,060 − 6,000) = 13,371 USD.
The figure to carry. 13,689 USD of apparent saving a head — 50.6 per cent — generated by arithmetic and nothing else. And it is an understatement, because the normal approximation has a thinner tail than real medical spending.
Why it matters. Every high-risk targeting programme in existence enrols on last year's spend. Every one of them will therefore report a large saving in year one. The only defence is a control group selected the same way.
You already know this because the second time you took the eye test you did worse, and the optician was not surprised.
The idea. Most workplace programmes are deployed in waves, because capacity is finite. The order of those waves is usually chosen by convenience. Choose it at random instead and the sites not yet reached become a randomised control group. The marginal cost is zero and the epistemic gain is total.
Worked example. The BJ's Wholesale Club trial did exactly this at scale: 160 worksites, 20 randomised to treatment and 140 to control, covering 32,974 employees — 4,037 at treatment sites, 28,937 at control sites, about 201.8 and 206.7 people per site. It found significantly higher self-reported exercise and weight management, and no significant difference in clinical measures, spending, utilisation, absenteeism, tenure or job performance, at eighteen months and again at three years.
The figure to carry. A staggered rollout with a random order is a trial. A staggered rollout with a convenient order is an anecdote. The difference in cost between them is 0.00 USD.
You already know this because you have already decided which region goes first, and you could not give a principled reason for it if asked.
The idea. A study can only detect an effect large relative to the noise around it. The sample needed grows with the square of the variance and falls with the square of the effect. Medical spending is extraordinarily noisy; turnover is not. So the choice of endpoint is not a matter of taste — it decides whether the study can exist.
Worked example. To detect a 300 USD change in annual spending with a standard deviation of 12,000 USD, at eighty per cent power and a five per cent two-sided test, you need 25,088 per arm — 50,176 employees. A firm of 5,000 is short by 10.04 x. To detect a four-point fall in a twenty per cent voluntary turnover rate you need 1,568 per arm, 3,136 in all, which the same firm clears 1.59 x over.
The figure to carry. 50,176 against 3,136. The same firm, the same money, one endpoint out of reach and one comfortably inside it.
The rule that follows. Do not buy a claim about an endpoint you could not have detected. Buy the claim you can hold the vendor to.
You already know this because you can tell whether your commute got longer this month, and you cannot tell whether your blood pressure did.
The idea. Presenteeism is productivity lost by people who are at work and unwell. It is measured by asking them, on scales such as the WHO Health and Work Performance Questionnaire or the Work Limitations Questionnaire, and then converted to money by multiplying reported lost time by the wage. Both steps are soft: the instruments agree with one another only moderately, and the multiplier linking an hour of lost output to an hour of wage is a property of the job, not a constant.
Worked example. A salary of 60,000 USD, reported impairment of 5.0 per cent of working time, a claimed relative reduction of 10.0 per cent:
multiplier 0.5 impairment 1,500.00 USD saving 150.00 USD 1.00 : 1
multiplier 1.0 impairment 3,000.00 USD saving 300.00 USD 2.00 : 1
multiplier 1.5 impairment 4,500.00 USD saving 450.00 USD 3.00 : 1
The figure to carry. The same programme returns anywhere from 1.00 : 1 to 3.00 : 1 — a spread of 3.00 x — on a parameter chosen in an appendix.
How to use it honestly. Print the multiplier beside the number and show the band. Mark Pauly and colleagues showed the ratio is high where work is timed and team-dependent and near or below one where it is not. A presenteeism figure without a stated multiplier is not a measurement.
You already know this because you have had a day where you were present and useless, and a day where you were ill and got more done than usual.
The idea. Absence is recorded by payroll for reasons that have nothing to do with health economics. It is therefore an administrative fact rather than a report: nobody is asked, nobody estimates, and the person collecting it has no stake in the answer. Its weaknesses are real — it misses short informal absence, it counts a day the same whoever took it, and a good programme can raise it by sending people to appointments — but they are weaknesses of coverage, not of credibility.
Worked example. The published meta-analytic claim is that absence costs fall by 2.73 USD per dollar spent, which alongside the 3.27 USD medical claim makes 6.00 USD of asserted return per dollar. Both randomised trials measured absence from administrative records. Neither found a significant effect.
The figure to carry. The medical half of that claim, applied to a 150.00 USD programme, asserts a saving of 490.50 USD an employee — against an employer single-coverage cost of 8,435 USD less a 1,401 USD worker contribution, or 7,034 USD, which is 6.97 per cent of the entire medical bill.
Why it matters. Given a soft big number and a hard small one, take the hard small one. It is the one that survives a second reader.
You already know this because you trust the odometer more than you trust your memory of the journey.
The idea. A quality-adjusted life year is time multiplied by a utility weight and summed:
QALY = SUM over states of ( time in state x utility of state )
utility 1.00 = full health, 0.00 = dead, negative values allowed
discounted: SUM u(h_t) x dt / (1 + r)^t
Utilities come from individual preference elicitation — time trade-off, standard gamble, or a preference-weighted instrument such as EQ-5D. QALYs are gained.
Worked example. Move 100 people from a utility of 0.78 to 0.85 — a gain of 0.07 — and hold it five years. The annuity factor at three per cent is 4.5797, so the gain is 100 × 0.07 × 4.5797 = 32.06 QALYs. At a programme cost of 600,000 USD that is 18,716 USD per QALY.
The figure to carry. Against a published threshold range of 20,000 to 30,000 GBP — 25,400 to 38,100 USD at a working rate of 1.27 — that is 1.36 x cheaper than the cheap end. Against the measured displacement figure of 12,936 GBP, or 16,429 USD, it is marginal. Both comparisons are legitimate and they answer different questions; Chapter V.01 sets out why.
You already know this because you have chosen a shorter recovery over a better outcome, or the reverse, and you knew you were trading two different goods against each other.
The idea. A disability-adjusted life year counts health lost:
DALY = YLL + YLD
YLL = deaths x reference remaining life expectancy at the age of death
YLD = prevalence x disability weight (0.00 = full health)
The scale runs the opposite way to a utility, and — this is the part routinely got wrong — the numbers are not complements. Disability weights come from population paired-comparison surveys; utilities come from individual preference elicitation. One minus a utility is not a disability weight. DALYs are averted; QALYs are gained. Since GBD 2010 the headline figures carry no age weighting and no discounting.
Worked example. Twelve workers injured, disability weight 0.20, duration 0.50 years: YLD = 12 × 0.20 × 0.50 = 1.20 DALYs. One fatality at 45 against a GBD reference life expectancy at birth of 88.87 years gives a YLL of 43.87 years, approximating the reference table by subtraction and saying so. Total: 45.07 DALYs.
The figure to carry. 43.87 of those 45.07 are one death. Fatality dominates injury in this arithmetic by more than thirty to one, which is why a safety case built on injury frequency alone understates itself badly.
You already know this because you already treat a near-miss on a ladder differently from a hundred paper cuts, and you were right to.
The idea. An employer funds health and receives the benefit only while the person stays. Model separations as exponential with hazard λ = ln 2 / median tenure. A perpetual benefit stream B is worth B / r to society and B / (r + λ) to the firm, so the firm's share is r / (r + λ).
Worked example. Median tenure is 3.90 years, so λ = 0.17773 a year. At a real discount rate of 0.03:
0.03 / (0.03 + 0.17773) = 0.03 / 0.20773 = 0.1444 = 14.44 %
hurdle multiple = 1 / 0.1444 = 6.92 x
The figure to carry. An employer keeps about 14.44 per cent of the long-run value of the health it creates, and therefore needs the investment to be 6.92 x better than society does before its own arithmetic approves it.
What follows from it. Move the owner. At 12.00 years of tenure in a trade rather than a firm, the hazard is 0.05776, capture rises to 34.18 per cent, the hurdle falls to 2.93 x, and the identical investment is 2.37 x more valuable to the party making it. Nothing about the medicine changed. This is why the chapter's instrument is a pooled fund rather than a company benefit.
You already know this because you have never repainted a flat you were leaving in six months, and you did not think of yourself as a bad person for it.