Haute Lumière
Commerce · V.10 · MMXXVI · daylight
For the person with a P&L, a signature limit, a board and a quarter. This workbook uses the language of the firm without apology, because the firm's own numbers already support most of what follows — they have simply never been arranged to show it.
You are almost certainly already buying an employee survey. It probably costs more than the instrument in this chapter, produces a score you cannot defend against a methods question, and has changed vendor at least once since it started — which means the series before the change and the series after it are two different series wearing one name.
The proposal here is not to spend more. It is to spend the same money on something that can be audited.
Three commercial facts, computed in full in the chapter and reproducible in lib/verify/V_10.py:
2σ²(1 − ρ) rather than 2σ².That third fact is the commercial argument, and it is unusual: the cheap thing is the one that works, and the expensive thing is the one you are doing.
Exercise 1.1 — Audit the instrument you already own (two hours, with HR).
Pull the last five years of your employee survey and record six facts per year:
vendor mode field month
item wording item order response rate
Then mark every year in which any one of the six changed. Each mark is a break in comparability, and the number you have been quoting across it is two numbers.
Most firms find at least one break. Some find a break in every second year. You have just discovered why the trend line never made sense, and the discovery cost two hours and nothing else.
Exercise 1.2 — Find the minimum detectable effect of what you already run.
Take your headcount, your response rate and your team size. Compute:
DEFF = 1 + (m − 1) × 0.05
MDE = √( 7.848879 × 2 × 1.9² × (1 − ρ) × DEFF / n )
For a firm of 1,200 with 65 percent completing both waves of a panel — 780 usable pairs, teams of 12, ρ = 0.60 — the answer is 0.21 points. Set that against a documented mode effect of 0.20 and the ratio is 1.06.
Write that ratio on one line and take it to the next management meeting. It says, in a form nobody can argue with, that your current instrument cannot distinguish its own signal from a change of vendor.
Exercise 1.3 — Find what is already working.
Two questions, asked out loud. Which measurement about people has this firm asked the same way for more than three years? and who has been quietly protecting that consistency? There is usually one, and it is usually somebody junior who has simply refused to change a form. Find that person. They are your protocol owner and they have already been doing the job.
Exercise 2.1 — Size the instrument for your own firm.
n per arm = 2 (z_α/2 + z_β)² σ² / δ² × DEFF
n pairs = (z_α/2 + z_β)² · 2σ²(1 − ρ) / δ² × DEFF
z_α/2 = 1.959964 z_β = 0.841621 σ = 1.9 points
ρ = 0.60 ρ_icc = 0.05
Decide δ first, and decide it as a business question: what size of change would make us do something differently? Most firms say a fifth of a point once they think about it. At δ = 0.20, teams of 12:
cross-section, both arms 4,392
panel 879
Exercise 2.2 — Build the budget line.
respondent time 879 × 9.5 min = 139.18 h × $47.20 = $ 6,569.06
platform + analyst = $18,000.00
---------------------------------------------------------------------
annual run cost = $24,569.06
per respondent per year = $ 27.95
first-year respondent cost (baseline + wave two) = $13,138.12
Compare against your current employee survey invoice. In most firms the honest instrument is cheaper, and the difference is the line that gets the proposal signed.
Exercise 2.3 — Value what you are detecting.
HM Treasury values one WELLBY — a point of life satisfaction, one person, one year — at £13,000 for public appraisal. Your detectable effect of 0.20 points is therefore worth £2,600 per person per year on that scale, against a cost of $27.95 to find out: roughly 93 times.
Say what this is and is not, in the board paper, in the same breath. It is a price check that shows the measurement is cheap relative to what it measures. It is not a booking entry; the Green Book value is for appraising public policy and no firm may recognise £2,600 a head on the strength of a survey. A board paper that makes that distinction itself will be trusted on everything else in it.
Exercise 2.4 — Write the sentence about what the number cannot do.
One paragraph, in the board paper, before any finding:
This instrument compares us to our own past on a frozen protocol. It does not support comparison to an industry benchmark collected by a different vendor with different wording in a different month, because the mode effect alone is worth 0.10 to 0.30 points and the benchmark gap under discussion is usually smaller than that.
Executives who write this paragraph stop being asked for benchmarks.
Exercise 3.1 — Freeze it.
Four items in ONS wording. One mode. One field month. One item order. One page of protocol, signed by the owner and the deputy. The signature is the control.
Exercise 3.2 — Negotiate the parallel-run covenant.
This is the commercially important half of the whole programme, and it goes into the vendor contract:
Any change to mode of administration, item wording, item order, field month or panel provider shall be preceded by a parallel run of no fewer than 1,000 respondents on both the incumbent and successor instruments, at the vendor's cost, with the bridge coefficient computed and delivered before the successor instrument is used for publication.
The run costs $7,473.33 in respondent time — 30.4 percent of one year's run cost — and it converts your largest measurement risk into a contractual obligation of the counterparty. Vendors sign it. It is a reasonable term and they know it.
Exercise 3.3 — Register the asset.
The run cost is expensed; an internally generated database is generally not capitalised under IAS 38 and nobody should pretend otherwise. So register the series in the data asset inventory at replacement cost: five years of protocol is $122,845.30 of spend and five years of calendar time.
Write both figures in the register, and write the second one in words as well: five years of elapsed time, which cannot be purchased. That line is what stops the series being cancelled in a thin year, because it is the only asset in the inventory that money cannot rebuild.
Exercise 3.4 — Fund the continuity reserve.
Three years at the annual run cost: $73,707.18, held outside the department that consumes it. This is not a contingency against overrun — the cost does not overrun, it is respondent time and a licence. It is a defence against discontinuation, which is the actual failure mode.
Exercise 4.1 — Field the baseline and publish nothing.
Write to the board before wave one confirming that the baseline will produce no findings. The pressure to find something in a baseline is what corrupts baselines, and pre-committing removes it.
Exercise 4.2 — Put it in the standing pack.
Anything reviewed on a cadence persists; anything reviewed by exception does not. One page: the four item means, the change, the MDE, the response rate, the protocol-adherence confirmation from internal audit.
Exercise 4.3 — Name the second owner.
One owner is a hobby; two is a practice. Change control requires both signatures plus a funded parallel run. Recruit the second owner by giving them the credit for the first release.
Exercise 4.4 — Decide, now, that it is not in anybody's bonus.
Put it in the governance page in writing, before anyone asks. A well-being item is unusually easy to move without moving anything real, because the respondent controls it entirely. The moment it becomes a target you have bought an expensive way of learning what people think you want to hear.
The commercially literate framing: this is an instrument, not a KPI. You do not put your thermometer in the bonus plan.
A vendor change nobody parallel-ran. The most common and most expensive. Prevented by the covenant in Exercise 3.2, and by nothing else.
A null read as no effect. The team ran a good programme, the series shows flat, and nobody computed the MDE — so an underpowered study becomes evidence against the programme. Prevented by printing the MDE before the finding, every time.
A flat year that was actually a good year. Response shift: people raised their standards while their lives improved. In the chapter's worked case the raw series shows 0.10 points when the true change was 0.80 — it hides 87.5 percent of it. Prevented by a then-test item, which costs one question.
A benchmark comparison in a board deck. Prevented by the paragraph in Exercise 2.4, written once and reused.
The quiet cancellation. Year four, thin budget, nothing dramatic in the numbers, and the series looks optional. Prevented by the reserve and by the replacement-cost line in the asset register.
| Day | Action | Artifact |
|---|---|---|
| 1–15 | Audit five years of the existing survey for breaks | The break register |
| 16–30 | Compute σ, DEFF and MDE for your own population | The power memo |
| 31–45 | Size and budget the panel; write the benchmark paragraph | The board paper |
| 46–60 | Negotiate the parallel-run covenant into the vendor contract | The signed covenant |
| 61–75 | Field wave one. Publish no findings | Baseline and methods note |
| 76–90 | Register at replacement cost; fund the reserve; name the deputy | The continuity facility |
One page. In this order, and no other.
That sixth line is the one that decides it, and it should be the last sentence on the page.
With the CFO. The framing that works is not wellbeing; it is series integrity. A CFO has seen a restatement and knows what a broken comparative costs. The sentence is: we have been reporting a trend across at least one undocumented methodology change, and there is a contractual term that prevents it recurring at the vendor's cost. Bring the break register from Exercise 1.1 and the covenant from Exercise 3.2, and nothing else.
With HR. The risk here is that this reads as a criticism of the engagement survey, and it is not — the engagement survey answers a different question and should continue. The distinction to draw: an engagement instrument is diagnostic and is meant to change often as the organisation's questions change; a wellbeing series is longitudinal and is worthless if it changes at all. They are different instruments with different rules, and the mistake almost everyone makes is running one instrument under both sets of rules. Propose them as two.
With the works council, union or employee forum. Lead with the governance page, not the arithmetic. The three lines they will care about are: it is anonymous and here is the mechanism; it is not attached to anybody's pay and that is in writing; and the result comes back to the people who gave it, within a month, including what it could not detect. Every one of those is also a precision improvement, because an instrument people trust is an instrument people answer, and non-response is the largest unmeasured error in the whole apparatus.
Three adjacent uses, once the panel exists and the protocol holds.
Programme evaluation for free. Any intervention that reaches part of the workforce and not another part — a shift-pattern change, a new manager development programme, a site refurbishment — is now evaluable against a within-person baseline that already exists. The marginal cost of evaluating it is zero, because the measurement was already running. Firms that build this find within two years that they stop commissioning bespoke evaluations.
A defensible input to the annual report. Not a score. A method: "we measure workforce wellbeing on a frozen protocol with a published minimum detectable effect and a contractual parallel-run requirement." That sentence survives assurance. A benchmark score does not.
A recruitment asset that is true. Organisations that publish what their instrument could not detect are rare enough that candidates notice. It is the same signal as a well-written methods note, pointed at a different audience.