Haute Lumière
Commerce · V.10 · MMXXVI · daylight
For the person studying this on their own time, with a notebook, a spreadsheet and no institutional budget. A term of practice, a term project, and a self-assessment you can actually mark.
Most of what is taught about surveys is taught as a list of things not to do. This one is built the other way round: you are going to run an instrument, on a population you can reach, for long enough that the interesting problems arrive on their own.
Everything in the chapter is reproducible with a spreadsheet. There is no software to buy. The three quantities that decide everything — the standard deviation of a 0–10 item (about 1.9 points), the design effect, and the test–retest correlation (take 0.60) — are three cells, and every sample-size question in this field is one formula away from them.
One instruction before you start, and it is the whole ethic of the chapter. You will measure people. Ask permission, say what it is for, keep it anonymous, and report back what you found. A respondent who never hears the result learns that answering was pointless, and they are right.
Exercise 1.1 — The four questions, from the source. Find the ONS personal well-being guidance and copy the four items exactly, including the introductory sentence and the response-scale labels. Then find the Cantril ladder wording in Gallup's World Poll methodology. Write them side by side.
Answer in writing: what does the ladder ask that "how satisfied are you with your life nowadays" does not?
Exercise 1.2 — The methods note audit. Take any published well-being statistic you can find — a newspaper figure, a company's engagement report, a national statistic. Find its methods note. For each, record six things:
mode month wording (verbatim)
response rate design effect minimum detectable effect
Most will not report all six. Note which are missing; that is the finding. A statistic that cannot tell you its response rate cannot tell you anything.
Exercise 1.3 — Answer it yourself, four ways. Answer the ONS life-satisfaction item four times over one week: once first thing in the morning, once last thing at night, once out loud to another person, once on a screen alone. Record the four numbers and the circumstances.
You will very likely produce a spread. That spread is the mode effect, the seasonality and the context effect, measured on a sample of one — and it is the single most useful hour in this workbook, because from here on you will never again read a national well-being figure as though it were a thermometer reading.
Exercise 2.1 — The power sheet. Build a spreadsheet with these cells and nothing else:
z_alpha 1.959964 two-sided 0.05
z_beta 0.841621 power 0.80
sigma 1.9 SD of a 0-10 item
delta (your input) effect to detect
m (your input) cluster size
rho_icc 0.05
CV_w (your input) 0 for a census
DEFF_c = 1 + (m - 1) * rho_icc
DEFF_w = 1 + CV_w^2
n_arm = 2 * (z_alpha + z_beta)^2 * sigma^2 / delta^2 * DEFF_c * DEFF_w
Check it against the chapter. With delta = 0.10, m = 10, CV = 0.50 you should get 5,667 before the design effect, a design effect of 1.8125, and 10,272 per arm — 20,543 both arms. If you do not reproduce those three numbers exactly, the error is in your sheet and finding it is the exercise.
Exercise 2.2 — The panel row. Add one row:
rho 0.60
n_pairs = (z_alpha + z_beta)^2 * 2 * sigma^2 * (1 - rho) / delta^2 * DEFF_c
With delta = 0.20 and m = 12 you should get 567 before the design effect and 879 after. Compare against the cross-sectional answer for the same delta — 4,392 both arms.
Write one sentence explaining why the panel needs a fifth as many people. If your sentence does not contain the phrase the same person, write it again.
Exercise 2.3 — Turn it round. Rearrange your sheet to solve for delta given n. This is the minimum detectable effect and it is the number you will use more than any other:
MDE = sqrt( (z_alpha + z_beta)^2 * 2 * sigma^2 * (1 - rho) * DEFF / n_pairs )
Check: 780 pairs, DEFF 1.55, rho 0.60 gives 0.21 points.
Exercise 2.4 — The ordinal reversal, by hand. Reproduce the chapter's worked reversal in your sheet. Populations A and B, raw means 5.80 and 5.40, difference +0.40. Apply the transform 3 → 0.0, 5 → 6.0, 9 → 7.0 and confirm you get 5.10, 5.55 and −0.45.
Then do the thing the chapter did not: find a different order-preserving transform that makes A's lead bigger. Both exist. That is the point.
Exercise 2.5 — The then-test arithmetic. Take the chapter's figures: 6.80 at wave one, 6.90 at wave two, 6.10 retrospectively. Confirm the observed change of 0.10, the true change of 0.80, the recalibration of 0.70, and that the raw series hides 87.5 percent of the change.
Now invert it: construct a case where recalibration makes a worsening look like an improvement. Write the three numbers.
Exercise 3.1 — Choose a population you can actually reach. A seminar group, a club, a shared house, a team, a course cohort, an online community you are part of. Thirty people is enough to learn everything and too few to conclude anything, and knowing the difference is the skill.
Write down, before you collect anything, the MDE for the population you have. It will be large. Write it on the front of the file.
Exercise 3.2 — Freeze the instrument. Four items, ONS wording, one screen, no preamble that hints at what you hope to find. Fix the day of the week and the time of day. Write a one-page protocol covering mode, month, wording, order and the anonymity promise, and sign it with the date.
Exercise 3.3 — Write three vignettes. Three short descriptions of imaginary people, spanning the range, each two or three sentences and each concrete. Not "somebody moderately happy" — a person with a named situation. Ask respondents to rate them on the same scale.
Exercise 3.4 — Wave one. Publish nothing. This is the hardest exercise in the workbook and it is one line: collect the baseline and report no findings from it. Write down, in advance, that you will not, and why.
Exercise 3.5 — Wave two, twelve weeks later. Same instrument, same day of week, same time. Add one then-test item: thinking back to when you first answered this, and using the scale as you use it today, how satisfied were you then?
Compute three numbers: observed change, true change, recalibration. Compare each to your MDE. Most of them will be inside it. Say so, in the report, in the first paragraph.
Exercise 4.1 — The one-page methods note. Mode, month, wording, order, response rate, design effect, minimum detectable effect, and — the line that makes the rest trustworthy — what you did not look at. One page. Written for somebody who did not do the study and is inclined to doubt it.
Exercise 4.2 — The sentence you cannot write. List three sentences your data does not support, and why. For example: "the seminar group is happier than the sports club" — a comparison of levels across populations whose reporting functions you observed only through your vignettes, with a gap almost certainly inside your MDE.
Learning which sentences you cannot write is most of what separates a person who can read this literature from a person who quotes it.
Exercise 4.3 — Report back. Give the respondents their result. One page, plain, no jargon, including the sentence about what you could not detect. Watch what happens to their willingness to answer next time. That is the response rate, being earned.
Build and run a two-wave well-being panel on a population of at least thirty people, and produce four artifacts.
Marked on: whether the protocol was frozen before collection (the single heaviest weight); whether the MDE appears before any finding; whether the then-test was run and interpreted; whether the report distinguishes a null result from an underpowered one; and whether the respondents got their result back.
Not marked on: whether you found anything. A well-run study that detected nothing and says clearly what it would have detected is a complete piece of work.
Mark yourself honestly on each, nought to three.
Under sixteen: reread Briefs 3, 4 and 7 and redo Exercises 2.1 to 2.3.
Three things worth doing after the term ends, in ascending order of value.
Keep answering your own four items, same day each month, for a year. You will have a personal series with a protocol, which is more than most organisations have, and by month nine you will have felt response shift from the inside.
Offer the instrument to a group that wants it. A student society, a small charity, a team. They will not have a protocol and you can write them one in an afternoon. This is a genuinely valuable thing to be able to do and almost nobody can do it.
Find one published well-being claim and check it properly. Locate the methods note, compute the MDE if it is computable, look for a mode change at the date the number moved. If you find one, write it up — a well-documented instrument artefact is a real contribution, and the field has more of them than it has found.
The two are the same thing here, and it is worth seeing why rather than being told.
An instrument that people trust gets answered. An instrument people do not trust gets answered by the people who have something to say and skipped by everyone else, which is non-response bias, which is the largest and least-discussed source of error in this entire field. Every ethical practice in the list below is also, independently, a precision improvement.
Anonymity, and say how it is achieved. Not "your responses are confidential" — "your answers are stored against a random identifier; the person analysing them cannot see who you are; the identifier lets us match your answer to your answer last year and nothing else." People who understand the mechanism answer more honestly than people who have been reassured about it.
No open text you will not read. A free-text box that nobody reads is a promise broken silently, and the people who wrote paragraphs into it will not answer next year.
Tell them the length before they start, and be accurate. Nine and a half minutes, said in advance and true, produces a higher completion rate than five minutes said in advance and untrue.
Give the result back within a month. This is the single largest determinant of your wave-two response rate, and wave-two response rate is the single largest determinant of whether you have a series at all.
And never ask a question you are not prepared to act on. If you ask about something you have no ability to change, you have converted a measurement into a grievance collection, and the people who told you the truth will watch nothing happen.
Write these five into your protocol page in Exercise 3.2, in your own words, before wave one. Then hold yourself to them, and notice at wave two what holding them bought you.
Three, and in this order, because each makes the next one legible.
The OECD Guidelines on Measuring Subjective Well-being (2013). Free, and the model questionnaires at the back are the practical core. Read the modules before the theory.
Bond and Lang, "The Sad Truth about Happiness Scales" (2019), with Kaiser and Vendrik's reply. Read them as a pair, in that order, and form your own view. This is the argument of the field, conducted properly, and watching two competent people disagree about something real is a better education than either paper alone.
Sprangers and Schwartz on response shift (1999). Short, clear, and it will change how you read every before-and-after study you encounter for the rest of your life.