Haute Lumière

Commerce · V.10 · MMXXVI · daylight

La Bourse  /  Volume V  /  Nº V.10  /  Workbook — the student

A watercolour portrait of a woman in a green shirt and camel jacket among leaves and sunflowers.
Plate V.10 · Workbook — the studentEleven Rungs.Every measurement of a life begins with a person translating a life into a number. The instrument is not the questionnaire. The instrument is that translation, and it is done by the respondent, alone, in about four seconds.

WORKBOOK — THE STUDENT

Chapter V.10 · The Thriving Survey

For the person studying this on their own time, with a notebook, a spreadsheet and no institutional budget. A term of practice, a term project, and a self-assessment you can actually mark.


WHY THIS WORKBOOK IS DIFFERENT

Most of what is taught about surveys is taught as a list of things not to do. This one is built the other way round: you are going to run an instrument, on a population you can reach, for long enough that the interesting problems arrive on their own.

Everything in the chapter is reproducible with a spreadsheet. There is no software to buy. The three quantities that decide everything — the standard deviation of a 0–10 item (about 1.9 points), the design effect, and the test–retest correlation (take 0.60) — are three cells, and every sample-size question in this field is one formula away from them.

One instruction before you start, and it is the whole ethic of the chapter. You will measure people. Ask permission, say what it is for, keep it anonymous, and report back what you found. A respondent who never hears the result learns that answering was pointless, and they are right.


PART ONE — DISCOVERY

Weeks 1–4: read the instruments that exist

Exercise 1.1 — The four questions, from the source. Find the ONS personal well-being guidance and copy the four items exactly, including the introductory sentence and the response-scale labels. Then find the Cantril ladder wording in Gallup's World Poll methodology. Write them side by side.

Answer in writing: what does the ladder ask that "how satisfied are you with your life nowadays" does not?

Exercise 1.2 — The methods note audit. Take any published well-being statistic you can find — a newspaper figure, a company's engagement report, a national statistic. Find its methods note. For each, record six things:

  mode            month           wording (verbatim)
  response rate   design effect   minimum detectable effect

Most will not report all six. Note which are missing; that is the finding. A statistic that cannot tell you its response rate cannot tell you anything.

Exercise 1.3 — Answer it yourself, four ways. Answer the ONS life-satisfaction item four times over one week: once first thing in the morning, once last thing at night, once out loud to another person, once on a screen alone. Record the four numbers and the circumstances.

You will very likely produce a spread. That spread is the mode effect, the seasonality and the context effect, measured on a sample of one — and it is the single most useful hour in this workbook, because from here on you will never again read a national well-being figure as though it were a thermometer reading.


PART TWO — THE ARITHMETIC

Weeks 5–8: build the calculator

Exercise 2.1 — The power sheet. Build a spreadsheet with these cells and nothing else:

  z_alpha      1.959964        two-sided 0.05
  z_beta       0.841621        power 0.80
  sigma        1.9             SD of a 0-10 item
  delta        (your input)    effect to detect
  m            (your input)    cluster size
  rho_icc      0.05
  CV_w         (your input)    0 for a census

  DEFF_c  = 1 + (m - 1) * rho_icc
  DEFF_w  = 1 + CV_w^2
  n_arm   = 2 * (z_alpha + z_beta)^2 * sigma^2 / delta^2 * DEFF_c * DEFF_w

Check it against the chapter. With delta = 0.10, m = 10, CV = 0.50 you should get 5,667 before the design effect, a design effect of 1.8125, and 10,272 per arm — 20,543 both arms. If you do not reproduce those three numbers exactly, the error is in your sheet and finding it is the exercise.

Exercise 2.2 — The panel row. Add one row:

  rho       0.60
  n_pairs = (z_alpha + z_beta)^2 * 2 * sigma^2 * (1 - rho) / delta^2 * DEFF_c

With delta = 0.20 and m = 12 you should get 567 before the design effect and 879 after. Compare against the cross-sectional answer for the same delta — 4,392 both arms.

Write one sentence explaining why the panel needs a fifth as many people. If your sentence does not contain the phrase the same person, write it again.

Exercise 2.3 — Turn it round. Rearrange your sheet to solve for delta given n. This is the minimum detectable effect and it is the number you will use more than any other:

  MDE = sqrt( (z_alpha + z_beta)^2 * 2 * sigma^2 * (1 - rho) * DEFF / n_pairs )

Check: 780 pairs, DEFF 1.55, rho 0.60 gives 0.21 points.

Exercise 2.4 — The ordinal reversal, by hand. Reproduce the chapter's worked reversal in your sheet. Populations A and B, raw means 5.80 and 5.40, difference +0.40. Apply the transform 3 → 0.0, 5 → 6.0, 9 → 7.0 and confirm you get 5.10, 5.55 and −0.45.

Then do the thing the chapter did not: find a different order-preserving transform that makes A's lead bigger. Both exist. That is the point.

Exercise 2.5 — The then-test arithmetic. Take the chapter's figures: 6.80 at wave one, 6.90 at wave two, 6.10 retrospectively. Confirm the observed change of 0.10, the true change of 0.80, the recalibration of 0.70, and that the raw series hides 87.5 percent of the change.

Now invert it: construct a case where recalibration makes a worsening look like an improvement. Write the three numbers.


PART THREE — DREAM AND DESIGN

Weeks 9–12: run something

Exercise 3.1 — Choose a population you can actually reach. A seminar group, a club, a shared house, a team, a course cohort, an online community you are part of. Thirty people is enough to learn everything and too few to conclude anything, and knowing the difference is the skill.

Write down, before you collect anything, the MDE for the population you have. It will be large. Write it on the front of the file.

Exercise 3.2 — Freeze the instrument. Four items, ONS wording, one screen, no preamble that hints at what you hope to find. Fix the day of the week and the time of day. Write a one-page protocol covering mode, month, wording, order and the anonymity promise, and sign it with the date.

Exercise 3.3 — Write three vignettes. Three short descriptions of imaginary people, spanning the range, each two or three sentences and each concrete. Not "somebody moderately happy" — a person with a named situation. Ask respondents to rate them on the same scale.

Exercise 3.4 — Wave one. Publish nothing. This is the hardest exercise in the workbook and it is one line: collect the baseline and report no findings from it. Write down, in advance, that you will not, and why.

Exercise 3.5 — Wave two, twelve weeks later. Same instrument, same day of week, same time. Add one then-test item: thinking back to when you first answered this, and using the scale as you use it today, how satisfied were you then?

Compute three numbers: observed change, true change, recalibration. Compare each to your MDE. Most of them will be inside it. Say so, in the report, in the first paragraph.


PART FOUR — DESTINY AND DELIGHT

Weeks 13–15: report honestly and notice what it felt like

Exercise 4.1 — The one-page methods note. Mode, month, wording, order, response rate, design effect, minimum detectable effect, and — the line that makes the rest trustworthy — what you did not look at. One page. Written for somebody who did not do the study and is inclined to doubt it.

Exercise 4.2 — The sentence you cannot write. List three sentences your data does not support, and why. For example: "the seminar group is happier than the sports club" — a comparison of levels across populations whose reporting functions you observed only through your vignettes, with a gap almost certainly inside your MDE.

Learning which sentences you cannot write is most of what separates a person who can read this literature from a person who quotes it.

Exercise 4.3 — Report back. Give the respondents their result. One page, plain, no jargon, including the sentence about what you could not detect. Watch what happens to their willingness to answer next time. That is the response rate, being earned.


THE TERM PROJECT

One instrument, carried the whole way

Build and run a two-wave well-being panel on a population of at least thirty people, and produce four artifacts.

  1. The frozen protocol, signed and dated before wave one — mode, month, wording, order, anonymity, and the pre-registered analysis (what you will compute, decided before you see the data).
  2. The power memo, one page: your σ, your design effect, your n, your MDE, and the honest sentence about what this study can and cannot see.
  3. The two waves, with the then-test at wave two and the vignettes on a rotating subset.
  4. The findings note, one page, opening with the MDE and closing with the three sentences your data does not support.

Marked on: whether the protocol was frozen before collection (the single heaviest weight); whether the MDE appears before any finding; whether the then-test was run and interpreted; whether the report distinguishes a null result from an underpowered one; and whether the respondents got their result back.

Not marked on: whether you found anything. A well-run study that detected nothing and says clearly what it would have detected is a complete piece of work.


SELF-ASSESSMENT

Mark yourself honestly on each, nought to three.

Under sixteen: reread Briefs 3, 4 and 7 and redo Exercises 2.1 to 2.3.


CARRYING IT FORWARD

Three things worth doing after the term ends, in ascending order of value.

Keep answering your own four items, same day each month, for a year. You will have a personal series with a protocol, which is more than most organisations have, and by month nine you will have felt response shift from the inside.

Offer the instrument to a group that wants it. A student society, a small charity, a team. They will not have a protocol and you can write them one in an afternoon. This is a genuinely valuable thing to be able to do and almost nobody can do it.

Find one published well-being claim and check it properly. Locate the methods note, compute the MDE if it is computable, look for a mode change at the date the number moved. If you find one, write it up — a well-documented instrument artefact is a real contribution, and the field has more of them than it has found.


APPRECIATIVE QUESTIONS FOR YOUR SEMINAR

  1. When has a number about people genuinely changed your mind — and what made it credible?
  2. Which of the four ONS items would tell you most about your own week, and why that one?
  3. What is the best-run measurement any of us has been on the receiving end of, and what made it feel worth answering?
  4. If we could ask our whole institution one question every year for twenty years, what should it be?
  5. Where in this room's experience has a flat result actually been good news?
  6. What would it take for the people we measure to want the result back?

A NOTE ON ETHICS, WHICH IS ALSO A NOTE ON METHOD

The two are the same thing here, and it is worth seeing why rather than being told.

An instrument that people trust gets answered. An instrument people do not trust gets answered by the people who have something to say and skipped by everyone else, which is non-response bias, which is the largest and least-discussed source of error in this entire field. Every ethical practice in the list below is also, independently, a precision improvement.

Anonymity, and say how it is achieved. Not "your responses are confidential" — "your answers are stored against a random identifier; the person analysing them cannot see who you are; the identifier lets us match your answer to your answer last year and nothing else." People who understand the mechanism answer more honestly than people who have been reassured about it.

No open text you will not read. A free-text box that nobody reads is a promise broken silently, and the people who wrote paragraphs into it will not answer next year.

Tell them the length before they start, and be accurate. Nine and a half minutes, said in advance and true, produces a higher completion rate than five minutes said in advance and untrue.

Give the result back within a month. This is the single largest determinant of your wave-two response rate, and wave-two response rate is the single largest determinant of whether you have a series at all.

And never ask a question you are not prepared to act on. If you ask about something you have no ability to change, you have converted a measurement into a grievance collection, and the people who told you the truth will watch nothing happen.

Write these five into your protocol page in Exercise 3.2, in your own words, before wave one. Then hold yourself to them, and notice at wave two what holding them bought you.

WHAT TO READ NEXT, IN ORDER

Three, and in this order, because each makes the next one legible.

The OECD Guidelines on Measuring Subjective Well-being (2013). Free, and the model questionnaires at the back are the practical core. Read the modules before the theory.

Bond and Lang, "The Sad Truth about Happiness Scales" (2019), with Kaiser and Vendrik's reply. Read them as a pair, in that order, and form your own view. This is the argument of the field, conducted properly, and watching two competent people disagree about something real is a better education than either paper alone.

Sprangers and Schwartz on response shift (1999). Short, clear, and it will change how you read every before-and-after study you encounter for the rest of your life.