Haute Lumière

Commerce · V.10 · MMXXVI · daylight

La Bourse  /  Volume V  /  Nº V.10  /  Quiz, reflection, essays

A watercolour portrait of a woman in a green shirt and camel jacket among leaves and sunflowers.
Plate V.10 · Quiz, reflection, essaysEleven Rungs.Every measurement of a life begins with a person translating a life into a number. The instrument is not the questionnaire. The instrument is that translation, and it is done by the respondent, alone, in about four seconds.

ASSESSMENT · Chapter V.10 — The Thriving Survey

Three instruments: a ten-point quiz, eight reflection questions, five essay prompts. The quiz checks comprehension rather than recall. The reflections are private and first-person. The essays are arguable from more than one side.


THE QUIZ — ten points

Four on recall.

1. Name the four ONS personal well-being items and say which construct each one measures.

Life satisfaction (evaluative), worthwhileness (eudaimonic), happiness yesterday and anxiety yesterday (affective). All four on a 0–10 scale. One mark for the four items, one for the construct split — the point of the split is that evaluation and affect move apart, which a single-item survey cannot see.

2. What is the difference between a Likert item and a Likert scale, and why does it matter for well-being measurement?

A Likert item is one question with ordered options; a Likert scale is the sum of several. Likert's 1932 instrument was summated, and it is the summing that earns any claim to interval-like behaviour. A single-item life-satisfaction measure inherits none of that protection — it is an ordinal item being used cardinally.

3. State Kish's two components of the design effect and what each one corrects for.

DEFF_c = 1 + (m − 1)·ρ_icc for clustering — people interviewed near each other resemble each other. DEFF_w = 1 + CV²(w) for unequal design weights. Total DEFF is their product. Credit an answer that notes a census removes the weighting term entirely, which is a free 25 percent.

4. Name the three mechanisms of response shift, and the instrument that detects the first of them.

Recalibration, reprioritisation, reconceptualisation (Sprangers and Schwartz, 1999). The then-test — a retrospective item at wave two asking about wave one on today's scale — detects recalibration.

Four on application.

5. A colleague reports that your employee well-being score fell three-tenths of a point this year and proposes an intervention. Before agreeing, what four things do you ask?

What was the minimum detectable effect — is three-tenths inside or outside it? Did the mode, month, wording or item order change since last year? Is this the same people or a fresh sample — because a fresh sample cannot separate a change in lives from a change in who answered? And is there a then-test, so we can tell a fall in lives from a rise in standards? Full marks require the mode question; it is the one that most often explains the whole movement.

6. Your board wants to benchmark your well-being score against an industry average published by a different vendor. What do you say?

That the comparison is a comparison of levels across populations whose reporting functions were never observed, which is the class of claim the ordinal problem destroys — and that the two instruments almost certainly differ in mode, wording, month and item order, any one of which is worth more than the gap being discussed. The constructive half of the answer: the defensible comparison is us against our own past on a frozen instrument, and if the board genuinely needs a between-group comparison, it needs anchoring vignettes.

7. A programme genuinely improved people's lives and the series shows a flat line. Give two distinct explanations and say how you would tell them apart.

Either the design was underpowered — the true effect is smaller than the MDE — or response shift recalibrated the scale so that a better life is being rated against a higher standard. Tell them apart by publishing the MDE (which settles the first) and by running a then-test (which settles the second). Credit a third: the programme moved affect and the instrument measured evaluation.

8. Why does the chapter insist the well-being series be attached to nobody's pay?

Because a well-being item is unusually easy to move without moving anything real — the respondent controls the number entirely. Once it is a target it measures the respondent's relationship with the target. Strathern's version of Goodhart's law applies with unusual force here. The stronger answer notes the consequence: publish it widely, act on it, and keep it out of the compensation formula — those are compatible.

Two that require the arithmetic to be done.

9. A firm of two thousand people wants to detect a 0.20-point change in life satisfaction. σ = 1.9, α = 0.05 two-sided, power 0.80, teams of 12 with ρ_icc = 0.05. Compute the two-arm cross-sectional requirement, then the panel requirement at ρ = 0.60. Which design can this firm actually run?

(z_α/2 + z_β)² = (1.959964 + 0.841621)² = 7.848879. Cross-section: n per arm = 2 × 7.848879 × 1.9² / 0.20² = 1,417. Design effect 1 + 11 × 0.05 = 1.55, so 2,196 per arm, 4,392 in total — more than the firm has. Panel: n pairs = 7.848879 × 2 × 1.9² × (1 − 0.60) / 0.20² = 567, × 1.55 = 879 people measured twice. The firm can run the panel and cannot run the cross-section. The reduction factor is (1 − ρ)/2 = 0.20 — a fifth of the people, same power. Credit any answer reaching 4,392 and 879, and full marks for naming the reason the panel is also the better instrument: within-person change differences out each respondent's private spacing of the scale.

10. That firm has 1,200 people, of whom 65 percent complete both waves. Compute the minimum detectable effect, compare it to a documented mode-of-administration effect of 0.20 points, and say what quadrupling the sample would buy.

Usable pairs = 1,200 × 0.65 = 780. MDE = √( 7.848879 × 2 × 1.9² × 0.40 × 1.55 / 780 ) = 0.21 points. Against a 0.20-point mode effect the ratio is 1.06 — the instrument can just barely resolve something the size of changing survey vendor. Quadrupling to 3,120 pairs halves the MDE to 0.11 points and raises the annual cost from $24,569 to $98,276, an extra $73,707 a year. The mode effect afterwards is 0.20 points; the change in it is 0.00. Full marks require the conclusion, not just the numbers: power buys precision and never buys unbiasedness. Bias is bought with protocol — fixed mode, month, wording and order, with a parallel run on any change — and a candidate who proposes spending the $73,707 has answered the arithmetic and missed the chapter.


REFLECTION — eight questions, for one person and a pen

These are not for a room. Write the answers by hand if you can; the slowness is the point.

  1. If someone asked you right now, on a scale of nought to ten, how satisfied are you with your life nowadays — what number would you give, and what did you compare yourself to in order to produce it? Write down the comparison, not the number.
  1. Think of a year you rated highly at the time and would rate lower now. What moved: the year, or the ruler?
  1. Where in your own life have you raised your standard and then experienced the improvement as no improvement at all? What did that cost you?
  1. What is something true about how you are doing that no survey you have ever been given could have detected? Is that a limit of surveys, or of the ones you have been given?
  1. Recall being asked a question by an organisation that plainly wanted a particular answer. What did you give them, and what did you actually think?
  1. Who in your life would you trust to ask you how you are and record the answer accurately? What is it about them that makes that true — and could an instrument have any of it?
  1. When you have described your own wellbeing to someone else, whose scale were you using? Name the person or the group you were implicitly rating yourself against.
  1. If you could have one honest number about your own life, measured the same way every year for twenty years, what would you want it to be a number about?

ESSAY PROMPTS — five

Each is arguable from more than one side. Each requires at least one source the chapter cites and at least one it does not.

1. Do the league tables survive? Bond and Lang (2019) argue that the ordering of group means in happiness research is frequently not robust to permissible monotone transformations of the reporting function; Kaiser and Vendrik (2020) reply that the reversing transformations are implausible. Take a position on whether cross-national well-being rankings should continue to be published as rankings. Engage both directly, and one source on ordinal-data methods or on the public use of rankings that the chapter does not cite.

2. The vignette that nobody runs. King and colleagues (2004) provided a working solution to cross-cultural comparability twenty years ago and almost no large well-being survey has adopted it. Argue either that this is an unforced failure of the field — and say what would change if the Gallup World Poll carried three vignettes — or that the costs, in respondent burden and in the assumptions vignettes themselves require, genuinely outweigh the gain. Use King et al. and Angelini et al. (2014), and one critique of the vignette method that the chapter does not cite.

3. Adaptation, and whether a flat line is good news. Oswald and Powdthavee (2008) find substantial but partial adaptation to disability. Response-shift theory says people recalibrate. Argue whether adaptation makes subjective well-being a poor measure of how a life is going — because it washes out real change — or a better one, because it tells you what is actually being lived rather than what an observer thinks ought to matter. Engage Sprangers and Schwartz (1999), and one source on adaptation, hedonic treadmill or capability theory that the chapter does not cite.

4. Pricing a point. HM Treasury values one WELLBY at £13,000. Write the case that putting a monetary value on a point of life satisfaction is what finally made well-being decision-relevant, and then the case that it subordinates flourishing to the accounting frame it was meant to correct. Conclude with which you find more persuasive and why. Use the Green Book guidance and Kahneman and Deaton (2010), and one source on monetary valuation of non-market goods that the chapter does not cite.

5. The measurement that becomes a target. The chapter argues for publishing the series widely and attaching it to nobody's pay. Argue the counter-case: that a measure with no consequences attached is a measure nobody acts on, and that the risk of gaming is a price worth paying for the risk of being ignored. Use Strathern (1997) or the chapter's governance material, and one source on performance measurement, gaming or incentive design that the chapter does not cite.