Haute Lumière
Commerce · V.10 · MMXXVI · daylight
Three instruments: a ten-point quiz, eight reflection questions, five essay prompts. The quiz checks comprehension rather than recall. The reflections are private and first-person. The essays are arguable from more than one side.
Four on recall.
1. Name the four ONS personal well-being items and say which construct each one measures.
Life satisfaction (evaluative), worthwhileness (eudaimonic), happiness yesterday and anxiety yesterday (affective). All four on a 0–10 scale. One mark for the four items, one for the construct split — the point of the split is that evaluation and affect move apart, which a single-item survey cannot see.
2. What is the difference between a Likert item and a Likert scale, and why does it matter for well-being measurement?
A Likert item is one question with ordered options; a Likert scale is the sum of several. Likert's 1932 instrument was summated, and it is the summing that earns any claim to interval-like behaviour. A single-item life-satisfaction measure inherits none of that protection — it is an ordinal item being used cardinally.
3. State Kish's two components of the design effect and what each one corrects for.
DEFF_c = 1 + (m − 1)·ρ_iccfor clustering — people interviewed near each other resemble each other.DEFF_w = 1 + CV²(w)for unequal design weights. Total DEFF is their product. Credit an answer that notes a census removes the weighting term entirely, which is a free 25 percent.
4. Name the three mechanisms of response shift, and the instrument that detects the first of them.
Recalibration, reprioritisation, reconceptualisation (Sprangers and Schwartz, 1999). The then-test — a retrospective item at wave two asking about wave one on today's scale — detects recalibration.
Four on application.
5. A colleague reports that your employee well-being score fell three-tenths of a point this year and proposes an intervention. Before agreeing, what four things do you ask?
What was the minimum detectable effect — is three-tenths inside or outside it? Did the mode, month, wording or item order change since last year? Is this the same people or a fresh sample — because a fresh sample cannot separate a change in lives from a change in who answered? And is there a then-test, so we can tell a fall in lives from a rise in standards? Full marks require the mode question; it is the one that most often explains the whole movement.
6. Your board wants to benchmark your well-being score against an industry average published by a different vendor. What do you say?
That the comparison is a comparison of levels across populations whose reporting functions were never observed, which is the class of claim the ordinal problem destroys — and that the two instruments almost certainly differ in mode, wording, month and item order, any one of which is worth more than the gap being discussed. The constructive half of the answer: the defensible comparison is us against our own past on a frozen instrument, and if the board genuinely needs a between-group comparison, it needs anchoring vignettes.
7. A programme genuinely improved people's lives and the series shows a flat line. Give two distinct explanations and say how you would tell them apart.
Either the design was underpowered — the true effect is smaller than the MDE — or response shift recalibrated the scale so that a better life is being rated against a higher standard. Tell them apart by publishing the MDE (which settles the first) and by running a then-test (which settles the second). Credit a third: the programme moved affect and the instrument measured evaluation.
8. Why does the chapter insist the well-being series be attached to nobody's pay?
Because a well-being item is unusually easy to move without moving anything real — the respondent controls the number entirely. Once it is a target it measures the respondent's relationship with the target. Strathern's version of Goodhart's law applies with unusual force here. The stronger answer notes the consequence: publish it widely, act on it, and keep it out of the compensation formula — those are compatible.
Two that require the arithmetic to be done.
9. A firm of two thousand people wants to detect a 0.20-point change in life satisfaction. σ = 1.9, α = 0.05 two-sided, power 0.80, teams of 12 with ρ_icc = 0.05. Compute the two-arm cross-sectional requirement, then the panel requirement at ρ = 0.60. Which design can this firm actually run?
(z_α/2 + z_β)² = (1.959964 + 0.841621)² = 7.848879. Cross-section:n per arm = 2 × 7.848879 × 1.9² / 0.20² = 1,417. Design effect1 + 11 × 0.05 = 1.55, so 2,196 per arm, 4,392 in total — more than the firm has. Panel:n pairs = 7.848879 × 2 × 1.9² × (1 − 0.60) / 0.20² = 567, × 1.55 = 879 people measured twice. The firm can run the panel and cannot run the cross-section. The reduction factor is (1 − ρ)/2 = 0.20 — a fifth of the people, same power. Credit any answer reaching 4,392 and 879, and full marks for naming the reason the panel is also the better instrument: within-person change differences out each respondent's private spacing of the scale.
10. That firm has 1,200 people, of whom 65 percent complete both waves. Compute the minimum detectable effect, compare it to a documented mode-of-administration effect of 0.20 points, and say what quadrupling the sample would buy.
Usable pairs = 1,200 × 0.65 = 780.
MDE = √( 7.848879 × 2 × 1.9² × 0.40 × 1.55 / 780 ) = 0.21 points.Against a 0.20-point mode effect the ratio is 1.06 — the instrument can just barely resolve something the size of changing survey vendor. Quadrupling to 3,120 pairs halves the MDE to 0.11 points and raises the annual cost from $24,569 to $98,276, an extra $73,707 a year. The mode effect afterwards is 0.20 points; the change in it is 0.00. Full marks require the conclusion, not just the numbers: power buys precision and never buys unbiasedness. Bias is bought with protocol — fixed mode, month, wording and order, with a parallel run on any change — and a candidate who proposes spending the $73,707 has answered the arithmetic and missed the chapter.
These are not for a room. Write the answers by hand if you can; the slowness is the point.
Each is arguable from more than one side. Each requires at least one source the chapter cites and at least one it does not.
1. Do the league tables survive? Bond and Lang (2019) argue that the ordering of group means in happiness research is frequently not robust to permissible monotone transformations of the reporting function; Kaiser and Vendrik (2020) reply that the reversing transformations are implausible. Take a position on whether cross-national well-being rankings should continue to be published as rankings. Engage both directly, and one source on ordinal-data methods or on the public use of rankings that the chapter does not cite.
2. The vignette that nobody runs. King and colleagues (2004) provided a working solution to cross-cultural comparability twenty years ago and almost no large well-being survey has adopted it. Argue either that this is an unforced failure of the field — and say what would change if the Gallup World Poll carried three vignettes — or that the costs, in respondent burden and in the assumptions vignettes themselves require, genuinely outweigh the gain. Use King et al. and Angelini et al. (2014), and one critique of the vignette method that the chapter does not cite.
3. Adaptation, and whether a flat line is good news. Oswald and Powdthavee (2008) find substantial but partial adaptation to disability. Response-shift theory says people recalibrate. Argue whether adaptation makes subjective well-being a poor measure of how a life is going — because it washes out real change — or a better one, because it tells you what is actually being lived rather than what an observer thinks ought to matter. Engage Sprangers and Schwartz (1999), and one source on adaptation, hedonic treadmill or capability theory that the chapter does not cite.
4. Pricing a point. HM Treasury values one WELLBY at £13,000. Write the case that putting a monetary value on a point of life satisfaction is what finally made well-being decision-relevant, and then the case that it subordinates flourishing to the accounting frame it was meant to correct. Conclude with which you find more persuasive and why. Use the Green Book guidance and Kahneman and Deaton (2010), and one source on monetary valuation of non-market goods that the chapter does not cite.
5. The measurement that becomes a target. The chapter argues for publishing the series widely and attaching it to nobody's pay. Argue the counter-case: that a measure with no consequences attached is a measure nobody acts on, and that the risk of gaming is a price worth paying for the risk of being ignored. Use Strathern (1997) or the chapter's governance material, and one source on performance measurement, gaming or incentive design that the chapter does not cite.