Haute Lumière

Commerce · V.10 · MMXXVI · daylight

La Bourse  /  Volume V  /  Nº V.10  /  Ten concept briefs

A watercolour portrait of a woman in a green shirt and camel jacket among leaves and sunflowers.
Plate V.10 · Ten concept briefsEleven Rungs.Every measurement of a life begins with a person translating a life into a number. The instrument is not the questionnaire. The instrument is that translation, and it is done by the respondent, alone, in about four seconds.

TEN CONCEPT BRIEFS · Chapter V.10 — The Thriving Survey

One page each. A reader who reads only these ten pages has the chapter.


BRIEF 1 — The Evaluative Item and the Cantril Ladder

The idea. The workhorse of global well-being measurement is one question with eleven answers.

Hadley Cantril's Self-Anchoring Striving Scale asks the respondent to imagine a ladder numbered nought at the bottom and ten at the top, define the top as the best possible life for you, and say which step they are on. The definition of the top is the respondent's own — that is the "self-anchoring" — which is the scale's great strength and its central difficulty in the same sentence.

The Gallup World Poll carries it to roughly 1,000 adults in each of 140 to 160 countries a year, covering about 98 percent of the world's adult population. The ONS runs a close relative — overall, how satisfied are you with your life nowadays? — to about 150,000 UK adults a year.

The distinction that matters. An evaluative item asks you to judge your life as a whole. An affective item asks how you felt yesterday. They are different constructs and they move apart: Kahneman and Deaton showed in 2010 that income tracks evaluation far more closely than it tracks daily affect. A survey carrying only one of them cannot see the divergence, and the divergence is usually the finding.

Worked example. The ONS asks four items and not one, precisely for this reason: life satisfaction, worthwhileness, happiness yesterday, anxiety yesterday. A year in which satisfaction holds steady and anxiety rises is a real event. A survey with one item reports "no change."

  evaluative   → life as a whole, judged        (satisfaction, ladder)
  eudaimonic   → life as meaningful             (worthwhile)
  affective    → yesterday, felt                (happy, anxious)

Why it matters. Almost every disappointing well-being programme result comes from a single evaluative item asked of an intervention that was only ever going to move affect. Ask the item that can move.

You already know this because you have had a week that was pleasant day by day and hollow in the round, and a week that was punishing day by day and one of the best of your life. You already know these are two measurements.


BRIEF 2 — What a Likert Scale Is, and Is Not

The idea. Rensis Likert's 1932 instrument was a summated rating scale — several items, each scored, added up. The addition was the point.

A single question with ordered response options is a Likert item. A Likert scale is the sum of several of them. It is the summing that gives the result any claim to behave like a measurement with equal intervals, and single-item well-being measures inherit none of that protection.

What the scale gives you. Order. If Amina answers eight and Ben answers six, Amina is reporting more than Ben. That is real information and it is all the instrument guarantees.

What it does not give you. Distance. A 0–10 item has 11 categories and 10 intervals, and the instrument establishes the size of none of them. The respondent decides what the gap between six and seven means, privately, in about four seconds, and nothing records the decision.

Worked example. Three people move from six to seven. For one it was a new job; for one it was a good week; for one it was noticing they had been under-rating themselves. All three produce +1. The mean rises by the same amount in each case, and the mean cannot tell you which happened.

Why it matters. Because the moment you take a mean, run a regression or build a league table, you have assumed those ten intervals are equal. That assumption is usually harmless and occasionally decisive, and Brief 3 is about telling the two apart.

You already know this because you have rated a restaurant four out of five and known perfectly well that your four and your friend's four were not the same quantity — and that averaging them was still the best available thing to do.


BRIEF 3 — The Ordinal–Cardinal Problem

The idea. If the spacing of a scale is unknown, then any transformation that preserves the order of the rungs is as legitimate as the raw numbers — and some of those transformations reverse the answer.

Worked example, done in full. Two populations on the same 0–10 item.

  population A    20% at 3,  50% at 5,  30% at 9      mean  5.80
  population B    10% at 3,  75% at 5,  15% at 9      mean  5.40
  A - B                                                     +0.40

A is ahead. Now suppose the true felt distance from rung three to rung five is large and from five to nine is small. Map 3 → 0.0, 5 → 6.0, 9 → 7.0. The order of the rungs is untouched.

  A transformed   5.10        B transformed   5.55        A - B   -0.45

Same data. Same ordering. Opposite league table.

The literature, in three lines. Bond and Lang (2019) demonstrated this on real happiness data and found many published orderings not robust. Kaiser and Vendrik (2020) replied that the reversing transformations are often implausible. Ferrer-i-Carbonell and Frijters (2004) established the reconciliation everyone actually uses: the ratios between estimated coefficients are stable under the transformation; the levels are not.

The operating rule.

ClaimSurvives?
Unemployment costs about three times what a long commute costsYes — a ratio, one sample, one instrument
Our engineers score 0.4 above our nursesNo — a comparison of levels across populations
Our score rose 0.4 this year, same people, same instrumentYes — within-person change
Country X is happier than country YNo — and this is the most-quoted claim in the field

You already know this because you have seen two teams both give their manager "7 out of 10" and known, without being able to prove it, that one of those sevens was generous and the other was grudging.


BRIEF 4 — Response Shift and the Then-Test

The idea. People change the scale while you are using it.

Sprangers and Schwartz named the three mechanisms in 1999. Recalibration — the same life now gets a different number, because the internal standard moved. Reprioritisation — the respondent weights the domains differently than before. Reconceptualisation — the respondent's idea of what a good life is has changed.

All three break the one assumption a time series depends on: that the ruler did not move between readings.

Worked example. A respondent says 6.80 at wave one. A year later they say 6.90 about now — an observed change of +0.10, which looks like nothing. Then you also ask, at wave two, what wave-one life felt like on today's scale. They say 6.10.

  observed change     T2 - T1              =  +0.10
  true change         T2 - retrospective   =  +0.80
  recalibration       T1 - retrospective   =  +0.70

  share of the true change the raw series hides       87.5 %

The programme worked. The series reported that it did not.

The fix: the then-test. One retrospective item at wave two about wave one. Howard and colleagues built it in 1979 for exactly this. It costs one question, takes about twenty seconds, and is almost never asked — which is why so many genuinely good interventions show flat.

The direction of the bias is the cruel part. Adaptation runs toward the mean, so a programme that raises people's standards while raising their lives reports smaller than it achieved, and a programme that lowers standards reports larger. Oswald and Powdthavee found substantial — though only partial — adaptation after the onset of disability, which is the same mechanism in a much harder setting.

You already know this because you can no longer feel the salary rise that delighted you three years ago, and if someone asked you to rate that year now you would rate it lower than you did at the time.


BRIEF 5 — Anchoring Vignettes and Scale Norming

The idea. You cannot see where somebody puts the cut-points of a scale — so give them something whose true value you already know, and watch where they put that.

Gary King, Christopher Murray, Joshua Salomon and Ajay Tandon built the method in

  1. Alongside the self-rating, the respondent rates several short descriptions

of third persons on the same scale. Because every respondent rates the same described people, the differences in those ratings estimate the differences in how they use the scale.

Worked example. Two respondents both rate themselves 7. Both are asked to rate "Yusuf, who has secure work he finds dull, good health and two close friends."

  respondent one rates Yusuf  5   →  self 7 is two rungs above Yusuf
  respondent two rates Yusuf  7   →  self 7 is level with Yusuf

Identical self-ratings. Different underlying positions. The vignette recovered the difference, and nothing else could have.

The documented payoff. King and colleagues' own case reversed a cross-national ordering of political efficacy once thresholds were corrected. Angelini and colleagues (2014) applied the machinery to European life satisfaction and found reporting-scale differences large enough to move the gaps between countries materially. Harzing's 26-country study of response styles shows the same thing from the other direction: how willing a population is to use the ends of a scale differs systematically by country.

Why it matters. This is the only established tool that makes a between-group comparison of levels defensible. Essentially no large well-being survey carries it, which is the single largest gap between what the field publishes and what the field knows how to do.

The cost, computed. Three vignettes add about 1.5 minutes per response. For a panel of 879 people across two waves that is $2,074.44 of respondent time. That is the price of being allowed to compare two groups.

You already know this because when a friend says a film was "a seven" you immediately ask what they gave something you have both seen. You are running a vignette.


BRIEF 6 — The Design Effect

The idea. A thousand people chosen by a real survey are worth fewer than a thousand people chosen at random, and the design effect is the exchange rate.

Real samples cluster — an interviewer visits a sampling point and interviews several people who live near each other, and neighbours resemble each other. Real samples are weighted — some respondents stand for more people than others, and unequal weights add variance. Leslie Kish gave both in 1965:

  clustering    DEFF_c  =  1 + (m - 1) · ρ_icc
  weighting     DEFF_w  =  1 + CV²(w)
  total         DEFF    =  DEFF_c × DEFF_w

Worked example, national. Ten interviews per sampling point, intra-cluster correlation 0.05, weight CV 0.50:

  clustering    1 + 9 × 0.05   =  1.45
  weighting     1 + 0.50²      =  1.25
  total                           1.8125

You need 81 percent more people than a textbook calculation says, for the same precision. A design effect ignored is a study that was underpowered from the day it was designed and will report a null.

Worked example, a firm. Teams of twelve, same intra-cluster correlation:

  1 + 11 × 0.05  =  1.55

Note that the firm's design effect is higher than the national clustering term, because teams are larger than sampling points — and then lower overall, because a census needs no design weights at all. Running a census removes DEFF_w entirely, which is a free 25 percent that most internal surveys throw away by sampling when they did not need to.

Why it matters. It is the most commonly omitted term in the most commonly performed calculation in this field.

You already know this because you know that asking ten people from the same team is not the same as asking ten people from ten teams, and you have never needed a formula to know it.


BRIEF 7 — Minimum Detectable Effect

The idea. Before you ask what a study found, ask what it was capable of finding.

        n per arm  =  2 (z_α/2 + z_β)² σ² / δ²

        MDE        =  δ such that the above equals the n you have

With α = 0.05 two-sided and 80 percent power, z_α/2 = 1.959964 and z_β = 0.841621, so (z_α/2 + z_β)² = 7.848879. The standard deviation of a 0–10 life-satisfaction item is about 1.9 points.

Worked example, national. To detect 0.10 points:

  n per arm, simple random                    5,667
  × design effect 1.8125                     10,272
  both arms                                  20,543
  issued cases at a 30% response rate        68,475

Worked example, a firm. A firm of 1,200 people, 65 percent completing both waves of a panel — 780 usable pairs, design effect 1.55, ρ = 0.60:

  MDE  =  0.21 points

Why it matters, and this is the whole use of the concept. A null result from an underpowered study is not evidence that nothing happened. It is evidence of nothing at all. Publishing the MDE alongside the finding converts "we found no effect" into the honest sentence: "we would have detected 0.21 points and did not; effects smaller than that are invisible to this instrument."

The line to put in every methods note. This release could have detected a change of X points. It did not look for anything smaller.

You already know this because you would not weigh a letter on a bathroom scale and then report that it weighs nothing.


BRIEF 8 — The Panel Discount

The idea. The cheapest statistical power available is a second measurement of somebody you have already measured.

Comparing two groups, the variance you fight is 2σ². Comparing the same people twice, it is 2σ²(1 − ρ), where ρ is the test–retest correlation. Krueger and Schkade put ρ for single-item life satisfaction at roughly 0.50 to 0.70; take 0.60.

  reduction factor  =  (1 - ρ) / 2  =  0.20

A fifth of the people, for the same answer.

Worked example. A firm wanting to detect 0.20 points, σ = 1.9, teams of twelve so DEFF = 1.55:

  two arms, measured once                    4,392 people
  one panel, measured twice                    879 people

879 is a number a mid-sized firm has. 4,392 is not. The organisation that concluded it was too small to measure well-being honestly had simply chosen the wrong design, and the correction is not a compromise — the panel is the better instrument, because within-person change is exactly the comparison that survives the ordinal problem of Brief 3. Each respondent's private spacing of the scale differences out against itself.

And it is cheaper twice over. Recontacting a known respondent costs less than recruiting an unknown one, and the second wave carries the then-test of Brief 4, which a cross-section structurally cannot.

The cost, computed. 879 people, 9.5 minutes a year, at the US Bureau of Labor Statistics loaded compensation figure of $47.20 an hour, plus $18,000 of platform and analyst time:

  annual cost of the series        $24,569.06
  per respondent per year              $27.95

You already know this because you can tell whether you are doing better than last year far more reliably than you can tell whether you are doing better than the person at the next desk.


BRIEF 9 — Mode Effects and the Break in a Series

The idea. How you ask changes the answer, by about as much as a real event does.

Dolan and Kavetsos (2016) put the effect of moving between interviewer-administered and self-completed collection at roughly 0.10 to 0.30 points on a 0–10 well-being item. People report higher to a person than to a screen.

Documented case one — the General Social Survey. The GSS has asked Americans the same happiness question since 1972.

  2018, in person       31 % "very happy"     13 % "not too happy"
  2021, mainly web      19 % "very happy"     24 % "not too happy"
  fall                  12 points  =  38.7 % of the 2018 level
  response rate         59.5 %  →  17.4 %     (29.2 % of the 2018 rate)

The mode changed in the same step. NORC published the change as a source of non-comparability. Nobody has an uncontaminated estimate of how much of that fall was American lives.

Documented case two — the ONS Annual Population Survey. Collection moved to telephone-only in March 2020, and ONS published the mode change as a break in comparability in its own personal well-being releases. In October 2023 it suspended Labour Force Survey estimates outright as response rates fell.

Documented case three — the interviewer in the room. Conti and Pudney (2011) showed that the much-replicated finding that women report higher job satisfaction than men is substantially larger under face-to-face collection than under self-completion. The result is partly about the room.

The arithmetic that follows. A national design powered to detect 0.10 points, exposed to a 0.20-point mode effect:

  mode effect / detectable effect  =  2.00 ×

The housekeeping is twice the signal.

The fix is a protocol, not a sample. One mode, for ever; and if it must change, a parallel run of both instruments on one wave with the bridge coefficient published. A series with a visible, dated, bridged seam is honest. A series with no seams has either never been improved or has been smoothed.

You already know this because you answer "how are you?" differently to your doctor, your colleague and a form.


BRIEF 10 — The WELLBY, and the Price of a Point

The idea. A government has put a number on one point of life satisfaction, and that makes well-being tradeable against money in a formal appraisal.

HM Treasury's Green Book supplementary guidance on wellbeing (2021) values one WELLBY — a one-point change on a 0–10 life-satisfaction scale, for one person, for one year — at £13,000 in 2019 prices, for use in public policy appraisal.

Worked example. An instrument powered to detect 0.20 points is therefore powered to detect a change worth, in Green Book terms:

  0.20 × £13,000  =  £2,600  per person per year

against a measurement cost of $27.95 per person per year — roughly 93 times the cost of finding out, taking the two currencies at parity for the comparison only.

What the WELLBY legitimately supports. Ranking public interventions against each other on a common scale. Making the implicit trade-off explicit, so that a decision which values a point of life satisfaction at £200 has to say so out loud.

What it does not support, and this matters. It is not a booking entry. A firm may not recognise £2,600 a head because its survey moved. It is an appraisal value, built for comparing public options ex ante, and it carries all the uncertainty of the studies behind it.

Why it matters anyway. Before the WELLBY, "we should care about well-being" and "we should build the bypass" were arguments in different currencies and the bypass always won, because it had a number. The WELLBY does not settle the argument. It makes the argument possible. That is what a unit of account is for, and it is the same move that made carbon comparable across sectors.

You already know this because you have already traded a pay rise against a shorter commute, and when you did it you were pricing a point of life satisfaction — you simply never wrote the price down.