Haute Lumière

Commerce · V.10 · MMXXVI · daylight

La Bourse  /  Volume V  /  Nº V.10

The Thriving Survey

Volume V — Labour, Value, Flourishing


THE PLATE

A watercolour portrait of a woman in a green shirt and camel jacket among leaves and sunflowers.
Plate V.10Eleven Rungs.Every measurement of a life begins with a person translating a life into a number. The instrument is not the questionnaire. The instrument is that translation, and it is done by the respondent, alone, in about four seconds.

THE LETTER

You want to know whether the people in your organisation are flourishing, and you want to know it the way you know your revenue: as a number you can put in front of somebody who does not share your intuitions.

That is a serious ambition and this chapter takes it seriously. It is a methods chapter. It will give you the instruments that already exist and what they actually cost, the measurement theory underneath them and where that theory gives way, the arithmetic that tells you how many people you need and what that comes to per head, and the honest account of what such an instrument can and cannot be made to support.

The short version, so you can decide now whether to read on. Measuring human flourishing at scale is cheaper than you expect and less conclusive than you want. The cheapness is genuine: the annual cost of an honest, properly powered well-being series inside a firm comes out at $27.95 per respondent per year, computed below from published survey design and the cost of the respondent's own time. The inconclusiveness is equally genuine: a large well-being survey reliably produces a number that moves for reasons that have nothing to do with the lives it is measuring, and we will walk through two national cases where exactly that happened and was traced.

Neither of those facts cancels the other. Together they say something quite precise about what to build: an instrument whose protocol is expensive and whose sample is cheap. That is the opposite of how almost every organisation approaches this, and it is the argument of the chapter.

One more thing, said plainly because it is what makes the rest usable. The people who built these instruments — Cantril, Likert, Kish, Diener, the OECD statisticians who wrote the 2013 guidelines, the ONS team who put four questions into a national survey and defended them for a decade — did careful, patient, honest work, and most of what we know about the limits of these measures we know because they went and found the limits themselves and published them. This chapter is downstream of that. It is not a critique of well-being measurement. It is the operating manual that the measurement's own literature wrote.

— The Editors


DISCOVERY

What is already working

Four instruments are running at scale right now, and each of them is good at a different thing. Knowing which is which is most of the skill.

The Gallup World Poll. Running since 2005, roughly 1,000 adults per country per year — larger frames of up to 4,000 in China, India and Russia — across 140 to 160 countries, together covering about 98 percent of the world's adult population. It carries the Cantril Self-Anchoring Striving Scale, which is the ladder: imagine a ladder with steps numbered from nought at the bottom to ten at the top; the top is the best possible life for you; on which step do you personally feel you stand? Telephone interviewing where landline and mobile coverage is high, face-to-face stratified multi-stage cluster samples elsewhere. The World Happiness Report has been built on it annually since 2012.

What it is genuinely good at: a consistent frame across a very large number of countries, held steady for two decades. Nobody else has that. The league table it produces is the most-quoted and least-defensible thing about it, and the people who run it know that better than their critics do.

The Global Flourishing Study. Wave one, collected between 2022 and 2024 and released in 2025: roughly 203,000 participants across 22 countries, with four further annual waves planned. Built by Baylor's Institute for Studies of Religion and Harvard's Human Flourishing Program with Gallup as the field house and the Center for Open Science as the data steward. Six domains — happiness and life satisfaction; mental and physical health; meaning and purpose; character and virtue; close social relationships; financial and material stability — each carried by two items on a nought-to-ten scale, twelve items in all.

What is working here is not the size. It is the preregistration and the open data. The analyses were lodged before the data came back and the microdata was released with the first papers. That is a well-being survey operating at the evidentiary standard of a clinical trial, and it is new.

The OECD apparatus. The Guidelines on Measuring Subjective Well-being (2013) are the reason national statistical offices ask the same questions in the same words: a core life-evaluation item, plus optional modules for affect and for eudaimonia, with model questionnaires a statistician can lift directly. The Better Life Index (2011) carries 11 topics and 24 indicators and lets the reader set the weights themselves, which was a quietly radical decision — it refuses to tell you what a good life weighs.

The ONS four questions. Introduced into the UK's Annual Population Survey in April 2011 and asked of adults ever since. Four items, each nought to ten:

Overall, how satisfied are you with your life nowadays? Overall, to what extent do you feel the things you do in your life are worthwhile? Overall, how happy did you feel yesterday? Overall, how anxious did you feel yesterday?

The Annual Population Survey issues on the order of 320,000 cases a year and the well-being items reach roughly 150,000 adult respondents. That is enough to publish at local-authority level, which is the thing that makes a council able to act. And in 2021 HM Treasury's Green Book supplementary guidance put a price on the output: one WELLBY — one point of life satisfaction, for one person, for one year — valued at £13,000 in 2019 prices, for use in public appraisal. Whatever you think of that number, a government that will trade a point of life satisfaction against a pound is a government that has taken the measurement seriously enough to be held to it.

And two more worth having in the room. Bhutan's Gross National Happiness Index — 9 domains, 33 indicators, 148 variables, aggregated by the Alkire–Foster method, surveyed in 2010, 2015 and 2022 — because it is the only national index built to identify who is not yet sufficient rather than to produce an average. And New Zealand's Living Standards Framework, which put well-being indicators into the 2019 Budget process itself, where allocation happens.

Here is the pattern across all six, and it is the appreciative finding of this movement: every one of them is strongest where it is most boring. The consistent frame held for two decades. The identical wording across offices. The preregistered analysis. The published methodology note. The local-authority granularity. Not one of them is strongest at its headline. The league table is the weakest output of the best instrument in the world, and the methods note is the strongest.


THE ARITHMETIC

What works, what does not, and where the line sits

What a Likert scale is, and what it is not

Rensis Likert's 1932 instrument was a summated rating scale: several items, each scored, added together. The addition was load-bearing. A single item is a Likert item; a Likert scale is the sum, and it is the sum that earns any claim to interval-like behaviour.

A nought-to-ten life-satisfaction item has 11 categories and therefore 10 intervals between them. Cardinal analysis — computing a mean, running a regression, comparing countries — assumes all ten intervals are the same size. The instrument establishes none of them. The respondent decides what the gap between six and seven is, in about four seconds, and nobody records that decision.

This is not a reason to stop. It is a reason to know exactly which claims survive it.

The ordinal–cardinal problem, worked

Take two populations answering the same item. Population A: 20 percent at rung three, 50 percent at rung five, 30 percent at rung nine. Population B: 10 percent at three, 75 percent at five, 15 percent at nine.

  raw means          A = 5.80        B = 5.40        A - B = +0.40

A is ahead. Now apply a transformation that preserves the order of the rungs — which is all an ordinal scale entitles you to assume — but not their spacing. Suppose the true felt distance from three to five is large, and from five to nine is small: map rung three to 0.0, rung five to 6.0, rung nine to 7.0.

  transformed        A = 5.10        B = 5.55        A - B = -0.45

Same data. Same ordering of the rungs. Opposite league table. Tim Bond and Kevin Lang did this properly on real data in 2019 and showed that the ordering of group means in published happiness research is frequently not robust to permissible monotone transformations of the reporting function. Christian Kaiser and Maarten Vendrik have argued back, with some force, that the transformations which reverse real orderings are often implausible ones — but the argument is about plausibility, not about whether the vulnerability exists.

The reconciliation is the teachable result, and it comes from Ada Ferrer-i-Carbonell and Paul Frijters in 2004: treating the scale as ordinal or cardinal barely changes the estimated ratios between coefficients, and changes the comparison of levels a great deal. So:

Almost every well-being claim that has ever embarrassed anybody is of the second kind.

Response shift, and the then-test

People rescale. Sprangers and Schwartz named the three mechanisms in 1999 — recalibration, reprioritisation, reconceptualisation — building on Howard's 1979 work on pretest–posttest self-report. The consequence for a series is severe, and it is arithmetic.

A respondent says 6.80 at wave one. A year later they say 6.90 about now. Observed change: +0.10. Then you also ask them, at wave two, to rate what wave-one life felt like on today's scale. They say 6.10.

  observed change     T2 - T1                =  +0.10
  true change         T2 - retrospective     =  +0.80
  recalibration       T1 - retrospective     =  +0.70

The programme improved lives by eight-tenths of a point and the series reports one-tenth. The raw series hides 87.5 percent of the real change, and it does so in the direction that makes good work look like nothing.

The fix is the then-test: a retrospective item, asked at wave two, about wave one. It costs one question. It is almost never asked.

Cross-cultural comparability, and where it actually fails

Gary King, Christopher Murray, Joshua Salomon and Ajay Tandon built the standard tool in 2004: anchoring vignettes. You ask the respondent to rate several described third persons on the same scale they rate themselves, which lets you estimate where their thresholds sit. Their own worked case reversed an ordering: on raw self-reports, one national sample rated its political efficacy above another's; after correcting for where each placed the scale's cut-points, the ordering flipped. Viola Angelini and colleagues ran the same machinery on European life satisfaction in 2014 and found reporting-scale differences large enough to move the gaps between countries substantially.

Anne-Wil Harzing's 26-country study of response styles adds the other half: acquiescence and extreme-responding differ systematically by country, so some of what a league table measures is how willing a population is to use the ends of a scale.

So the honest statement about cross-cultural comparison is narrow and it is usable: a fixed instrument compares a population to its own past reliably, and compares two populations to each other only as well as its vignettes allow — and almost no instrument in the field carries vignettes.

How many people, and what it costs

This is the computation that decides whether an organisation can measure its own people honestly. Everything below is in lib/verify/V_10.py and reproducible.

Standard two-sample power, equal arms:

        n per arm  =  2 (z_α/2 + z_β)² σ² / δ²

  z_α/2 = 1.959964   two-sided α = 0.05
  z_β   = 0.841621   power = 0.80
  σ     = 1.9 points  SD of a 0–10 life-satisfaction item
  δ     = the effect you want to detect

For a national comparison at δ = 0.10 points:

  (z_α/2 + z_β)²                           7.848879
  n per arm, simple random                    5,667

Real national surveys are not simple random samples. Multiply by the design effect, which has two parts (Kish, 1965):

  clustering   1 + (m − 1)·ρ_icc   with m = 10, ρ_icc = 0.05   =  1.45
  weighting    1 + CV²(w)          with CV = 0.50              =  1.25
  ------------------------------------------------------------------
  design effect                                                 1.8125

  n per arm, clustered                       10,272
  both arms                                  20,543

At a 30 percent response rate — ordinary for national fieldwork now, against a European Social Survey target of 70 percent, a shortfall of 40 points — you must issue 68,475 cases to complete 20,543 interviews.

At planning rates per completed interview:

  probability web panel      $12 ×  20,543  =  $   246,516
  dual-frame telephone       $45 ×  20,543  =  $   924,435
  in-person field team       $85 ×  20,543  =  $ 1,746,155

Telephone works out at $13.50 per issued case and $9,244,350 per full point of detectable effect. Those are the real orders of magnitude of national well-being measurement, and they are why it is done by states and by Gallup and by nobody else.

Now the firm, which is the case that matters to you

Same σ. A firm would not act on a tenth of a point; set δ = 0.20. Teams of 12, so the design effect is 1 + 11 × 0.05 = 1.55.

  n per arm, simple random                    1,417
  n per arm, clustered                        2,196
  both arms                                   4,392

Most firms do not have 4,392 people. For them the cross-sectional design is not expensive. It is unavailable, and the usual conclusion — we are too small to measure this properly — follows, and is wrong.

Here is why it is wrong, and it is the cut this chapter is built around.

A person's life satisfaction is correlated with their own life satisfaction a year ago. Krueger and Schkade put the test–retest correlation of single-item life satisfaction at roughly 0.50 to 0.70 over short intervals; take ρ = 0.60. Measure the same people twice and you are testing a difference whose variance is not 2σ² but 2σ²(1 − ρ):

  variance of the change   2σ²(1 − ρ)        =   2.888
  SD of the change                           =  1.6994 points

  n pairs, simple random                          567
  n pairs, clustered                              879

879 people, measured twice, replaces 4,392 people measured once. Five times fewer, from the same power, at the same effect size, on the same scale. The reduction factor is (1 − ρ)/2 = 0.20, and it is paid for entirely by the fact that people's lives are correlated with themselves.

And it is not merely smaller. It is better: a panel identifies within-person change, which is precisely the comparison that survives the ordinal problem, because each respondent's reporting function is differenced out against itself.

The cost, in the only currency a firm actually pays — respondent time, at the US Bureau of Labor Statistics figure for total compensation of $47.20 per hour:

  instrument: 8.0 min core + 1.5 min vignettes = 9.5 min
  first year, 879 people × 2 waves = 1,758 responses
    core       14,064 min = 234.40 h × $47.20  =  $11,063.68
    vignettes   2,637 min                      =  $ 2,074.44
    ----------------------------------------------------------
    first year, respondent cost                =  $13,138.12
    per person                                 =  $    14.95

  steady state, one wave a year
    respondent time  139.18 h                  =  $ 6,569.06
    platform and analyst                       =  $18,000.00
    ----------------------------------------------------------
    ANNUAL COST OF THE SERIES                  =  $24,569.06
    COST PER RESPONDENT PER YEAR               =  $    27.95

Twenty-eight dollars a head a year. That is the number that decides whether an organisation can measure its own people honestly, and it is smaller than the coffee budget. Against the Green Book's £13,000 per WELLBY, the 0.20 points the instrument is powered to detect is worth £2,600 per person per year — about 93 times the cost of finding out. That comparison is a price check, not a booking entry: the Green Book value exists for appraising public policy, and no firm may recognise £2,600 a head in its accounts on the strength of it.

The honest negative, and it is the load-bearing part of the chapter

A large well-being survey reliably produces a number that moves for reasons that are not the thing being measured. Two documented cases, both national, both traced.

The General Social Survey, 2018 to 2021. The GSS has asked Americans the same happiness question since 1972. In 2018, collected in person: 31 percent "very happy", 13 percent "not too happy". In 2021, collected predominantly by self-administered web because of the pandemic: 19 percent very happy — a fall of 12 points, or 38.7 percent of the 2018 level, the lowest reading in the series' history — and 24 percent not too happy, a rise of 11 points, leaving 57 percent in the middle. The response rate fell in the same step, from 59.5 percent to 17.4 percent: the 2021 rate is 29.2 percent of the 2018 rate. NORC published the mode change as a source of non-comparability. Nobody has an uncontaminated estimate of how much of that 12-point fall was American lives and how much was the absence of an interviewer in the room.

The ONS Annual Population Survey, March 2020. Collection moved to telephone-only, and ONS published the change as a break in comparability in its own personal well-being releases. Paul Dolan and Georgios Kavetsos measured the size of this class of effect directly: subjective well-being reported to an interviewer runs higher than the same construct self-completed, by something on the order of 0.10 to 0.30 points on a nought-to-ten scale. Take the midpoint, 0.20. Set it against the national design above, which is powered to detect 0.10:

  mode effect / detectable effect            2.00 ×

The housekeeping is twice the size of the signal. And for the firm, the minimum detectable effect with 1,200 people on the payroll and 65 percent completing both waves — 780 usable pairs — is 0.21 points, against a mode effect of 0.20. Ratio: 1.06. A fully panelled firm of 1,200 can just barely resolve something the size of changing survey vendor.

And here is the part that does not yield to money. Quadruple the sample — 3,120 pairs — and the minimum detectable effect halves exactly, from 0.21 to 0.11 points. The annual cost goes from $24,569 to $98,276, an extra $73,707 a year. The mode effect after you have spent it: 0.20 points. The change in the mode effect: 0.00.

  four times the money  →  half the precision term
                        →  the bias term exactly where it was

Statistical power buys precision and never buys unbiasedness. Bias is bought with protocol, not with sample — fixed mode, fixed month, fixed wording, fixed position in the questionnaire, and a parallel run whenever any of those changes. There is a third case worth carrying for the same reason: Gabriella Conti and Stephen Pudney showed in 2011 that the much-replicated finding that women report higher job satisfaction than men is substantially larger when an interviewer is present than under self-completion. That famous result is partly a finding about the room.


DREAM

What becomes ordinary

In the organisation that has built this properly, the well-being series is kept the way the general ledger is kept: by protocol, by the same person, in the same month, in the same words, and any change to it is an announced event with its own paperwork.

The instrument is four items, not one, because life satisfaction, worthwhileness, yesterday's happiness and yesterday's anxiety move differently and a single number hides the divergence. Three anchoring vignettes ride on a rotating third of the sample, so the organisation knows where its own people place the scale's cut-points and can tell a real improvement from a shifted threshold. A then-test rides on another rotating third, so a year in which everyone quietly raised their standard is visible as exactly that rather than as a flat line.

The methods note is published internally alongside every release and it is one page. It states the mode, the month, the wording, the response rate, the design effect, the minimum detectable effect, and — this is the line that makes the rest trustworthy — what the release did not look at. People read the methods note. It is short, so they do.

Nobody in this organisation quotes a league table. They do not compare their score to an industry benchmark collected by a different vendor with a different instrument in a different month, because they know what that comparison is worth and they have the arithmetic to say so in one sentence. What they compare is themselves to themselves, which is the comparison the instrument can actually support, and which is also the only one that tells them whether what they did worked.

When a change is made to the instrument — and changes are made, because instruments improve — the old and the new are run side by side on one wave, the bridge coefficient is computed and published, and the series carries a visible seam at that date for ever. The seam is the honesty. A series with no seams has either never been improved or has been quietly smoothed, and the second is more common.

And the series is old. That is its most valuable property and the one nobody can buy. A five-year panel with a clean protocol is worth more than a fifty-thousand- person cross-section taken once, because it can answer the only question that matters to a person deciding what to do next: did the thing we changed change anything?


DESIGN

The instrument, built

The items. Four evaluative and affective items on nought to ten, in the ONS wording, because using somebody else's validated wording is free and inventing your own costs you every external comparison you will ever want. Add domain items only if you will act on them separately; each one costs respondent time and respondent time is the real budget.

The frame. A census invitation to the whole population, not a sample. Below about five thousand people, sampling buys you nothing and costs you the goodwill of the people you did not ask.

The design. A panel with a stable identifier, measured annually in the same month. Baseline plus one wave is the minimum viable series. Target 879 completed pairs for a 0.20-point effect at 80 percent power, or compute your own from the formula above with your own σ.

The protocol, which is the expensive part and the part that works. Five rules, and each one is there because breaking it has produced a false national finding somewhere:

  1. One mode, for ever. If you must change it, parallel-run.
  2. One month, for ever. Seasonality in affect items is real and large.
  3. One wording, for ever. Including the introduction, including the order.
  4. One position in the questionnaire. Items that precede a well-being item change it; a block of questions about workload before a life-satisfaction item is not a neutral context.
  5. Every change is announced, dated and bridged. A change log lives beside the series and is published with it.

The instruments that protect the instrument. Three anchoring vignettes on a rotating third of respondents. A then-test item at each wave after the first, on a different rotating third. Both together add 1.5 minutes and about $2,074 in the first year, and they are the difference between a series you can defend and a series you can only publish.

The governance. A named owner, a named deputy, and a change-control rule that requires both signatures plus a funded parallel run before any item is altered. This is the piece most organisations skip, and it is why most organisational well-being series are three years old at most: somebody changed vendor, the number moved, nobody could say why, and the series was quietly retired.

The sequence. Wave one is a baseline and produces no findings — say so loudly and in advance, because the pressure to find something in the baseline is what corrupts the baseline. Wave two produces the first change estimate and the first then-test. Wave three is where the series starts to be worth its protocol, and it is also the first point at which somebody will propose changing the instrument. That is the moment the governance exists for.


DESTINY

How it holds when you stop pushing

A well-being series sustains itself on exactly one property: continuity, and continuity is not a technical achievement, it is a governance one.

Three things keep it alive. It is in the standing pack, reported at the same cadence as anything else that is reported at all. It has two owners, because one owner is a hobby. And it has a funded continuity reserve, so that a bad budget year cannot kill five years of baseline — which is the failure mode, because the series always looks optional in the year nothing happened.

Now the honest part, and it is specific. Here is how this dies.

It dies when somebody changes the survey vendor for good commercial reasons and nobody runs the parallel wave, and the number moves, and from that day the series before the change and the series after it are two different series wearing one name. It dies when a good year produces a flat reading because everybody rescaled, and a leader who was promised a return concludes the instrument does not work — which is why the then-test is not optional. It dies when the minimum detectable effect is never computed, so a null result is read as evidence of no effect rather than as evidence of an underpowered design. And it dies, most often, when it becomes an incentive: the moment a manager's compensation moves with a well-being score, the score stops measuring well-being and starts measuring the manager's relationship with the score. Marilyn Strathern's version of Goodhart's law is the one to keep on the wall — when a measure becomes a target, it ceases to be a good measure — and a well-being item is unusually easy to move without moving anything real, because the respondent controls it entirely.

The rule that follows is uncomfortable and it is right: publish the series widely and attach it to nobody's pay. It is an instrument, not a target. The moment it becomes a target you have bought a very expensive way of finding out what people think you want to hear.


DELIGHT

What it feels like

There is a particular quiet in a room where the second wave has come back and somebody has just put the then-test beside the raw change. The flat line becomes a rise. Nothing about the year changed; what changed is that the room can now see what the year did.

And there is a better one after it. Someone asks a question the series was not built to answer — is it the new shift pattern or the new manager? — and the answer is there, because the panel has the same people in it both times and the comparison is within them. The instrument turns out to be more generous than its designer was. That happens with good instruments and almost never with questionnaires.

The smallest pleasure is the truest one. Four items, nine and a half minutes, once a year, twenty-eight dollars a head — and a person who has answered it three times will tell you, unprompted, that it is the only thing the organisation asks them that does not want anything from them. A question that does not want anything is rare enough to be felt. That is what people are responding to, and it is why the response rate holds.


OPERATIONALIZE THIS

At the level of finance

The instrument: a three-year funded Measurement Continuity Facility, with a parallel-run covenant in the vendor contract.

You are not buying a survey. You are building a time series, and a time series is an asset whose whole value is its continuity — which means the thing to finance is not the fieldwork, it is the protection of the fieldwork from your own budget cycle.

The structure. A ring-fenced three-year commitment, drawn annually, held outside the department that consumes it. Annual run cost $24,569.06; continuity reserve at three years $73,707.18. The reserve is not a contingency against cost overrun — the cost does not overrun, it is respondent time and a licence — it is a defence against discontinuation.

The balance-sheet treatment, honestly. An internally generated database is generally not capitalised under IAS 38, and nobody should pretend otherwise. So the run cost is expensed, and the series is registered in the data asset inventory at its replacement cost: five years of protocol is $122,845.30 of spend and five years of calendar time, and the calendar time cannot be bought at any price. State both figures in the register. The second one is what stops the series being cancelled, because it is the only line in the inventory with an asset that money cannot rebuild.

The counterparty. The survey vendor, with internal audit as the verifier of protocol adherence — not of the findings, of the protocol. Audit does not need to know what a design effect is to check that the mode, month and wording matched the last wave.

The covenant that makes it an instrument rather than a subscription. Written into the vendor contract, in these terms:

Any change to mode of administration, item wording, item order, field month or panel provider shall be preceded by a parallel run of no fewer than 1,000 respondents on both the incumbent and successor instruments, at the vendor's cost, with the bridge coefficient computed and delivered before the successor instrument is used for publication.

That parallel run costs $7,473.33 in respondent time — 30.4 percent of one year's run cost — and it is the single clause that prevents the most common and most expensive failure in the whole field: a national-scale series turning into two incomparable series without anybody noticing until years later.

The first ninety days.

DayActionArtifact
1–15Fix the four items in ONS wording; fix the month; fix the positionThe frozen instrument
16–30Compute σ, the design effect and the minimum detectable effect for your populationThe power memo
31–45Name owner and deputy; write the change-control ruleThe governance page
46–60Negotiate the parallel-run covenant into the vendor contractThe signed covenant
61–75Field wave one. Publish no findingsBaseline and methods note
76–90Register the series in the data inventory at replacement cost; fund the reserveThe continuity facility

The number that decides it. One ratio, and it belongs on the front page of the proposal:

        minimum detectable effect
   --------------------------------------  ≥  2.0
    largest documented nuisance effect
      on the same scale and instrument

For the worked firm above the ratio is 1.06 — the instrument cannot distinguish its own signal from its own housekeeping, and the honest response is not to buy more sample (which would move that ratio the wrong way in cost and the right way only slowly: four times the money takes it to 0.53) but to eliminate the nuisance by fixing the protocol, at which point the mode effect contributes nothing to the change and the ratio stops being the binding constraint.

If you cannot hold the protocol, do not buy the sample. An unprotocolled well-being series is a way of paying $27.95 per person per year to generate a number you will later have to explain.


APPRECIATIVE QUESTIONS

Twelve, for a room

Discovery — what is already working

  1. Where in this organisation do we already collect something about people's experience that has been asked the same way for more than two years — and who has been quietly protecting that consistency?
  2. Think of a time a number about people changed someone's mind here. What made that number credible to the person who changed their mind?
  3. Which question do our people answer most willingly, and what is it about it that they trust?

Dream — what becomes possible

  1. If we knew honestly, once a year, whether our people's lives were going better, what is the first decision we would make differently?
  2. Imagine a methods note that our own people read voluntarily. What would be on it?
  3. If we could answer is it the shift pattern or the manager with evidence, what else would suddenly become answerable?

Design — what we build

  1. What would we have to freeze — wording, month, mode, order — and who here has the standing to freeze it?
  2. What is the largest nuisance effect our current instrument is exposed to, and how would we find out its size?
  3. Which third of our people would we ask the vignettes of, and how would we explain to them why we are asking?

Destiny — how it holds

  1. What would have to be true for this series to still be running, unbroken, in ten years?
  2. Who is the second owner, and what would make the series genuinely theirs?
  3. If somebody proposed attaching this number to a bonus, who in this room would be the one to say no, and what would we want them to be able to point at?

WORKS CITED

Likert, R. (1932). "A Technique for the Measurement of Attitudes." Archives of Psychology, 140, 1–55.

Cantril, H. (1965). The Pattern of Human Concerns. Rutgers University Press.

Kish, L. (1965). Survey Sampling. Wiley.

Groves, R. M. (1989). Survey Errors and Survey Costs. Wiley.

Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences, 2nd edn. Lawrence Erlbaum.

Diener, E., Emmons, R. A., Larsen, R. J. and Griffin, S. (1985). "The Satisfaction with Life Scale." Journal of Personality Assessment, 49(1), 71–75.

Gallup Inc. World Poll Methodology. Successive editions.

Helliwell, J. F., Layard, R., Sachs, J. D., De Neve, J.-E., Aknin, L. B. and Wang, S. (eds). World Happiness Report. Annual, 2012–.

VanderWeele, T. J. (2017). "On the promotion of human flourishing." Proceedings of the National Academy of Sciences, 114(31), 8148–8156.

Global Flourishing Study (2025). Wave 1 Data and Codebook. Baylor University Institute for Studies of Religion, Harvard Human Flourishing Program, Gallup and the Center for Open Science.

OECD (2013). OECD Guidelines on Measuring Subjective Well-being. OECD Publishing.

OECD. How's Life? Measuring Well-being. OECD Publishing, biennial, 2011–.

Office for National Statistics. Personal Well-being in the UK. Annual statistical bulletins, 2012–, with accompanying user guidance and quality and methodology information.

HM Treasury (2021). Green Book Supplementary Guidance: Wellbeing. HM Government.

Ura, K., Alkire, S., Zangmo, T. and Wangdi, K. (2012). A Short Guide to Gross National Happiness Index. Centre for Bhutan Studies.

Alkire, S. and Foster, J. (2011). "Counting and multidimensional poverty measurement." Journal of Public Economics, 95(7–8), 476–487.

New Zealand Treasury (2019). The Wellbeing Budget 2019; and Our Living Standards Framework.

Ferrer-i-Carbonell, A. and Frijters, P. (2004). "How Important is Methodology for the Estimates of the Determinants of Happiness?" Economic Journal, 114(497), 641–659.

Bond, T. N. and Lang, K. (2019). "The Sad Truth about Happiness Scales." Journal of Political Economy, 127(4), 1629–1640.

Kaiser, C. and Vendrik, M. C. M. (2020). "How Threatening Are Transformations of Happiness Scales to Subjective Wellbeing Research?" IZA Discussion Paper No. 13905.

Oswald, A. J. (2008). "On the curvature of the reporting function from objective reality to subjective feelings." Economics Letters, 100(3), 369–372.

King, G., Murray, C. J. L., Salomon, J. A. and Tandon, A. (2004). "Enhancing the Validity and Cross-cultural Comparability of Measurement in Survey Research." American Political Science Review, 98(1), 191–207.

Angelini, V., Cavapozzi, D., Corazzini, L. and Paccagnella, O. (2014). "Do Danes and Italians Rate Life Satisfaction in the Same Way?" Oxford Bulletin of Economics and Statistics, 76(5), 643–666.

Harzing, A.-W. (2006). "Response Styles in Cross-national Survey Research: A 26-country Study." International Journal of Cross Cultural Management, 6(2), 243–266.

Uchida, Y. and Kitayama, S. (2009). "Happiness and unhappiness in East and West: themes and variations." Emotion, 9(4), 441–456.

Sprangers, M. A. G. and Schwartz, C. E. (1999). "Integrating response shift into health-related quality of life research: a theoretical model." Social Science & Medicine, 48(11), 1507–1515.

Howard, G. S., Ralph, K. M., Gulanick, N. A., Maxwell, S. E., Nance, D. W. and Gerber, S. K. (1979). "Internal invalidity in pretest–posttest self-report evaluations." Applied Psychological Measurement, 3(1), 1–23.

Oswald, A. J. and Powdthavee, N. (2008). "Does happiness adapt? A longitudinal study of disability with implications for economists and judges." Journal of Public Economics, 92(5–6), 1061–1077.

Krueger, A. B. and Schkade, D. A. (2008). "The reliability of subjective well-being measures." Journal of Public Economics, 92(8–9), 1833–1845.

Dolan, P. and Kavetsos, G. (2016). "Happy Talk: Mode of Administration Effects on Subjective Well-Being." Journal of Happiness Studies, 17(3), 1273–1291.

Conti, G. and Pudney, S. (2011). "Survey Design and the Analysis of Satisfaction." Review of Economics and Statistics, 93(3), 1087–1093.

Smith, T. W., Davern, M., Freese, J. and Morgan, S. L. General Social Survey 1972–2022 Cumulative Codebook. NORC at the University of Chicago.

NORC at the University of Chicago (2021). General Social Survey 2021 Cross-section: Methodological Primer.

Stone, A. A. and Mackie, C. (eds) (2013). Subjective Well-Being: Measuring Happiness, Suffering, and Other Dimensions of Experience. National Research Council, National Academies Press.

Kahneman, D. and Deaton, A. (2010). "High income improves evaluation of life but not emotional well-being." Proceedings of the National Academy of Sciences, 107(38), 16489–16493.

Killingsworth, M. A., Kahneman, D. and Mellers, B. (2023). "Income and emotional well-being: A conflict resolved." Proceedings of the National Academy of Sciences, 120(10), e2208661120.

Strathern, M. (1997). "'Improving ratings': audit in the British University system." European Review, 5(3), 305–321.

Cooperrider, D. L. and Whitney, D. (2005). Appreciative Inquiry: A Positive Revolution in Change. Berrett-Koehler.

Ostrom, E. (1990). Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press.

Bureau of Labor Statistics. Employer Costs for Employee Compensation. Quarterly news release, US Department of Labor.

Note on figures. Every figure in this chapter is computed in lib/verify/V_10.py and printed with its inputs. Published survey parameters — sample sizes, coverage, response rates, the GSS happiness shares, the WELLBY — are labelled CITED there. Field rates per completed interview are labelled PLANNING RATES: substitute a real quotation and the arithmetic re-runs. The mode-of-administration effect is taken as the 0.20-point midpoint of the 0.10-to-0.30 band in Dolan and Kavetsos (2016).