Haute Lumière
Commerce · VI.02 · MMXXVI · daylight
One page each. A reader who reads only these ten pages has the chapter.
The idea. Elinor Ostrom's eight design principles are not a list of good practices. They are the conditions she found present in institutions that had governed a shared resource for centuries, and absent or partial in the ones that collapsed.
From Table 3.1 of Governing the Commons, page 90:
The three things the popular paraphrase drops. Principle 4 is a compound claim — audit both things, and be answerable to the audited; an external inspectorate fails it. Principle 5's graduation is by context, not by repeat offence. Principle 7 is a negative — the authority need not bless the arrangement, only refrain from attacking it, which is a far weaker and far more attainable condition.
Worked example. Törbel, in the Swiss Alps, charter of 1483: no villager may send more cows to the summer alp than they can overwinter on hay from their own land. That single rule satisfies principle 2 exactly — appropriation pegged to provision capacity — and gives principle 4 away free, because every neighbour can count the hay.
You already know this because you have worked somewhere with an unwritten rule that everyone followed and nobody enforced, and you could have named, if asked, exactly what made it stick.
The idea. In 2010 Michael Cox, Gwen Arnold and Sergio Villamayor-Tomás went and checked, and the checking is more useful than the list.
studies coded 91
cases coded 77
mean support, on a 1-5 scale 3.73 = 74.6 % of the scale
Fisher's exact test between presence-of-principle and reported success was significant at 5 per cent for every principle except the eighth, which reached 10 per cent. Every principle had at least twice as many supportive as unsupportive cases.
The split that matters. They sorted studies by method:
overt studies (name the principles) 60 mean 3.60
implicit studies (never mention them) 31 mean 3.97
abstract, non-empirical studies 9 median 2.0
overt, with those 9 removed mean 3.82
Nearly all the criticism came from the nine studies that were not empirical. Remove them and the empirical literature is uniformly supportive — including the thirty-one studies that were not trying to test the principles and did not know they were.
Why it matters. Evidence that comes strongest from the people who were not looking for it is the most valuable kind there is, because it cannot have been produced by wanting the answer. This is the reason the principles can be leaned on at all.
You already know this because you trust a compliment from someone who did not know you were listening more than one delivered to your face.
The idea. Cox and colleagues found that four of Ostrom's eight were each carrying two separable claims, and separating them made every one of them more useful.
| Ostrom | Cox et al. |
|---|---|
| 1 boundaries | 1A user boundaries · 1B resource boundaries · 1C congruence between the two |
| 2 congruence | 2A congruent with local conditions · 2B appropriation congruent with provision |
| 4 monitoring | 4A social monitoring · 4B environmental monitoring |
| 3, 5, 6, 7, 8 | unchanged |
Eight principles become twelve. Principles 3, 5, 6, 7 and 8 they left exactly as they were.
The trap, and it catches nearly everybody. In the coding for the meta-analysis, 4A meant the presence of monitors and 4B meant their accountability to users. In the published reformulation, principle 4 splits along a completely different line — social versus environmental monitoring.
So when the paper reports that 4B was very strongly supported, that is accountable monitors — not environmental monitoring. The letters in the scored table and the letters in the reformulated list are not the same objects.
Why it matters. If you have ever seen a slide claiming environmental monitoring is the best-evidenced of the principles, it is reading the wrong table. Cite the finding, not the letter.
You already know this because you have seen a footnote's meaning survive a copy-and-paste while its referent quietly did not.
The idea. Baggio and eleven colleagues asked in 2016 whether the principles combine, and the answer reshapes how the list can be used.
They re-coded 69 cases from the Cox dataset — 25 forestry, 24 irrigation, 20 fishery — and ran qualitative comparative analysis on configurations rather than components.
cases 69
complete cases (every principle coded) 27 = 39.1 %
principles in the analysis 11
logically possible configurations 2,048 ( 2^11 )
configuration coverage 1.32 %
The findings.
Why it matters. You cannot score seven out of twelve and declare yourself most of the way there. The list is not additive, and the honest reading of the evidence is that you want the cluster, not the count.
You already know this because you have seen a recipe with one ingredient missing produce something that was not seven-eighths of a cake.
The idea. The principles become an instrument the moment you attach a rubric with a stated denominator.
0 absent
1 present in form only
2 present and operating
3 present, operating, and measurable from outside without permission
Worked example — the English-language Wikipedia, census read live from the MediaWiki API on 17 September 2026.
| Principle | Score | Evidence |
|---|---|---|
| 1A user boundaries | 2 | No boundary at the door; 34 rights groups above it |
| 1B resource boundaries | 3 | 7,240,827 articles inside 66,269,160 pages |
| 1C congruent boundaries | 1 | Resource bounded, user set unbounded |
| 2A rules fit conditions | 3 | Protection tiers scale to threat |
| 2B appropriation ≈ provision | 1 | 262,745 active of 54,539,294 registered — 0.4818 % |
| 3 collective choice | 3 | Open request for comment; 13 elected arbitrators |
| 4A social monitoring | 3 | Every edit emits a public diff; 308 bots |
| 4B environmental monitoring | 3 | Quality assessment, public pageview data |
| 5 graduated sanctions | 3 | Warning → block → topic ban → indefinite, appealable |
| 6 conflict resolution | 3 | Talk → noticeboard → arbitration; cost is time only |
| 7 right to organise | 1 | A foundation holds servers, marks and terms |
| 8 nested enterprises | 3 | Article → project → noticeboard → arbitration → foundation |
full score 29 of 36 = 80.6 %
the core (1B 2A 2B 4B 5 6) 16 of 18 = 88.9 %
the non-core 13 of 18 = 72.2 %
Why it matters. The core reads higher than the whole, because both weak scores sit outside the load-bearing set. A single total would have hidden that, and the diagnosis is the only part of the score that is any use.
You already know this because you have seen an average conceal the one number that mattered.
The idea. A score tells you how many conditions are met. It does not tell you what meeting them is worth. For that you need a likelihood ratio.
W = log2 [ P(principle present | endures) / P(present | fails) ]
evidence = sum over i of (score_i / 3) x W_i
For the Wikipedia score, the sum of score/3 is 9.667. The multiplier — the likelihood ratio — has never been measured for a digital commons.
Sweep it. Prior odds 0.4658 (a 31.78 % base rate, from Brief 7):
LR = 1.2 (weak) 2.54 bits posterior 73.08 %
LR = 2.0 (moderate) 9.67 bits posterior 99.74 %
LR = 4.0 (strong) 19.33 bits posterior 99.9997 %
The same score, the same institution, the same day — and a band 16.79 bits wide. And that assumes the twelve are independent, which Brief 4 says they are not, so this is the optimistic uncertainty.
What a score licenses. A diagnosis. You know exactly which conditions are unmet and can go and fix them. That is real and it is most of the value.
What it does not license. Any probability of survival. Any comparison against a commons scored by a different hand. Any claim that raising the score raises the odds. The conversion factor from points to odds does not exist yet, and a ratio you have not measured cannot be assumed to be one.
You already know this because you have seen a credit score used as if it were a probability, by someone who had never seen the default curve behind it.
The idea. The principles are conditions associated with durability within a sample. Two different selection faults sit underneath, and they are usually run together.
Ostrom's 1990 sample is selected on survival. It is explicitly a study of long-enduring institutions. Every case is a survivor by construction. That design cannot produce a base rate and never claimed to.
Cox's sample is selected on documentation. Failures are in it — the success variable takes zero — but a case enters by having been written about, and the failures that get written about are the instructive ones. A commons founded on a Tuesday and forgotten by Christmas produces no study.
What the record looks like when somebody counts the dead. Schweik and English classified the whole SourceForge population:
projects, 2009 census 174,333 (107,747 in 2006)
success in growth 24,899 14.28 % (paper: 14)
abandoned in growth 53,450 30.66 % (paper: 31)
determinate growth outcomes 78,349
success share 31.78 %
Roughly two in three die. And TeBlunthuis, Shaw and Hill studied 740 wikis, stating plainly they were the top one per cent by unique article editors — implying a population of about 74,000, of which 73,260 have never been looked at.
The scale comparison:
Cox's entire case base 77 cases
one digital-commons census 174,333 projects
ratio 0.0442 % 2,264 projects per case
Why it matters. Every claim of the form "commons with these principles survive" is conditioned on a sample whose denominator excludes almost everything that died.
You already know this because you have noticed that every published founder memoir is written by somebody whose company still exists.
The idea. The missing quantity is a likelihood ratio, and getting one requires a cohort assembled at birth rather than a sample assembled at fame.
The design. An inception cohort with blind coding.
The size of it, at 80 per cent power, Bonferroni-corrected across eleven principles, to detect a prevalence difference of 0.30:
alpha unadjusted 0.0500 z = 1.9600
alpha over 11 tests 0.004545 z = 2.8376
n per group, unadjusted 42
n per group, adjusted 73
at a 31.78 % base rate, enrol 230 founded commons
failures that then arrive 157
Two hundred and thirty. That is the whole study.
Why it matters. It would turn 80.6 per cent from a description into a prediction, and it is well within the reach of a single research group or a single consortium of consortia. It has not been run.
You already know this because every drug you have taken was licensed on a cohort, not on a collection of grateful letters.
The idea. In every durable commons, the evidence of a violation is produced automatically by the act of appropriation itself. Where monitoring is a separate funded job, it is the first line cut in a bad year.
Worked example. Wikipedia's entire enforcement record — every edit anyone has ever made, with the diff that shows exactly what changed:
edits, lifetime 1,370,633,170
stored bytes per revision delta 500 B ILLUSTRATIVE
total 685.3 GB
object storage at $0.023/GB/month
annual cost of holding all of it $189.15 / yr
per edit per year $1.38e-07
The evidentiary base of the largest commons ever assembled, held for about the price of one dinner a year. In a Valencian huerta the same principle costs a salaried ditch-rider, every year, forever.
The design lesson. Look for the arrangement in which using the resource generates the record. Törbel pegged cows to hay, so the barn is the inspection. The Peruvian contiguous irrigation order puts each farmer in the next one's sight. Take that design even at some cost in convenience, because it is the one that survives a budget round.
The warning. Cheap monitoring can become total monitoring. The same property that makes principle 4 free in a digital commons makes newcomer rejection instantaneous and impersonal — and the record shows Wikipedia's transition from growth to decline running alongside the maturation of its automated quality control. A commons can enforce itself into sterility.
You already know this because you have watched a team's code review go from a conversation to a gate, and noticed who stopped contributing.
The idea. A commons is an institutional technology with an overhead. Below a certain number of parties it is an expensive way to do something a few contracts would do better, and saying so is the most credible thing in any proposal.
The arithmetic. Bilateral agreements grow as n(n-1)/2. A rulebook grows as n. So there is a crossover, and it is a member count.
Worked example, a nine-member consortium sharing a reference dataset — an illustrative scenario; only the arithmetic on it is computed:
value appropriated per year £4,500,000
attestation + secretariat + panel reserve
£45,000 + £180,000 + £30,000 £255,000 5.67 %
------------------------------------------------------------
bilateral: 36 pairs at £12,000 each £432,000 9.60 %
------------------------------------------------------------
saving £177,000 41.0 %
advantage at nine members 1.69 x
The crossover:
n = 7 21 pairs £252,000 < £255,000 bilateral still wins
n = 8 28 pairs £336,000 > £255,000 the rulebook wins
Eight members. Below it, write contracts. Above it, write a rulebook, and by fifteen members it is not close.
The one number on the front page. Annual governance cost divided by annual value appropriated, compared against the same ratio for the bilateral alternative. If it clears, the commons is not an ethical structure — it is the cheaper structure, and it should be presented in exactly those words.
You already know this because you have been in an organisation that replaced thirty side agreements with one policy, and felt the relief before anyone costed it.
All figures in these briefs are computed in lib/verify/VI_02.py and sourced in the chapter's Works Cited. The Wikipedia census is a live API read dated 17 September 2026; the consortium scenario and the likelihood ratios are labelled illustrative and assumed in that module, because nobody has measured them.