Haute Lumière
Commerce · VI.02 · MMXXVI · daylight
Volume VI — Governance and the Commons
Twelve principles, one commons.
You have almost certainly met the eight design principles already. They travel well — a slide with eight bullets, a workshop handout, a paragraph in a strategy document. They are the most portable finding in the whole of institutional economics, and portability is exactly what has been done to them.
What has mostly not travelled is the second half of the work: the thirty-five years of people going out and checking whether the eight principles actually predict anything. That literature exists. It is large, it is careful, and it is more interesting than the list. It found that the principles hold up. It also found that four of them were two principles wearing one coat, that the letters in the standard reformulation do not mean what most people citing them think they mean, that no single principle is either necessary or sufficient, and that the whole apparatus rests on a denominator nobody has ever assembled.
This chapter does three things in order. It states the eight in Elinor Ostrom's own terms, from page ninety of Governing the Commons, because the popular paraphrase has drifted and the drift matters. It walks the empirical record — Cox, Arnold and Villamayor-Tomás on ninety-one studies, Baggio and eleven colleagues on sixty-nine cases — and reports what was supported, what was not, and what got rebuilt. And then it does the thing almost nobody does: it takes a real modern commons that is not a fishery and not an irrigation system, scores it principle by principle with numbers you can re-fetch yourself, and says precisely what that score licenses and what it does not.
The last part will be the most useful to you and it is also the part that has the sharpest edge in it. A high score is worth something. Exactly how much is a question with an arithmetic answer, and the arithmetic answer is not comfortable.
You are not being handed a checklist. You are being handed an instrument with its calibration certificate attached, which is a considerably better thing to own.
— The Editors
Start where Ostrom started, which is with institutions that have simply worked for a very long time and were not designed by anybody clever.
The huertas of Valencia. Irrigation communities on the Turia and the Júcar have allocated water among thousands of smallholders since at least the fifteenth century and plausibly much earlier. Disputes go to the Tribunal de las Aguas, which sits in the open air outside the Cathedral of Valencia on Thursdays at noon. The judges are elected irrigators. Proceedings are oral, in Valencian, and typically last minutes. There is no appeal and no written judgment. Ostrom's point about this case is not that it is charming; it is that the cost of resolving a conflict is close to zero for the person bringing it, and a conflict-resolution mechanism that is expensive to use is a conflict-resolution mechanism that does not exist.
The Törbel commons, Switzerland. A village charter dated 1483 governs the alpine meadow, and the rule that carries the load is beautifully simple: no villager may send more cows to the summer pasture than they can overwinter on hay from their own land. The appropriation right is pegged to a provision capacity that is already visible to every neighbour. Nobody has to be inspected, because everybody's winter barn is the inspection.
The Japanese iriaichi. Village common lands, some governed continuously for three centuries and more, with opening dates for the harvest set by the village, detailed graded penalties, and monitors — the saruta — drawn from the villagers themselves and rotated.
The Philippine zanjera. Farmer-built irrigation federations in which a member's water share is explicitly proportional to the labour contributed to building and maintaining the canal. Appropriation and provision are the same number seen from two sides.
Ostrom went looking for the common structure in these and found eight conditions. Here they are in her own language, from Table 3.1 of Governing the Commons, page ninety. Read them slowly; the popular version has sanded most of the qualifiers off.
Three things in that wording are routinely lost and each one is load-bearing.
Principle 4 is not "have monitoring." It is a compound claim: that the monitors audit both the condition of the resource and the behaviour of appropriators, and that the monitors are accountable to the people being monitored or are those same people. An external inspectorate satisfies the paraphrase and fails the principle.
Principle 5 says "depending on the seriousness and context of the offence." Graduation is not a ladder of escalating penalties for repeat offences. It is proportionality to circumstance — the first offence in a drought year is not the first offence in a wet one.
Principle 7 is a negative. It does not require that an external authority grant or bless the arrangement. It requires only that the authority not challenge it. That is a far weaker and far more achievable condition, and it is why so many commons survive under states that have never heard of them.
And the appreciative finding underneath all four cases: none of these institutions was designed. They were accreted, by people with no theory, paying attention to what kept working. The principles are a description of convergent evolution. That is a much stronger kind of evidence than a design brief, and it is also the reason the principles cannot be read as a recipe — which is the argument of the next movement.
In 2010 Michael Cox, Gwen Arnold and Sergio Villamayor-Tomás published the review the field had been waiting for. They coded 91 studies and 77 cases against the principles and scored each study's support on a one-to-five scale.
studies coded 91
cases coded 77
mean evaluation, all studies 3.73 / 5 = 74.6 %
3.73 out of 5 — the paper's own gloss is "slightly below the value for moderately supportive." Fisher's exact test between presence-of-principle and reported success was significant at the five per cent level for every principle except the eighth, which reached ten per cent. Every one of the principles had at least twice as many supportive as unsupportive cases.
Now the part that decides how much weight to put on this. They split the studies by method, and the split is the finding:
overt studies (name the principles) 60 mean 3.60
implicit studies (do not name them) 31 mean 3.97
difference +0.37 (P = 0.343)
abstract, non-empirical studies 9 median 2.0
overt, with the abstract studies removed mean 3.82 (P = 0.842)
Nine of the ninety-one studies — 9.9 per cent — were abstract rather than empirical, and those nine carried nearly all of the criticism. Remove them and the empirical literature is uniformly supportive, whether or not it had heard of the principles. The principles are best supported by the studies that were not trying to test them, which is the shape of evidence you want, because it is the shape that cannot have been produced by wanting the answer.
Their verdicts, principle by principle, in their own words:
| Principle | Verdict in Cox et al. |
|---|---|
| 1A user boundaries | strong evidence |
| 1B resource boundaries | moderate evidence |
| 2A congruence with local conditions | supported |
| 2B appropriation/provision congruence | supported |
| 3 collective choice | moderately well supported |
| 4A presence of monitors | moderately well supported |
| 4B monitors accountable to users | very strongly supported |
| 5 graduated sanctions | moderately well supported |
| 6 conflict resolution | moderately well supported |
| 7 right to organise | moderately supportive |
| 8 nested enterprises | moderately supportive |
And here is the trap that almost every citation of this paper walks into. In the coding, 4A meant the presence of monitors and 4B meant their accountability to users. In the reformulation they proposed at the end of the paper, principle 4 splits along a different line entirely — into users monitoring each other's behaviour, and users monitoring the condition of the resource. The "very strongly supported 4B" is accountable monitors. It is not environmental monitoring. The letters in the scored table and the letters in the published reformulation are not the same objects. If you have ever seen a slide claiming that environmental monitoring is the best-evidenced of Ostrom's principles, that slide is reading the wrong table.
The reformulation itself takes eight principles to twelve: principle 1 into user boundaries, resource boundaries and the congruence between them; principle 2 into congruence with local conditions and congruence of appropriation with provision; principle 4 into social and environmental monitoring. Principles 3, 5, 6, 7 and 8 they left alone.
Cox and colleagues tested each principle against outcome separately. Jacopo Baggio and eleven co-authors asked the obvious next question in 2016: do the principles combine?
They re-coded 69 cases from the Cox dataset — 25 forestry, 24 irrigation, 20 fishery — under a stricter definition of success, and ran qualitative comparative analysis on the configurations rather than the components.
cases re-coded 69
complete cases (every principle coded) 27 = 39.1 %
incomplete 42
principles carried into the analysis 11
logically possible configurations 2,048 ( 2^11 )
configuration coverage 1.32 % ( 27 of 2,048 )
Their findings, and each one changes how you should use the list:
That last point is the one to carry into any non-natural commons. Which principles carry the load depends on the physics of the thing being shared. A mobile resource makes boundaries expensive and social definition cheap. A static one reverses it.
And note the honest constraint in their own numbers: twenty-seven complete cases against two thousand and forty-eight possible configurations is 1.32 per cent coverage. They say so themselves, and they built a reliability metric specifically because they could not pretend otherwise.
Now apply it. Not to a fishery. To the English-language Wikipedia — a digital commons with a bounded resource, an unbounded user population, a written constitution, an elected judiciary, and a complete public record of every act of appropriation and every act of enforcement since 2001.
Here is the census, read live from the MediaWiki API on 17 September 2026. Every one of these moves by the minute; a census without a timestamp is a rumour.
articles 7,240,827
pages, all namespaces 66,269,160 articles are 10.93 %
edits, lifetime 1,370,633,170 189.29 per article
registered accounts 54,539,294
active accounts (30 days) 262,745 0.4818 % of registered
administrators 809 324.8 active per admin
bot accounts 308 8,950.3 articles per admin
bureaucrats · checkusers · interface admins 15 · 47 · 15
arbitrators 13
distinct rights groups 34
Those officer counts overlap — an arbitrator is nearly always an administrator — so they are never summed here. A summed officer count would be a fiction.
The scoring rubric, stated before the scores so you can disagree with the instrument rather than the result. Each of the twelve reformulated principles takes 0 to 3:
0 absent
1 present in form only
2 present and operating
3 present, operating, and measurable from outside without permission
| # | Principle | Score | The countable evidence |
|---|---|---|---|
| 1A | user boundaries defined | 2 | No boundary at the door — anyone may edit. But 34 rights groups form a real lattice of graded appropriation rights above that door. |
| 1B | resource boundaries defined | 3 | Exact. 7,240,827 articles inside 66,269,160 pages; namespaces are the boundary; every page has a title and a full history. |
| 1C | user and resource boundaries congruent | 1 | They do not line up at all. The resource is bounded and the user set is the species. |
| 2A | rules congruent with local conditions | 3 | Protection tiers scale to threat — semi, extended-confirmed, full — and sourcing rules tighten by subject class. |
| 2B | appropriation congruent with provision | 1 | The honest failure. Appropriation is reading, and it is free and unbounded. Provision is done by 262,745 accounts out of 54,539,294 — 0.4818 per cent. |
| 3 | most affected can modify the rules | 3 | Policy changes by open request for comment; 13 arbitrators elected by the editing community. |
| 4A | social monitoring | 3 | Every edit emits a public diff, automatically. Watchlists, 308 bots, edit filters. |
| 4B | environmental monitoring | 3 | Article quality assessment, featured and good article review, public pageview and quality data. |
| 5 | graduated sanctions | 3 | Warning → short block → escalating block → topic ban → indefinite; duration explicitly tied to likelihood of repetition; appealable throughout. |
| 6 | low-cost conflict resolution | 3 | Talk page → third opinion → noticeboard → arbitration. Cost to the complainant is time, and nothing else. |
| 7 | right to organise unchallenged | 1 | The servers, the trademark and the terms of use belong to a foundation. On 10 June 2019 that foundation imposed a one-year project-specific ban over the community's head. |
| 8 | nested enterprises | 3 | Article → project → noticeboard → arbitration committee → foundation → global movement. |
sum of the twelve 29 of 36 = 80.6 %
the core (1B 2A 2B 4B 5 6) 16 of 18 = 88.9 %
the non-core 13 of 18 = 72.2 %
The core reads higher than the whole, because both of the weak scores — 1C and 7 — sit outside the set Baggio and colleagues found load-bearing. That is a real finding about this commons, and it is also a warning about single totals: an aggregate that hides which principle failed has thrown away the only part of the diagnosis that was any use.
Here is the cut, and it applies to every commons scorecard you will ever be handed, including this one.
Eighty per cent sounds like a lot. Ask what it is worth. The evidential weight of observing a principle is not the point you awarded it. It is the likelihood ratio — how much more often that principle is present in commons that endure than in commons that die — expressed in bits:
W = log2 [ P(principle present | endures) / P(present | fails) ]
evidence = sum over i of (score_i / 3) x W_i
The sum of score/3 across the twelve is 9.667. So the arithmetic is one multiplication. And the multiplier — the likelihood ratio — has never been measured for a digital commons, because measuring it requires the dead, and nobody has assembled the dead.
Sweep it across three plausible values and watch what happens. The prior odds come from the one population census of a digital commons that counts its failures, which we come to in a moment: 31.78 per cent of growth-stage open-source projects with a determinate outcome succeed, giving prior odds of 0.4658.
likelihood ratio bits of evidence posterior probability of enduring
----------------------------------------------------------------------------
1.2 (weak) 2.54 bits 73.08 %
2.0 (moderate) 9.67 bits 99.74 %
4.0 (strong) 19.33 bits 99.9997 %
The same score, on the same institution, on the same day, supports a survival probability anywhere from seventy-three per cent to five nines. A band 16.79 bits wide. And that calculation assumes the twelve principles are independent, which Baggio and colleagues demonstrated they are not — so the band above is the optimistic version of the uncertainty.
This is what a score licenses and what it does not.
It licenses: a diagnosis. You now know precisely which two conditions this commons does not satisfy, and you know they are 1C and 7, and you can go and work on them. That is real and it is most of the value.
It does not license: any statement about probability of survival, any comparison against a different commons scored by a different hand, and any claim that raising the score raises the odds. The conversion factor from points to odds is unmeasured, and a ratio you have not measured cannot be assumed to be one.
The principles are conditions associated with durability within a sample of documented commons. They are not a recipe, and the reason is a fault in the denominator that no amount of careful coding inside the sample can repair.
Two different selection faults, usually run together and worth separating.
Ostrom's own 1990 sample is selected on survival. It is explicitly a study of long-enduring institutions. Every case in it is, by construction, a survivor. That design cannot estimate a base rate, and Ostrom never claimed it could.
Cox et al.'s sample is selected on documentation, which is different and subtler. Their success variable does take the value zero — failures are in there. But a case enters the sample by having been written about, and the failures that get written about are the instructive ones. A commons that was founded on a Tuesday and forgotten by Christmas produces no study.
What the base rate actually looks like when somebody counts the dead: Charles Schweik and Robert English classified the entire SourceForge population.
projects, 2006 census 107,747
projects, 2009 census 174,333 1.62 x in three years
success in growth 24,899 14.28 % (paper: 14)
abandoned in growth 53,450 30.66 % (paper: 31)
indeterminate, initiation / growth 16,806 / 12,052
------------------------------------------------------------------
determinate growth-stage outcomes 78,349
success share of determinate 31.78 %
Roughly two in three die. And for wikis specifically, Nathan TeBlunthuis, Aaron Shaw and Benjamin Mako Hill studied 740 — stating plainly that these are the top one per cent of the host's wikis by unique registered article editors. That implies a population of about 74,000, of which 73,260 have never been examined by anybody.
Set the scales beside each other and the problem is arithmetic rather than rhetorical:
Cox's entire case base 77 cases
one digital-commons census 174,333 projects
ratio 0.0442 %
SourceForge projects per Cox case 2,264
The commons literature's whole coded case base is four parts in ten thousand of a single population census — and the census is the one that contains the failures.
An inception cohort with blind coding. It is not expensive and it has not been run.
The size of it, at 80 per cent power, Bonferroni-corrected across eleven principles, to detect a prevalence difference of 0.30 between endurers and failures:
alpha, unadjusted 0.0500 z = 1.9600
alpha, adjusted over 11 tests 0.004545 z = 2.8376
n per group, unadjusted 42
n per group, adjusted 73
at a 31.78 % base rate, enrol 230 founded commons
failures that would then arrive 157
Two hundred and thirty. That is the whole study — a cohort of a few hundred, coded at birth, followed until each one is alive or dead. It would produce, for the first time, the likelihood ratio that the entire scorecard depends on, and it would turn 80.6 per cent from a description into a prediction.
Until someone runs it, score your commons, read the diagnosis, fix what the diagnosis names — and do not convert the total into a probability, because the conversion factor does not exist yet.
In the version of this that has already happened, a shared resource arrives with its governance already legible, the way a building arrives with its structural calculations.
The consortium agreement for a pooled dataset opens with a schedule that names each of the twelve conditions and states, in a sentence each, how this arrangement satisfies it and where it does not. The two it does not satisfy are named on the first page rather than discovered in year four. Nobody finds this unusual; it reads the way an environmental report reads, and for the same reason.
Attestation is routine. Once a year an independent practitioner walks the governance the way an auditor walks a stock count — reading the conflict log, timing the median resolution, sampling the sanctions for graduation, checking that the people who monitor are answerable to the people they monitor — and issues a score with a stated scope and a stated denominator. The score does not claim to predict survival. It claims to describe structure, which is what it can honestly do, and that turns out to be enough to price a subscription.
Because the score is standard, the cohort assembles itself. Every consortium that attests is a coded case, coded at founding and before anybody knows how it ends, and after a few years there are several hundred of them with outcomes attached. The likelihood ratios get published. The scorecard acquires a calibration curve. The instrument that was a heuristic in 1990 and a meta-analysis in 2010 becomes, somewhere in the 2030s, an actuarial table, and the people who built the commons that failed are the ones who made that possible, which is a better legacy than anonymity.
And the language settles. People stop saying tragedy of the commons to mean "shared things get wrecked" and start saying it the way it is actually true: that an unowned, unmonitored, unsanctioned, unbounded resource gets wrecked, and that each of those four adjectives names a condition somebody can put right this quarter.
Here is how to build the thing, in the order the conditions have to be met.
Stage one — decide whether it should be a commons at all.
Before any of the twelve, one prior question: is the cost of governing this shared thing less than the value it produces? A commons is an institutional technology with an overhead, and where the overhead exceeds the gain the honest answer is a contract or a firm. The arithmetic is in the last movement and it resolves to a member count.
Stage two — build the four the record says carry the load.
Baggio and colleagues found the absence of 2A, 2B, 4B and 5 most strongly associated with failure. Build those first, and build them in this order, because each makes the next cheaper.
Stage three — buy the boundaries you can afford.
1B is usually cheap: name the resource precisely. 1A is usually expensive in a digital commons and often undesirable, because the openness is the point. Where you cannot bound users at the door, bound them above it — a lattice of graded rights earned by demonstrated provision. That is what 34 rights groups are for, and it is how an unbounded commons recovers most of what principle 1A was protecting.
Stage four — make conflict cheap and the ladder short.
Principle 6 is measured in hours and money, not in the existence of a policy. Set a service level: a first response inside a stated window, a decision inside a stated window, a capped cost to the complainant. Then publish the median. A conflict mechanism whose median resolution time is unpublished is a mechanism nobody can trust before they need it.
Stage five — write principle 7 down, because you will not be given it.
This is the one nobody builds because it feels like somebody else's decision. It is not. If a sponsor, a foundation, a platform or a parent company holds the servers, the trademark or the terms, then the community's right to make its own rules exists at that party's pleasure — until it is written into a reserved matters schedule with a notice period, a published reason requirement, and an appeal. On 10 June 2019 the largest digital commons in the world discovered exactly what it had not written down. Write it down.
Stage six — nest it, and make the nesting real.
Principle 8 fails most often as a diagram that no dispute has ever travelled up. Test it by sending one. An escalation path that has never carried traffic is a drawing.
Three things make a commons self-sustaining and they are not the same three that make a project self-sustaining.
Monitoring has to be a byproduct, not a job. The Törbel barn, the contiguous irrigation order that puts each farmer in the next one's sight, the public diff — in every durable case the evidence of a violation is produced automatically by the act of appropriation itself. Where monitoring is a separate funded role, it is the first line cut in a bad year, and the commons unravels from there. If you can find a design where using the resource generates the record, take it even at some cost in convenience.
The rules have to be amendable by the people bound by them. Principle 3 is the one that converts a commons from a static arrangement into a living one. A rulebook that cannot be changed from inside will be either obeyed into irrelevance or abandoned.
Somebody has to be able to lose. A commons with no enforced sanctions is not a commons; it is an open-access resource with good manners, and good manners are not a boundary condition.
Now the failure modes, named honestly, because the principles were formulated to specify collapse and it would be a poor tribute to soften them.
It fails when appropriation and provision drift apart. This is the most common failure in a digital commons and it is slow. Readers multiply; editors do not. Users grow; maintainers do not. The 0.4818 per cent figure is not a scandal — it is the structural condition of every knowledge commons, and it means the provision side has to be actively cultivated or it thins to nothing while every other indicator still reads healthy.
It fails when monitoring becomes so cheap that it becomes total. The same property that makes principle 4 free in a digital commons — every act leaves a record — makes newcomer rejection instantaneous and impersonal. The record shows Wikipedia's own transition from growth to decline running alongside the maturation of its automated quality control. A commons can enforce itself into sterility.
It fails when principle 7 was never written down and the external party exercises a right everybody had forgotten it held.
And it fails where the arithmetic never supported it. Below the member count at which shared governance beats bilateral contracting, a commons is an expensive way to do something simpler. That threshold is the subject of the next movement, and it is a number.
There is a specific pleasure in a rule that enforces itself, and everybody who has built one knows it. The Törbel hay rule is the purest example in the literature: you cannot cheat, not because you would be caught but because the thing you would have to cheat with is standing in your own barn where your neighbours can count it. It is an elegant piece of engineering and it was built by farmers with no word for engineering.
There is a second pleasure, quieter, in the Thursday tribunal. Minutes, not months. Outdoors, in the language people actually speak, in front of anyone who cares to stand there. No filings, no fees, no transcript. A thousand years of water disputes have been settled by seven elected farmers in the time it takes to drink a coffee, and the reason it works is not tradition — it is that the cost of bringing a complaint is low enough that complaints get brought early, while they are still small.
And a third, which is the one this chapter is really about: the pleasure of finding out that the folk wisdom was checked. Somebody went and read ninety-one studies. Somebody else re-coded sixty-nine cases and found the combinations. The eight became twelve because the evidence asked for it. The best thing about Ostrom's principles is not that they are true. It is that the field treated them as a hypothesis, which is the highest compliment one set of researchers can pay another, and it is the reason you can lean on them today.
Everything above becomes real when it acquires an instrument. Here is the instrument, in the form a general counsel and a treasurer will both recognise.
The structure: a shared-infrastructure consortium under a governance deed with an annual attestation and a fee ratchet.
Not a joint venture and not a contract web. A company limited by guarantee (or its local equivalent) holding the shared resource, with members admitted by class, a published rulebook, and — the part that is new — a governance schedule drafted as twelve numbered covenants, one per principle, each with a named measurable.
The mechanics.
The counterparty. Always the members first. An external funder or platform joins as a member under the same deed or does not join.
The number that decides it. One figure, on the front page, and it is a ratio:
annual governance cost the same ratio for the
----------------------------------- < bilateral alternative
annual value appropriated
Worked, on a nine-member consortium sharing a reference dataset — an illustrative scenario; every input below is a scenario figure and only the arithmetic on them is computed:
members 9
value appropriated per year £4,500,000
------------------------------------------------------------
attestation £45,000
secretariat £180,000
conflict panel reserve £30,000
governance cost £255,000 5.67 %
------------------------------------------------------------
the alternative: bilateral agreements
pairs, n(n-1)/2 36
upkeep at £12,000 each £432,000 9.60 %
------------------------------------------------------------
saving £177,000 41.0 %
advantage at nine members 1.69 x
And the crossover, which is the number to actually remember: eight members.
n = 7 21 pairs £252,000 < £255,000 — bilateral still wins
n = 8 28 pairs £336,000 > £255,000 — the rulebook wins
Bilateral agreements grow as n²; a rulebook grows as n. Below eight members a commons is an expensive way to do something a few contracts would do better, and saying so is the most credible thing in the proposal. Above it the advantage compounds, and by fifteen members it is not close.
The first ninety days.
| Day | Action | Artifact |
|---|---|---|
| 1–15 | Count the members and run the crossover | The one-page ratio |
| 16–30 | Write covenants 1B, 2A and 2B — the resource and the contribution formula | The data dictionary and the equation |
| 31–45 | Write covenants 4A, 4B and 5 — who monitors whom, and the ladder | The monitoring schedule |
| 46–60 | Write covenant 7 and negotiate it with whoever holds the infrastructure | The reserved matters schedule |
| 61–75 | Stand up the conflict panel and publish its service level | The panel terms |
| 76–90 | Commission the first attestation and set the baseline score | The attested twelve, dated |
The attestation in month three is the irreversible commitment. A published score with a date on it cannot be quietly walked back, and next year's score will be compared to it by people you have not met.
Discovery — what is already working
Dream — what becomes possible
Design — what we build
Destiny — how it holds
Agrawal, A. (2002). "Common Resources and Institutional Sustainability." In E. Ostrom et al. (eds), The Drama of the Commons. National Academy Press, 41–85.
Baggio, J. A., Barnett, A. J., Perez-Ibarra, I., Brady, U., Ratajczyk, E., Rollins, N., Rubiños, C., Shin, H. C., Yu, D. J., Aggarwal, R., Anderies, J. M. and Janssen, M. A. (2016). "Explaining Success and Failure in the Commons: The Configural Nature of Ostrom's Institutional Design Principles." International Journal of the Commons, 10(2), 417–439. DOI 10.18352/ijc.634.
Cox, M., Arnold, G. and Villamayor-Tomás, S. (2010). "A Review of Design Principles for Community-based Natural Resource Management." Ecology and Society, 15(4), 38.
Crowston, K., Wei, K., Howison, J. and Wiggins, A. (2012). "Free/Libre Open-Source Software Development: What We Know and What We Do Not Know." ACM Computing Surveys, 44(2), 1–35.
Forte, A., Larco, V. and Bruckman, A. (2009). "Decentralization in Wikipedia Governance." Journal of Management Information Systems, 26(1), 49–72.
Halfaker, A., Geiger, R. S., Morgan, J. T. and Riedl, J. (2013). "The Rise and Decline of an Open Collaboration System: How Wikipedia's Reaction to Popularity Is Causing Its Decline." American Behavioral Scientist, 57(5), 664–688.
Hess, C. and Ostrom, E. (eds) (2007). Understanding Knowledge as a Commons: From Theory to Practice. MIT Press.
Netting, R. McC. (1981). Balancing on an Alp: Ecological Change and Continuity in a Swiss Mountain Community. Cambridge University Press.
Ostrom, E. (1990). Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press. Design principles at Table 3.1, p. 90.
Ostrom, E. (2005). Understanding Institutional Diversity. Princeton University Press.
Ostrom, E. (2007). "A Diagnostic Approach for Going Beyond Panaceas." Proceedings of the National Academy of Sciences, 104(39), 15181–15187.
Ostrom, E. (2009). "A General Framework for Analyzing Sustainability of Social-Ecological Systems." Science, 325(5939), 419–422.
Ostrom, E. (2010). "Beyond Markets and States: Polycentric Governance of Complex Economic Systems." American Economic Review, 100(3), 641–672.
Ostrom, E. and Cox, M. (2010). "Moving Beyond Panaceas: A Multi-tiered Diagnostic Approach for Social-Ecological Analysis." Environmental Conservation, 37(4), 451–463.
Ratajczyk, E., Brady, U., Baggio, J. A., Barnett, A. J., Perez-Ibarra, I., Rollins, N., Rubiños, C., Shin, H. C., Yu, D. J., Aggarwal, R., Anderies, J. M. and Janssen, M. A. (2016). "Challenges and Opportunities in Coding the Commons: Problems, Procedures, and Potential Solutions in Large-N Comparative Case Studies." International Journal of the Commons, 10(2), 440–466.
Schweik, C. M. and English, R. C. (2012). Internet Success: A Study of Open-Source Software Commons. MIT Press.
Schweik, C. M. and English, R. (2013). "Preliminary Steps Toward a General Theory of Internet-based Collective-Action in Digital Information Commons: Findings from a Study of Open Source Software Projects." International Journal of the Commons, 7(2), 234–254. Classification counts at Table 1.
Shaw, A. and Hill, B. M. (2014). "Laboratories of Oligarchy? How the Iron Law Extends to Peer Production." Journal of Communication, 64(2), 215–238.
TeBlunthuis, N., Shaw, A. and Hill, B. M. (2018). "Revisiting 'The Rise and Decline' in a Population of Peer Production Projects." Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, Paper 355.
Wilson, D. S., Ostrom, E. and Cox, M. E. (2013). "Generalizing the Core Design Principles for the Efficacy of Groups." Journal of Economic Behavior and Organization, 90(Supplement), S21–S32.
Note on figures. The English Wikipedia census is a live read of the MediaWiki API (action=query&meta=siteinfo&siprop=statistics) on 17 September 2026 and is pinned with that date in lib/verify/VI_02.py; those figures move by the minute. The Cox, Baggio, Schweik–English and TeBlunthuis figures are from the published papers as cited. The twelve-point scoring rubric, the likelihood-ratio sweep, the cohort sizing and the consortium scenario are computed in lib/verify/VI_02.py and reproducible there; the consortium's inputs are labelled illustrative in that module, as are the assumed likelihood ratios, because the whole argument of the Arithmetic is that nobody has measured them.