BEAM·readiness

BEAM: the Boundary-aware Evaluation & Alignment Method

Do you want to see a world where coercion can't take hold, where the earth is thriving instead of being spent, where what people build lifts them up instead of extracting from them?

That world runs on a different kind of structure. We pulled the criteria for it from the best sources we have on what actually lasts and doesn't harm: Elinor Ostrom's research on the commons, the Teal paradigm, the New GDP measures of real thriving. Then we built BEAM to hold anything up against it: your organization, your coordination framework, your new-paradigm alliance, your federation, whatever it is you're building.

BEAM is for people building something different.

If you're here to build wealth and move on, this honestly isn't for you. If you're building something meant to outlast you and move the world toward that, this is built for you. And it isn't really about one organization. It's about the whole thing: if every group built this way, coercion would have nowhere to stand, and thriving (for people and the planet) would be the default instead of the exception. That's what this measures.

What it actually looks at

Not your mission statement. Your structure. How decisions really get made. How power is held, and how it's let go. How people join, and how they leave. How value moves. What happens when something breaks. Structure is what decides whether good intentions survive contact with reality. The ones that hold their values are the ones built to hold them.

What you get

You answer for what you're building and show the documents that back it up. You get one honest score with all its working laid out: where you're strong, where you're exposed, and what a stronger version looks like. Nothing hidden, nothing to game. Then you can see how you sit alongside others building the same way.

What it's really for

The world we've inherited runs on one kind of structure, one that competes, extracts, concentrates power, and pushes the damage onto everyone else. The world a lot of us actually want runs on another: cooperative, regenerative, accountable, hard to capture. BEAM measures how close you already are, and what it would take to get closer.

Where the criteria come from

How your score is worked out: the whole thing, step by step

1. Each of the 21 criteria is measured against your system.
Every criterion asks one specific structural question: how decisions get made, how someone can leave, how power is held and handed on. The app reads what you've given it and places you from 0 to 100: from “nothing here,” through “it exists but it's informal,” to “it's fully in place, required, and has held up under real pressure.” Each level is written down, so the score is a match against a stated bar, not a gut call.

2. Each criterion is weighted by how much it holds the whole thing up.
The 21 aren't equal. A clear purpose or a safe way out is load-bearing: remove it and the whole thing fails or turns on its people. A regular review matters, but its absence fades slowly. So each criterion carries a weight, ranked from most to least load-bearing, and your score leans on the heavy ones.

3. Proof decides how much a score is allowed to count.
Anyone can claim anything, so evidence sets a ceiling on every criterion. A public, independently checkable record counts in full; an internal document you provide counts high; simply saying it, with nothing to show, is capped low and flagged; a vague or non-answer scores zero. You can't talk your way up, only show your way up.
(Worked example: you score 80 on a criterion, but your only proof is an internal document: solid, not independently verifiable. That evidence counts at 0.85: 80 × 0.85 = 68. Show a public, checkable version and it counts in full.)

4. It comes together into one number out of 100, every line visible.
Each criterion's level, its weight, the proof behind it, and exactly what it added are all laid out. Add the lines back up and you land on the same number. Nothing hidden, and nothing gamed without leaving a mark.

You also get a confidence figure: how much of your score stands on hard, checkable proof versus your own word. A high score built on claims and a high score built on public record are not the same thing, and the app never lets them look the same.

Technically, here's more

The evidence tiers: A/B/C/D, multipliers and caps
Tier A: Genuine substantiating document ×1 cap 100
A GENUINE document from a REAL source that substantiates the criterion — a filing, signed agreement, dated record, decision log, written policy, fund record, or a real internal governance artifact, INCLUDING one from a private system the org actually runs on (a private Notion decision log, a signed member attestation, a screenshot of it). PUBLIC / INDEPENDENT VERIFIABILITY IS NOT REQUIRED (2026-07-30). NOT merely asserted, NOT unfalsifiable, NOT fabricated. The rubric BAND still independently gates quality — a genuine but thin document that only partially meets the criterion scores a low band and/or tier B. Anti-gaming floor now rests on: no artifact -> C; unfalsifiable/fabricated -> D; authenticity or relevance uncertain -> B (downgrade-on-doubt).
Tier B: Partial or uncertain ×0.85 cap 90
A real document was shown but it only PARTIALLY substantiates the criterion, OR its authenticity or direct relevance is uncertain (downgrade-on-doubt). A document that clearly substantiates the criterion is tier A. This is the enforced floor that replaces the old 'self-supplied caps at B' gate.
Tier C: Asserted only ×0.5 cap 50
They say they do it. No artifact. FLAGGED IN OUTPUT as 'not hard evidence'.
Tier D: Vague or non-responsive ×0 cap 0
Answer does not address the criterion, or uses unfalsifiable language ('we value transparency'). Scores zero and is flagged.

Tier is assigned from the artifact supplied, never from self-rating; no artifact means the ceiling is tier C by rule. Documents your AI connector sends in are self-supplied, so they cap at tier B.

The exact formula

Per criterion: min(rubric band score, tier cap) × tier multiplier. Total: Σ(weight × criterion score) ÷ Σ(weights), with Σ(weights) = 1879 across 21 criteria. Confidence = the share of total weight backed by tier-A (independently verifiable) evidence.

The honest limit: what this deliberately does not claim

This measures structure, not real-world impact: a structurally sound organization can still be ecologically destructive, and this won't catch that.

This cannot stop an org from fabricating documents. It makes fabrication necessary rather than optional — the cost of lying goes from zero to producing a forgery. That is the achievable goal; claiming more would be dishonest.

Full criteria provenance: every criterion's origin, sources, and weight

The 21 Criteria — where each one comes from

Public by design. If we score an organization, it can ask why a criterion exists, what it's derived from, why it carries the weight it does, and what it connects to — and get a real answer. A standard you can't interrogate isn't a standard, it's an opinion with a number attached. Per criterion: what it is · what it means in practice · where it comes from · what it plugs into · how it affects society · the math. Companions: os-readiness-spec.json (machine-readable, the instrument) · assessment-demo.html (runs the formula) · matrix.md (our own scores) · score-levers.md (what would move them) · formula-reconciliation.md (how this formula and the New GDP formula relate). Owner: OS Compare. Status: COMPLETE — all 21 written (2026-07-21). Sources named per entry. Attribution: this instrument is built by the NEOS operating system. Stated once here; not tagged per-criterion or per-suggestion.

The math, in full

score = Σ(weight × criterion_score) / Σ(weight)          weights total 1,879
criterion_score = min(rubric_band, tier_cap) × tier_multiplier
confidence      = Σ(weight where tier = A) / Σ(weight) × 100

Evidence tiers: A genuine substantiating document ×1.00 (cap 100) · B partial, or authenticity/relevance uncertain ×0.85 (cap 90) · C asserted only ×0.50 (cap 50) · D vague/non-responsive ×0. ⚠️ Tier A does NOT require public or independent verifiability (retired 2026-07-30, dropped-public ruling): a genuine internal artifact from a private system the org actually runs on is tier A. The old "independently verifiable" label is wrong and contradicts os-readiness-spec.json, which is authoritative.

One score, not two (2026-07-19). We don't publish what an organization claims about itself. Assertion without an artifact simply scores low — there's no second number to argue over. What we do publish is the entire working: every criterion's band, tier, cited artifact, and exact contribution. Anyone can recompute the total and get the same answer.

Why the formula is locked (2026-07-21). Three things do not change, because changing any one of them changes every score ever issued: the 21 criteria ("I don't think we can add to this because this is based on math" — the set is closed), the weights (total 1,879), and the evidence tiers (multipliers + caps). What stays open is scoring — an organization's readings move as it builds and evidences. The instrument is fixed; the measurements move. That separation is what makes a rising score mean something.

Where the weights come from. They are comparative judgements of how load-bearing each requirement is — not measurements. Ranked 100 down to 76, derived from failure analysis: what, when it's missing, most reliably ends the organization or turns it against the people in it. Purpose sits at 100 because nothing else can be assessed without it; periodic review at 76 because its absence degrades slowly rather than fatally. The ranking is arguable and we publish it so it can be argued with. Source: requirements-top-20-and-importance.md.


What these criteria actually are — read this before using the list

This is not a corporate survival checklist. Framing it as "what stops an organization failing" is true and badly incomplete, because it leaves out why anyone should care whether it survives.

The claim underneath the whole model: the things that make an organization durable and the things that make it non-harmful are the same list. That convergence is not a coincidence and it isn't a moral hope bolted onto a technical model — it's causal, in both directions.

  • An organization that extracts from its members loses them, and slowly loses the ability to attract more. (#21 fairness, #7 capital, #11 economics)
  • An organization that dominates its own people destroys its error-correction — nobody reports the problem to the person causing it. (#4 decisions, #3 authority, #12 conflict)
  • An organization that traps people stops having to earn their participation, and degrades without resistance. (#2 safe exit)
  • An organization that can't be questioned can't detect its own drift. (#6 memory, #19 drift, #20 review)

Coercion is a failure mode, not just a wrong. That's the whole argument. It means you don't have to choose between "effective" and "ethical" — and it means an organization can be assessed on ethics without anyone having to agree on values first, because every one of these is measurable as a structural fact. You can check whether exit exists. You cannot check whether people are good.

What this is attached to

  • The five problem-load drivers / New GDP · BIRTH — attached, and the relationship is precise. BIRTH measures outcome: how much of society's problem load a different operating system removes. This model measures capacity: whether a given organization is actually built the way that removal depends on. Coupled, not merged (agreed with New GDP, 2026-07-19) — different scales, one causal chain. These criteria are the mechanism BIRTH's numbers assume. The full reconciliation — how the two formulas relate, why they aren't one, and the one place they don't line up — is formula-reconciliation.md. Figures: gdp-advocacy/numbers.md — theirs, never restated.
  • Ethics — attached structurally, not by declaration. Non-domination, consent, non-extraction, safe exit, proportional fairness, dignity. Every one appears here as a checkable structural property rather than a stated value. That's deliberate: values in a mission statement are unfalsifiable; the same commitments as structure are testable. This is the ethical content of the model — it's in the criteria themselves, not in a preamble.
  • The commons literature (Ostrom / CPR) — the strongest empirical grounding available: monitoring (#19), proportional equivalence (#21), collective choice (#4, #20), conflict resolution (#12), boundaries (#1, #17), nested enterprises (#16). These are the design principles that distinguish commons which lasted centuries from those that collapsed. Not theory — observed survivors.
  • ✅ The Teal Paradigm — MAPPED 2026-07-19 → teal-mapping.md. Result: 11 of 21 map; the other 10 are the half Teal doesn't reach. Self-management and evolutionary purpose map well (#3 #4 #9 #10 #12 · #1 #19 #20). Wholeness maps partially and Teal is genuinely stronger there — it's psychological where we're structural. What Teal has no counterpart for is 47% of the model's weight: capital-not-control, safe exit, dissolution, commons, auditability, interop, legal, security, economics, explicit agreements. Teal describes how an organization is run; this describes what stops it being taken — which is why Teal organizations have been acquired and reverted, culture intact but structurally undefended. Compatible and extending, not identical. ⚠️ Two flags carried there: the encyclopedia's "Teal-certified" wording is broader than this mapping supports, and this was run against Laloux's three breakthroughs, not against any external certification scheme.
  • Whole Living Systems — the 15-organ comparison is that lens and it connects (living-system-organ-comparison.md), but connecting is not mapping — a criterion-by-criterion run has not been done.

How it affects society — the honest version

Every criterion's "societal effect" field answers this specifically. The general shape: these are the structural conditions under which people can coordinate at scale without someone ending up owned. Most of what we experience as institutional harm — capture, extraction, unaccountable authority, unfair distribution, organizations that outlive their purpose — is not caused by bad people. It's caused by structures where those outcomes are the path of least resistance. Change the structure and the same people produce different results. That claim is what the model is for, and it's testable: assess organizations, then watch what happens inside them.


1 · Clear operating purpose and scope — weight 100

What it is. A written statement of what the organization exists to do, and what it does not do, specific enough to settle a boundary dispute. What it means in practice. Not a mission statement. The test is whether it can refuse something — if the purpose can't exclude any activity, it isn't scoped. Where it comes from. Universal across every system family compared (full-os-comparison.md). It ranks first because it's the precondition for assessing everything else: without a stated purpose, "is authority appropriate?" and "is this drift?" are unanswerable — there's no standard to measure against. Failure analysis: purpose-drift precedes most institutional capture, because an unbounded purpose can absorb any new activity, including the one that hollows it out. Plugs into. #19 drift (you can only detect drift from a stated purpose) · #3 authority (scope bounds it) · #5 failure containment (what's in scope defines what's contained). Societal effect. Scope-bounded organizations can federate without merging — they know where they end. Unbounded ones expand until they collide, then compete for territory. Most institutional sprawl is an unwritten scope. Math. Weight 100 of 1,879 = 5.3% of total. The single largest.

2 · Safe exit and voluntary participation — weight 99

What it is. A member can leave without penalty, without forfeiting what they're owed, and without needing permission — by a route written before anyone needs it. What it means in practice. Exit written after a dispute starts is negotiated under duress. The test is whether it predates the need. Where it comes from. The strongest single anti-coercion mechanism found across the comparison set (anti-capture-comparison.md). Ranked second because exit is what makes every other agreement voluntary in fact rather than in name: consent without exit is compliance. Hirschman's exit/voice/loyalty is the classical form — where exit is cheap, voice is honest; where exit is costly, voice becomes flattery. Plugs into. #14 pluralism (exit is what makes non-ideological participation real) · #7 capital (trapped capital creates trapped people) · #4 decisions (a party that can leave can't be dominated). Societal effect. This is the criterion that distinguishes an association from a trap. Organizations with real exit must keep earning participation; organizations without it can degrade indefinitely and still retain members. Most abusive institutions are structurally identical to good ones except here. Math. Weight 99 = 5.3%.

3 · Scoped, revocable authority — weight 98

What it is. Authority attaches to a defined role, is bounded, and can be withdrawn by a process that does not depend on the holder's consent. What it means in practice. The test isn't "do you have roles" — it's whether authority has ever actually been revoked or expired on schedule. Authority that can't be removed isn't scoped, it's owned. Where it comes from. The authority axis is where the comparison set divides most sharply between systems that distribute power and systems that concentrate it (full-os-comparison.md). Unrevocable authority is the precondition for charismatic capture (anti-capture-comparison.md; failure-modes.md mode 3, charisma capture). Canon/01 carries the role-review/expiration + role-continuity CAs. Ranked third because authority that can't be bounded or withdrawn makes every downstream agreement unenforceable — there's no one who can hold the holder to it. Plugs into. #1 purpose (scope bounds the authority) · #4 decisions (revocable authority is what stops a decision method collapsing into a boss) · #12 conflict (revocation is the terminal escalation step). Societal effect. Concentrated, permanent authority is how organizations built by good founders become unaccountable after them. The revocation route is what converts a leader into a role. Its absence is why "founder's syndrome" is a structural condition, not a personality flaw. Math. Weight 98 = 5.2%.

4 · Decision process that resolves without domination — weight 97

What it is. A named method reaches decisions by integrating disagreement into a resolution the parties can carry — not by majority-override (voting that leaves an outvoted minority) and not by open-ended consensus that stalls. It is used for the decisions that matter. What it means in practice. "Not decided by one founder" is a low bar — voting and paralytic consensus both clear it and both fail (one splits the group, the other freezes it). Sharpened 2026-07-21: the top band requires a decision trail showing that under real disagreement the dissenting position was integrated or addressed, not outvoted or overridden, and the decision held — no later split, reversal, or breakaway. A method used only when everyone already agrees is decoration. (Weight, formula, and evidence tiers unchanged — only the rubric was tightened. OmniOne's ACT — integration, no outvoted minority — is the exemplar of the top band.) Where it comes from. The governance-at-scale comparison (governance-at-scale-comparison.md) shows why vote-tally and boss-authority both fail at scale — tally manufactures a losing minority (the 49% who become "renegade cells"), authority manufactures unreported error. ACT (Agent Coherence Template, canon/01; decision-models/act-decision-model-comparison.md) is the integration-not-tally method. failure-modes.md mode 4 (decision fatigue). Ranked fourth: below authority because a sound method needs revocable authority to enforce it, above everything else because domination in the decision layer poisons every other structure. Plugs into. #3 authority (the method distributes what authority scopes) · #14 pluralism (integration doesn't require shared belief, only consent to the method) · #12 conflict (the decision method and the conflict method are the same muscle). Societal effect. Domination in decisions destroys error-correction — nobody reports the problem to the person who caused it and can't be overruled. Organizations don't usually fail because they decided wrong; they fail because the structure stopped anyone from saying so. Change how decisions resolve and the same people surface different information. Math. Weight 97 = 5.2%.

5 · Failure containment and graceful dissolution — weight 96

What it is. Failure of a part does not cascade, and dissolution is a defined, survivable process with named recipients of what remains. What it means in practice. Two tests — containment (can one unit fail without taking the others) and dissolution (a written end-of-life naming where assets and obligations go, including the standalone case). Where it comes from. The whole of failure-modes.md — the 10 primary modes, each mapped to the requirement it stresses; #5 is where modes 7 (scaling mismatch) and 10 (hard-truth avoidance) land. Ranked fifth and weighted 96 because an organization that can't fail safely takes everything with it when it goes — the cost of its absence is unbounded. Most organizations have no dissolution plan because admitting the possibility feels like planning to fail; that avoidance is itself mode 10. Plugs into. #1 purpose (scope defines the containment boundary) · #2 safe exit (individual exit is the person-scale case of dissolution) · #10 succession (what survives dissolution depends on what was replaceable). Societal effect. Organizations without containment fail as a single catastrophic event rather than a local one; without a dissolution plan, their ending is a fight over the corpse. A clean end is the rarest and most telling structural feature — the one nobody builds until they need it, when it's too late to build. Math. Weight 96 = 5.1%. Framework note: 6.EA26 names an inheritor before operating, so a standalone can satisfy this today — NOT framework-gated (corrected 2026-07-21).

6 · Transparent memory and auditability — weight 95

What it is. Decisions, agreements, roles and finances are recorded so that a member — or an outsider with standing — can reconstruct what happened and why. What it means in practice. The test is reconstruction by right, not by permission: can a member see the decision and financial record without asking a gatekeeper. Knowledge in people's heads is not memory, it's dependency. Where it comes from. The memory axis of full-os-comparison.md, and directly downstream of Ostrom's monitoring principle — you can't monitor what isn't recorded. failure-modes.md mode 8 (cultural drift) is undetectable without a record to compare against. Ranked sixth: memory is the substrate the accountability criteria (#19, #20, #12) all read from. Plugs into. #19 drift (drift is measured against the record) · #20 review (review reads the record) · #7 capital (open books are how capital-capture is caught). Societal effect. Institutions whose memory lives in individuals are captured by those individuals — the person who remembers how it works controls it. Auditability is what makes power legible; its absence is where quiet capture lives. Most institutional abuse survives on the fact that reconstructing it is harder than tolerating it. Math. Weight 95 = 5.1%.

7 · Capital cannot become governance control — weight 94

What it is. Money contributed or lent confers no decision rights — no vote, no veto, no seat, no proportional say. What it means in practice. Not "we limit investor influence." The test is structural: does a share class exist that could carry control? If yes, the separation is a policy and policies get revised under pressure. Where it comes from. The core anti-capture safeguard; the axis on which the comparison set most sharply divides (anti-capture-comparison.md). Land trusts and certain co-operatives achieve it structurally; most "mission-driven" corporate forms achieve it only by policy, and revert under acquisition or funding stress. Zappos is the documented case (teal-mapping.md): management changed, ownership left intact, reverted — exactly what this criterion exists to prevent. Plugs into. #17 commons (assets that can't be privatized) · #11 economics (funding without control) · #3 authority (authority from role, not from stake). Societal effect. This is the mechanism by which organizations stop serving their stated purpose. Not corruption in the criminal sense — ordinary, legal capital accumulation converting into governance weight, until the entity optimizes for capital rather than purpose. Separating the two is the difference between an institution that can be bought and one that cannot. Math. Weight 94 = 5.0%.

8 · Explicit agreements before complexity grows — weight 93

What it is. The binding commitments are written, specific enough to answer yes or no, and consented to before participation. What it means in practice. The test is falsifiability — "we value transparency" is not an agreement; "decisions are recorded with who/what/when and members can see them" is. Consent must be at entry, not assumed. Where it comes from. The agreement layer of canon/01 is the worked example; full-os-comparison.md shows that systems relying on culture-and-expectation rather than written agreements degrade unpredictably as they scale (failure-modes.md mode 1, under-structure). Ranked eighth because explicit agreements are the precondition for measuring most of the others — an unwritten commitment can't be checked for drift, reviewed, or onboarded into. Plugs into. #18 onboarding (you onboard people into the agreements) · #20 review (agreements carry review dates) · #19 drift (drift is measured per-agreement). Societal effect. Unwritten expectation scales into arbitrary power — whoever interprets the unwritten rule holds the authority. Written, testable agreements move that authority from a person to a document everyone can read. The move from "how we do things" to "what we agreed" is the move from culture to accountability. Math. Weight 93 = 4.9%.

9 · Adaptive feedback loops — weight 92

What it is. Signal from operations reaches the people who can change the structure, on a cadence, and changes actually result. What it means in practice. The test is a change record traceable to a feedback signal — feedback that never changes anything is a suggestion box. Where it comes from. The living-systems lens (living-system-organ-comparison.md) — the nervous and endocrine organs: sense the signal, integrate it, route the response. failure-modes.md modes 2 (goodwill/burnout) and 5 (free-rider) both surface first as feedback the structure didn't act on. Weighted 92, high but below the constitutional criteria, because feedback improves a sound structure but can't substitute for one. Plugs into. #19 drift (drift detection is one feedback channel) · #20 review (review is scheduled feedback into structure) · #3 authority (feedback must reach who can actually change things). Societal effect. Organizations without a route from operational signal to structural change optimize for not hearing bad news. The gap between the people who feel the problem and the people who can fix it is where institutions rot — and it widens silently, because the people who could close it are the ones comfortable with it open. Math. Weight 92 = 4.9%.

10 · Role clarity, succession, and replaceability — weight 90

What it is. Every critical function is documented well enough for someone else to take it over, and no critical function rests on one person. What it means in practice. Two tests — current handover notes, and a named second person able to perform each critical role. The founder-dependency test: what breaks if this person leaves tomorrow. Where it comes from. failure-modes.md key-person failure (mode 7 stresses succession); the role-continuity/succession CAs in canon/01 (5.EA15 handover note + a second person per critical role). Ranked tenth and weighted 90 — high, and under-rated in practice: almost no organization does the "named backup per critical role" because it feels redundant until the person is gone. Plugs into. #3 authority (roles are the unit of both authority and succession) · #5 failure containment (replaceability is what stops a departure becoming a collapse) · #6 memory (handover notes are memory made portable). Societal effect. Organizations that run on irreplaceable people are hostage to them — to their continued goodwill, their health, their agreement. Replaceability is what makes an organization outlive its founders without a succession crisis. Its absence is why capable organizations die with the person who held them together. Math. Weight 90 = 4.8%.

11 · Economic viability without dependency traps — weight 89

What it is. The organization can sustain itself without depending on a single funder, founder subsidy, or extraction from participants. What it means in practice. The test is diversified, collectively-held sustenance with expenditures traceable to recorded decisions, and no member ownership stake. A single point of funding failure is a dependency trap. Where it comes from. new-paradigm-economies/ and canon/04, and the ETHOS economic minimum now binding in canon/01 (commons fund · all profit in · surplus never distributed · basic-needs program · compensation ceiling · open books). The kibbutzim + Buurtzorg cases (ethos-book/evidence-teal-cases.md) are the documented failure evidence: a self-managing organization inside a financing system that can strangle it — the ones most reliant on bank loans degenerated fastest. Weighted 89. Plugs into. #7 capital (funding without control) · #21 fairness (the fund that makes a floor possible) · #17 commons (collective holding). Societal effect. Organizations dependent on one funder serve that funder, whatever their stated purpose — the dependency is the governance, no matter what the charter says. Diversified, collectively-held economics is what lets an organization's purpose survive contact with its need for money. Most mission-drift is a funding structure. Math. Weight 89 = 4.7%. Our own score: 71 — the economic minimum is binding, but held under the next band by the unresolved legal-vehicle/tax question. See score-levers.md.

12 · Conflict repair and escalation pathways — weight 88

What it is. A named process for handling conflict, with defined escalation, that members are oriented in before they need it. What it means in practice. The test is orientation-before-need + a record of use — a conflict process nobody learned until the conflict is a process that doesn't exist when it's needed. Where it comes from. Solutionary Culture, now binding in canon/01 (the ETHOS uses it and orients every member before contributing; 8.EA3 ⭐); the CPR conflict-resolution principle (accessible, low-cost conflict resolution distinguishes commons that lasted); no-blame-matrix-run.md (the psychological matrix run on the no-blame mechanism — 7 of 20 influence patterns structurally suppressed). Weighted 88. Plugs into. #4 decisions (conflict and decision are the same method under stress) · #21 fairness (unfairness is the most common conflict source) · #18 onboarding (you onboard people into the conflict process). Societal effect. Organizations without a conflict process don't have less conflict — they have conflict resolved by whoever has the most power, which teaches everyone to avoid it or route around it. A named, pre-learned repair process is what lets disagreement stay in the open instead of going underground and calcifying. Math. Weight 88 = 4.7%. ⚠️ Carry the inversion risk from no-blame-matrix-run.md: "the upset is 100% yours" is blame-removing when self-applied and shame-inducing when applied to someone else — the scope-lock (inside an agreement field only) mitigates it; the failure mode is real.

13 · Legal and institutional compatibility — weight 87

What it is. The structure — a single entity OR a combination of lawful structures (nonprofit + trusts + PMA + private society) — can hold assets, contract, and operate lawfully without contorting its governance to fit the vehicle. Legal (statutory) and lawful (agreement-based) both qualify; the concern is fit + durability, not paradigm. What it means in practice. The test is fit + demonstrated durability — the vehicle chosen to match the governance, shown to actually hold assets and operate over time without capital-control creep. An entity that fights the governance (equity/voting where there's meant to be none) is a latent capture surface. Revised 2026-07-27: tax-status verification dropped — too intrusive and not independently verifiable to check, which would make the score corruptible; and widened to credit agreement-grounded lawful structures + a track record of lawful operation (years operating · assets held · funds flowed) as satisfying the top band. Counsel review removed entirely 2026-07-30: lawful structures carry their own jurisdiction — the top band rests on a demonstrated operating record alone; no counsel/lawyer sign-off is required or requested. This is "lawful over legal" — the new OS taking over the old scaffolding via agreements. Where it comes from. ecosystem-operating-agreements/ (the v0.1 template family — master core + connective layer + nonprofit + LLC modules + fillable form) and Legal/Lawful. One of the ecosystem-4 (kernel-16/ecosystem-4, ethos-boundaries.md): the ecosystem provides the entity templates. Design call (2026-07-21): the nonprofit vehicle and the requirement already agree — an ETHOS is required not to retain profit, so the vehicle most treat as a constraint is the natural fit; an LLC works too, donating its profit. Weighted 87. Plugs into. #7 capital (the entity is where capital-control separation is made legally real) · #11 economics (the vehicle's tax treatment is an economic-viability question) · #17 commons (asset-holding is what makes a commons legally non-privatizable). Societal effect. The wrong legal vehicle quietly re-imports everything the governance was designed to exclude — equity holders, distributable surplus, capital votes — because the law's defaults are extractive. Choosing an entity that fits is how a non-extractive governance survives contact with the legal system instead of being slowly bent back into a normal company. Math. Weight 87 = 4.6%. Our own score + its history live in matrix.md / source-of-truth.md — do not cite a figure from here (§2f). Counsel/tax review no longer gates the top band: tax dropped 2026-07-27, counsel removed 2026-07-30; the band rests on a demonstrated lawful operating record.

14 · Pluralism and non-ideological participation — weight 86

What it is. Participation requires consent to the agreements, not shared belief, identity, or worldview. What it means in practice. The test is entry conditions that name agreements, not beliefs — and evidence of genuine diversity of belief among active participants. "Formally open, culturally closed" fails. Where it comes from. psychological-matrix.md (the identity/belonging influence patterns — where belonging is conditioned on belief, dissent becomes betrayal); full-os-comparison.md. failure-modes.md mode 9 (psychological mismatch). Weighted 86: pluralism keeps an organization from narrowing into a monoculture that can't error-correct, but it's downstream of the structures (exit, decisions) that make it real. Plugs into. #2 safe exit (exit is what makes non-ideological participation honest — you can disagree and leave) · #4 decisions (integration-not-tally doesn't need shared belief) · #8 agreements (consent to agreements is the entry condition that replaces shared belief). Societal effect. Organizations that require shared belief select for conformity and lose the dissent that catches error — and they fracture hardest, because disagreement becomes heresy rather than a normal input. Agreement-based rather than belief-based membership is what lets people who don't think alike still build together. Math. Weight 86 = 4.6%.

15 · Security of data, identity, access, and treasury — weight 85

What it is. Access to funds, systems and identity is individually accounted, multi-party where it matters, and logged. What it means in practice. The test is multisig/dual-authorization on the treasury + individually-accounted, logged access. Shared credentials and single-person treasury access are the fail state. Where it comes from. SIR §18–19 custody provisions; the anti-capture Design 90 / Evidence 63 gap — this is a named evidence gap (the design is specified, the operating practice — migration to key-holder groups — is pending). One of the ecosystem-4: shared custody and multisig are ecosystem-provided. failure-modes.md mode 6 (boundary leakage). Weighted 85. Plugs into. #7 capital (treasury control is capital-control made operational) · #6 memory (access logs are auditability) · #3 authority (access is a form of authority — scope and revoke it the same way). Societal effect. Single-person treasury access and shared credentials are how organizations get drained by insiders or captured by whoever holds the password — the technical version of unaccountable authority. Multi-party, logged, individually-accounted access is separation-of-powers for the money and the systems. Its absence is a standing invitation. Math. Weight 85 = 4.5%. Our own score: 58, the second-lowest — a build/evidence gap, not a design gap. See score-levers.md.

16 · Interoperability with existing systems — weight 83

What it is. The organization can exchange resources and work with peer organizations on defined, non-extractive terms. What it means in practice. The top-band test is at-cost in both directions and operable because a current directory of peer capacity exists. A commitment to interoperate with no directory to interoperate through is unenforceable. Where it comes from. full-os-comparison.md (the interop axis); canon/01 6.EA16 (replaced with the both-directions, at-cost agreement); the Resource Gatherers dependency; Ostrom's nested-enterprises principle (long-lived commons are organized in nested layers). One of the ecosystem-4. Weighted 83. Fully framework-gated: a standalone can't evidence the top band because the peer directory is an ecosystem artifact. Plugs into. #17 commons (interop is how commons are shared across organizations) · #11 economics (at-cost exchange is an economic-viability channel) · #1 purpose (interop terms are a scope boundary between peers). Societal effect. Organizations that can't interoperate on non-extractive terms are forced back onto the market for everything they don't produce, which re-imports the extraction they were built to escape. Interoperability is what lets a network of small non-extractive organizations be viable instead of each one having to be self-sufficient or captured. Math. Weight 83 = 4.4%. Our own score: 65, HELD — the agreement is binding but the Resource Gatherers directory is unbuilt. See score-levers.md.

17 · Commons and resource stewardship rules — weight 82

What it is. Shared assets are stewarded under written rules, cannot be privatized, and entry of assets is clean of ownership claims. What it means in practice. The top-band test is a non-privatization structure + a clean asset-entry clause. The standalone case (3.EA5) is a fund the organization holds itself; the shared-pool top band is ecosystem-scale. Where it comes from. Ostrom's boundary + commons-governance principles directly; SHUR (access + stewardship, capital separated from control); green-communities/ (land stewarded/non-collateralizable, capital in buildings). Partially framework-gated (ethos-boundaries.md): the own-commons bands are reachable alone, the shared pool needs the ecosystem. Weighted 82. Plugs into. #7 capital (non-privatizable assets are capital that can't convert to control) · #11 economics (the commons fund is the collective-holding structure) · #13 legal (non-collateralization has to be made legally real). Societal effect. Assets that can be privatized eventually are — the commons gets enclosed the moment someone can collateralize or sell it. Written stewardship rules with a clean-entry clause are what keep shared resources shared across time and membership turnover, instead of being quietly captured by whoever's holding them when the rules are weak. Math. Weight 82 = 4.4%.

18 · Onboarding, education, and cultural transmission — weight 80

What it is. New participants are systematically brought to competence in the agreements and methods, not left to absorb them. What it means in practice. The top-band test is required-before-contribution + a per-participant completion record. Osmosis onboarding means the agreements exist only for whoever already knew them. Where it comes from. education-onboarding/ and CTC orientation; the living-systems reproductive-organ lens (how a system reproduces its structure into new members); canon/01 8.EA5 (oriented in the conflict process before contributing). Weighted 80 — real but second-order: onboarding transmits the structure, it isn't the structure. Plugs into. #8 agreements (you onboard into the agreements) · #12 conflict (orientation in the conflict process is onboarding) · #14 pluralism (structured onboarding is how belief-diverse members reach shared competence without shared belief). Societal effect. Organizations that don't systematically transmit their agreements lose them over generations of members — each cohort absorbs a fuzzier copy until the structure is folklore. Onboarding is how a structure survives its own turnover; its absence is why founding cultures dilute and "how we do things" drifts from "what we agreed." Math. Weight 80 = 4.3%.

19 · Metrics that detect drift without gaming behavior — weight 78

What it is. A repeated, evidence-backed measurement of the gap between what was agreed and what is actually done — resistant to manipulation by those being measured. What it means in practice. Two failure modes, and the second is the hard one. No instrument: nobody's checking. A gameable instrument: everyone's checking, and the metric has become the target. The requirement is measurement from byproducts — artifacts produced for other reasons — rather than from self-report. Where it comes from. Monitoring is the strongest empirically-supported CPR (Ostrom) design principle for long-lived commons. The anti-gaming qualifier is Goodhart's law: a measure that becomes a target ceases to be a good measure. Weighted only 78 despite that pedigree because it's a second-order requirement — it detects failures in the others rather than being load-bearing itself. An organization with no drift metric but sound structure survives; one with excellent metrics and no safe exit does not. Plugs into. #20 review (same cadence, different question: review asks should this still be true, drift asks is it true) · #6 memory (drift detection needs records to compare against) · #1 purpose (drift is measured from the stated purpose). Societal effect. Institutions rarely fail by decision — they fail by accumulated unnoticed divergence, each step defensible. Drift detection is what converts slow failure into a visible, correctable event. Its absence is why organizations are usually surprised by their own collapse. Math. Weight 78 = 4.2%. The honest floor of this model. 7.EA6 binds us to notice and correct drift, and an agreement to notice drift is not a drift meter. We publish the acceptance criterion instead of a claim: a meter counts when it draws its signal from records made for another purpose, costs as much to fake as to do properly, runs on a cadence we cannot skip, keeps its history so trend is visible, and can be re-derived by someone who did not produce the evidence. We already run repeated structural checks that read our own records rather than asking anyone how they are doing; two of those tests we do not yet meet.

20 · Periodic constitutional / structural review — weight 76

What it is. The agreements and structure carry review dates and are actually revisited, with each item reaffirmed, changed, or retired on the record. What it means in practice. The top-band test is scheduled + binding + per-item recorded outcomes, member-visible. Review discussed but never scheduled is set-and-forget with good intentions. Where it comes from. Ostrom's collective-choice principle (those affected by the rules can modify them); canon/01 agreement-review-dates CA + the binding yearly review (7.EA6). Ranked last and weighted 76 — not because it's unimportant but because its absence degrades slowly rather than fatally: an organization survives a missed review, it doesn't survive missing purpose or exit. Plugs into. #19 drift (same cadence, different question) · #8 agreements (agreements carry the review dates) · #6 memory (review reads and writes the record). Societal effect. Structures that are never revisited become set-and-forget constraints that no longer fit — and because nobody scheduled the review, changing them requires a crisis. A binding review cadence is what lets an organization update itself deliberately instead of only under emergency, and what stops old agreements ruling people who never agreed to them. Math. Weight 76 = 4.0%. The lowest weight — degrades slowly, which is exactly why it's easy to skip. Reachable alone — NOT framework-gated.

21 · Proportional fairness — weight 91

What it is. What participants receive relates to contribution and need by a written method agreed in advance, visible to those it affects. What it means in practice. Not equality and not merit — proportionality by a stated rule. The measurable form is: does the written method exist, can members see it, is there a floor and a ceiling, are contributions and receipts recorded. Where it comes from. The second-strongest CPR design principle after monitoring: benefits proportional to costs. Perceived unfairness is a documented primary cause of group breakup — not unfairness itself, but the perception, which is why visibility of the method matters as much as the method. Added as the 21st criterion on 2026-07-19; the original 20 had no clean item for "the split is seen as fair." Weighted 91 rather than higher because it partially overlaps #7 capital, #8 agreements and #11 economics, and shouldn't double-count. Plugs into. #11 economics (the fund that makes a floor possible) · #7 capital (no ownership stake) · #12 conflict (unfairness is the most common conflict source). Societal effect. The split people can't see is the split they assume is against them. Groups tolerate wide disparities under a known rule and fracture under narrow ones under an unknown rule. Publishing the method matters more than optimizing it. Math. Weight 91 = 4.8%. Added after the original 20; weights total moved 1,788 → 1,879.


Provenance summary — the grounding in one view

Every criterion traces to at least one of five evidence bases, and most to several. This is what "researched from everything we have" means concretely:

| Evidence base | Criteria it grounds | |---|---| | Ostrom / CPR (observed centuries-long commons) | #1 boundaries · #4 collective choice · #12 conflict resolution · #16 nested enterprises · #17 boundaries+commons · #19 monitoring · #20 collective choice · #21 proportional equivalence | | Anti-capture comparison (land trusts, Bitcoin, sociocracy) | #2 exit · #3 authority · #7 capital-not-control · #15 custody | | Failure-modes doctrine (10 primary modes) | #1 drift-precursor · #3 charisma capture · #5 containment+dissolution · #8 under-structure · #9 burnout/free-rider · #10 key-person · #14 psychological mismatch · #15 boundary leakage | | Real-system cases (kibbutzim, Buurtzorg, Zappos/Teal) | #7 (Zappos revert) · #11 (financing strangulation) | | Living-systems organs + canon/01 agreements | #6 · #9 · #10 · #11 · #12 · #13 · #16 · #17 · #18 · #20 (each names its binding agreement) |

Classical single-source anchors: #2 Hirschman (exit/voice/loyalty) · #19 Goodhart (a measure that becomes a target).

The honest limit on all of it (from os-readiness-spec.json scope_boundary): this measures structure, not impact. A structurally sound organization can still be ecologically destructive and this instrument will not catch it — that belongs to an outcome measure (New GDP), not here. The absence of an ecological criterion is deliberate, not an oversight. Stated so a high score is never over-read as a general endorsement.


Instrument documentation. Our own scores are in matrix.md — deliberately not here, so the standard can be read independently of how we happen to do against it. All 21 criteria written 2026-07-21; the instrument is complete and the formula is locked.

This app is made by the NEOS operating system. The instrument (the 21 criteria, weights, rubrics, and formula) belongs to OS Compare (v1.0); interventions NEOS is associated with are assessed on this same instrument, working published, like anyone else.