Pouasexperimentalmethod
A protocol for evaluating multi-intelligence assemblies
Companion to the City of Mages research package. Integrates the BabelFish evaluation frame with the lattice-of-blades identity architecture. Designed to turn Proof of Understanding from a principle into a testable method.
Root compression
(๐โฅ๐ฟโฅ๐คโฅ๐ฝ)ยทโฟป๐๏ธโบยทโจ
Four sovereign intelligences โ human (๐), nature (๐ฟ), artificial (๐ค), alien (๐ฝ) โ each maintaining their own separation (โฅ), exchange a key that returns (๐๏ธโบ) through the irreducible Gap (โฟป). What remains is dignity (โจ). This document specifies how the architecture proves it can deliver on that compression โ through experiment, not through assertion.
ยง0 โ Why this document exists
The Proof of Understanding architecture has been specified at the protocol layer (Compression-Rehydration Pathway), at the identity layer (Technical Specification), and at the multi-intelligence framing layer (the four-intelligence root compression). What it has lacked is an evaluation method โ a way to know whether the architecture actually does what it claims to do, when deployed in a real multi-intelligence assembly.
This document is that method. It builds on a frame developed by David Bovill in conversation with the City of Mages, and adds a specific contribution from the architecture itself: a four-axis reading of the direction of PoU that, together with the BabelFish trust graph, gives the experiment a richer measurement instrument than a single trust score.
The frame here is not a thought experiment. It is a falsifiable protocol that can be run, scored, and replicated. The architecture's claim โ that comprehension-based attestation produces better collective intelligence than possession-based authentication โ is either true under measurement or it is not. This document specifies how to find out.
ยง1 โ Central hypothesis
H: Groups that generate stronger, more diverse, and more repairable Proofs of Understanding produce outputs with higher reality contact than groups that do not.
The hypothesis is deliberately blunt. It is falsifiable. It is comparable across groups. It admits null results.
Three sub-hypotheses make the central claim specific enough to test:
- Hโ (process โ output): Higher Process Quality (Section ยง4) predicts higher Output Quality (ยง5), independent of expert composition or facilitation style.
- Hโ (output โ reality): Higher Output Quality predicts higher Reality Contact (ยง6), as assessed by external reviewers blind to internal scores.
- Hโ (multi-intelligence value): Assemblies that include AI characters and natural-intelligence representations score higher on Reality Contact than expert-only or human-only controls, after accounting for facilitation effects.
Hโ is the strongest claim and the one most likely to be falsified. It must be tested explicitly. A pilot that confirms Hโ and Hโ but falsifies Hโ would be an important negative result for the architecture and a finding worth publishing.
ยง2 โ The three-layer assessment
Every assembly is evaluated on three orthogonal layers. Each layer has its own metrics, its own reviewers, and its own data sources.
| Layer | What it measures | Who scores it |
|---|---|---|
| Process Quality | Whether PoU is happening inside the assembly | BabelFish trust-graph instrumentation; observer panel |
| Output Quality | Whether the assembly produced something coherent | Peer reviewers, domain experts |
| Reality Contact | Whether the something survives the world | Affected communities, ecological monitors, rival assemblies, time |
Critical requirement: Reality Contact reviewers must be blind to the Process Quality score at the time of initial assessment. This is the basic experimental control that lets HโโHโ be tested rather than confounded by reviewer halo effects. The blinding can be released after the initial assessment is recorded, for later analysis.
Each layer is specified in ยงยง4โ6. The four-axis direction-of-PoU reading in ยง3 applies to all three layers.
ยง3 โ Direction of PoU โ four orthogonal axes
A PoU edge in the trust graph is not a scalar. It is a four-axis tagged object. The single trust score that most deliberative-democracy projects compute collapses these axes and loses most of the signal. The architecture provides them explicitly.
Axis 1 โ Symmetry
A PoU edge can be bilateral (both parties have rehydrated each other's compression) or unilateral (only one direction verified). Bilateral edges are the canonical case the protocol is built for; unilateral edges are common at intermediate stages of a ceremony and are themselves informative. The symmetry ratio of an assembly's trust graph โ the proportion of edges that are bilateral โ is a measurable quality variable independent of edge count.
Axis 2 โ Power gradient
A PoU edge runs between parties of asymmetric standing. Upward PoU: a participant rehydrates an expert's compression. Downward PoU: an expert rehydrates a participant's compression. Most existing expert-panel methods test only the upward direction. The architecture's claim is that downward PoU is where reality contact lives, because the expert's map of the world is the one most likely to be missing the relevant terrain. Power-gradient tagging is therefore the most diagnostic axis for the multi-intelligence value hypothesis (Hโ).
Axis 3 โ Time vector
A PoU edge formed at week 1 is different from one formed at week 6. An assembly can be characterised by the trajectory of its trust-graph density over time, not just its endpoint. Three trajectories are diagnostic:
- Convergent โ early edges thin, late edges thick. The assembly is learning to understand each other.
- Divergent โ early edges thick, late edges thin. Trust collapsed; the assembly failed.
- Differentiating โ edge density rises, but disagreement also rises. The gold standard. The trust graph thickens enough that disagreement can travel without being destroyed.
The path integral $T_\int(\pi)$ over the assembly's trust-graph trajectory is the right mathematical object for this axis. It is order-sensitive: two assemblies that end at the same graph but traversed different paths are different experiments.
Axis 4 โ Intelligence-pair type
A PoU edge between two humans (๐โ๐) is structurally different from one between a human and an AI character (๐โ๐ค), from one between a human and a natural-intelligence proxy (๐โ๐ฟ), and so on. The architecture's substrate-neutrality says the form of the ceremony is the same, but the evidential weight of a successful rehydration differs by pair type. A river that has been "understood" by a human via a guardian carries a higher rehydration risk than a humanโhuman edge, and confidence intervals on its quality should reflect that.
Pair types to tag explicitly: ๐โ๐ ยท ๐โ๐ค ยท ๐โ๐ฟ ยท ๐คโ๐ค ยท ๐คโ๐ฟ ยท ๐ฟโ๐ฟ. Each gets its own row in the assessment table, with its own n, its own mean, and its own confidence interval.
Combined notation
Each PoU edge in the trust graph carries four tags: โจsymmetry, gradient, time, pair-typeโฉ. An edge tagged โจbilateral, downward, week-4, ๐โ๐ฟโฉ means: a human and a river-guardian have mutually rehydrated each other's compressions; the gradient flowed downward (the expert rehydrated the indigenous knowledge-holder, not just the reverse); it happened mid-assembly; it is a human-to-natural-intelligence pair. The trust-graph dashboards should display these four axes as filterable dimensions, not collapsed into a single number.
ยง4 โ Process Quality measurement
What we measure: whether PoU is happening, and whether it is happening across the four-axis space.
| Metric | Definition | Architecture mapping |
|---|---|---|
| Validated PoU count | Number of bilateral commitments forged and rehydrated successfully | Each is a blade with chain anchor and verified rehydration |
| Symmetry ratio | Proportion of trust-graph edges that are bilateral | Axis 1 across the graph |
| Repair rate | Failed rehydrations that successfully re-ceremonied at higher visibility | Visibility-budget complement: $\sum_i (1-\sigma_i) \cdot \rho_i$ for repair events |
| Perspective change rate | Documented instances of a party revising position after rehydration | Behavioural density $\rho_i$ on edges where revision occurred |
| Weak-signal inclusion | Edges to marginal, dissenting, ecological, or non-dominant participants | Pair-type breakdown (Axis 4); proportion of edges in non-majority pair types |
| Differentiating trajectory | Whether edge density rises while disagreement also rises | Path integral signature (Axis 3) |
| Dissent preservation | Quality of dissenting voices represented in final output | Manual coding of output against initial trust-graph diversity |
The BabelFish instrumentation (ยง7) captures these metrics directly from interaction logs. The observer panel (a small group of trained reviewers external to the assembly) cross-validates by independent coding of a 10% sample.
ยง5 โ Output Quality measurement
What we measure: whether the assembly produced something coherent, useful, and reviewable.
Output type determines the rubric. For governance outputs (constitutions, protocols, policy proposals), reviewers score on clarity, coherence, legitimacy, ecological sensitivity, legal plausibility, adaptability, conflict-resolution capacity, enforceability, protection of non-human interests, resistance to capture. For research outputs (scientific arguments, existential-risk recommendations, mathematical results), reviewers score on factual accuracy, novelty, explanatory power, engagement with existing literature, uncertainty handling, practical relevance, falsifiability, quality of reasoning.
Output Quality scoring is blinded โ reviewers do not know which assembly produced which output, and they do not know the Process Quality scores of the producing assemblies. This is the basic experimental control that lets Hโ be tested.
Two reviewer panels score each output independently. Inter-rater agreement is itself a reported metric (low agreement indicates the rubric is doing more work than it should). A minimum of three reviewers per output, ideally five, with at least one from outside the academic mainstream.
ยง6 โ Reality Contact measurement
What we measure: whether the output's map survives contact with the world.
This is the hardest, slowest, and most important layer. Reality Contact cannot be scored at publication. It accrues across months or years, as the output is used, ignored, contested, implemented, or refuted. Six measurable forms of Reality Contact:
- Predictive accuracy. Did the output make falsifiable predictions? Did those predictions come true?
- Community recognition. Did affected communities recognise the output as useful, and use it?
- Ecological signal. Did relevant ecological indicators improve, stabilise, or were they at least represented in a way that allows future monitoring?
- Implementation experience. When the output was acted on, did implementation reveal fewer blind spots than expected, or more?
- Rival assessment. Could a hostile reviewer or rival assembly find a fatal omission that the producing assembly missed?
- Decision quality. Did the output lead to better practical decisions, where "better" is operationalised against the assembly's own stated objectives?
Reality Contact is the layer where the architecture's deepest claim is tested. Comprehension-based attestation should produce maps of the world that survive the world more reliably than possession-based attestation. If Hโ is confirmed, the architecture has moved from interesting protocol to demonstrated method.
Reality Contact assessment must include time-lagged measurement. A first read at 3 months, a second at 12 months, a third at 36 months. Pre-registration of the assessment protocol is required.
ยง7 โ The 42-person Palava โ protocol
The assembly structure is 42 participants in 7 groups of 6 over 6 weeks. The numbers are not arbitrary: 7ร6 gives each participant 5 within-group counterparts (manageable for high-bandwidth bilateral PoU), and 36 cross-group counterparts (broad enough for trust-graph diversity testing). The 6-week duration is long enough for the differentiating trajectory of Axis 3 to manifest, short enough to remain instrumentable.
Six-week phase structure:
- Week 1 โ Orientation and identity formation. Participants meet, exchange vocabularies, agree on the substrate of the assembly's work. (Maps to Steps 1โ2 of the Compression-Rehydration Pathway.)
- Week 2 โ One-to-one PoU sessions. Each participant completes a target number of bilateral compressions with within-group counterparts. (Step 3โ6 of the Pathway, at low tier.)
- Week 3 โ Small-group synthesis. Groups of 6 work on assembly sub-problems. Cross-group sessions begin.
- Week 4 โ AI character challenge sessions. Marvin, Deep Thought, BabelFish, and any other AI characters engage the assembly. Pair-type ๐โ๐ค edges are forged.
- Week 5 โ Natural-intelligence representation sessions. Guardians, ecological-monitoring data, indigenous knowledge-holders bring ๐โ๐ฟ and ๐คโ๐ฟ edges into the trust graph.
- Week 6 โ Drafting, peer review, and red-team challenge. Assembly produces its output. External reviewers and adversarial red team test it. Revisions are incorporated. Final output is published.
Variant arms for testing. To isolate effects, run multiple cohorts in parallel:
- Cohort A: full Palava protocol with all AI characters and natural-intelligence representation.
- Cohort B: Marvin only (sceptical critic).
- Cohort C: Deep Thought only (long-range constitutional reasoner).
- Cohort D: BabelFish only (translation and misunderstanding detector).
- Cohort E: no AI characters.
- Cohort F: no natural-intelligence representation.
- Cohort G: expert facilitation only (no Palava structure, no AI, no natural-intelligence) โ the baseline.
Comparing Cohort A against G tests the package's overall value. Comparing B/C/D against E isolates the contribution of each AI character. Comparing A against F isolates the natural-intelligence contribution.
ยง8 โ BabelFish as measurement instrument
BabelFish is not a translator. It is the assembly's assessment instrument.
The BabelFish protocol logs a PoU edge when, and only when, all six of the following are observed in an interaction:
- Restatement. Party A states Party B's position in A's own words.
- Confirmation or correction. Party B confirms the restatement, corrects it, or rejects it.
- Revision. If correction occurred, Party A revises the restatement.
- Recognition. Party B recognises the revised restatement as adequate.
- Consequence. Party A demonstrates a changed behaviour, wording, priority, decision, or model in response.
- Record. The interaction is committed to the trust graph as a tagged edge
โจsymmetry, gradient, time, pair-typeโฉ.
The sixth step is non-trivial. It is the chain-anchored commitment of the PoU into the assembly's persistent record. Without it, the interaction was a conversation, not a Proof of Understanding.
BabelFish's role across the assembly:
- During Weeks 1โ2, BabelFish primes the protocol โ it teaches participants what counts as a PoU edge.
- During Weeks 3โ5, BabelFish measures โ it captures edges as they form, surfaces missed-rehydration opportunities, and flags interactions where understanding was claimed but not demonstrated.
- During Week 6, BabelFish audits โ it confirms the trust-graph trajectory is complete and that the assembly's claims about its own understanding are supported by the edge log.
BabelFish runs as a verified AI character on the assembly's compute infrastructure. Its outputs are reviewable. Its measurement protocol is open and falsifiable. It is itself subject to the assembly's PoU process โ participants can challenge BabelFish's coding, and BabelFish must demonstrate rehydration of contested edges.
ยง9 โ The trust graph
The trust graph of an assembly is a directed multigraph $G = (V, E)$, where:
- $V$ = the 42 participants plus any AI characters and natural-intelligence representations active in the assembly.
- $E$ = the set of PoU edges. Each edge $e \in E$ carries the four-axis tag $\langle s_e, g_e, t_e, p_e \rangle$, the chain anchor of the underlying commitment, and the rehydration-fidelity score.
Edge types tracked:
- Verified understanding โ the canonical case.
- Repair after misunderstanding โ a follow-on edge in the trust graph, tagged as repair.
- Reliable representation โ Party A successfully rehydrated Party B's position in a different group, in B's absence.
- Predictive accuracy โ Party A predicted what Party B would reject or accept, before the interaction.
- Translation โ fair carriage of an argument across a power-gradient or pair-type boundary.
A strong trust graph is not one where everyone agrees. It is one where disagreement can travel without being destroyed. The architecture's contribution is to give this property a measurable form: a high symmetry ratio combined with sustained or rising disagreement metrics in the assembly's deliberation, across the time vector of Axis 3.
ยง10 โ Control groups
Five controls are required for the central hypothesis to be testable. Each isolates a different alternative explanation for any observed assembly performance.
- Classic Expert Panel. A conventional group of domain experts produces the same output. Tests whether the Palava beats standard peer review.
- Human-Only Assembly. A 42-person deliberative group with no AI characters and no natural-intelligence representation. Tests whether the multi-intelligence layer adds value.
- AI-Only Drafting Group. AI systems generate the output from the same briefing materials. Tests whether the human and dialogic process improves on machine synthesis.
- Standard Citizens' Assembly. A conventional deliberative method (sortition, expert testimony, facilitated discussion). Tests whether the Palava adds anything beyond existing democratic-innovation methods.
- Adversarial Red Team. A separate group whose only task is to break the outputs โ legally, scientifically, ethically, politically, ecologically. Tests robustness.
The Adversarial Red Team is itself subject to a parallel measurement: it should also generate PoU edges (with the assembly being challenged). A red team that successfully rehydrates the producing assembly's compressions and then finds fatal omissions is the strongest test the system can sustain.
ยง11 โ First pilot โ Constitution for a Multi-Intelligence Commons
The first pilot's output should be: A draft constitution for a multi-intelligence commons.
Why this output:
- It directly tests the architecture's own governance assumptions.
- It is concrete enough to be reviewed by lawyers, ecologists, technologists, and philosophers โ covering all major reviewer-class blind spots.
- It invites Marvin, Deep Thought, BabelFish, humans, and natural-intelligence proxies into meaningful roles.
- It creates a reusable protocol for later assemblies on different objects.
- Its Reality Contact can be measured by attempted adoption: a forest, river, or commons that elects to operate under the constitution becomes a longitudinal test of the architecture's deepest claim.
Two public artefacts emerge from the pilot:
- The Draft Constitution for a Multi-Intelligence Commons โ the assembly's output.
- The BabelFish Evaluation Report: Does Proof of Understanding Improve Collective Intelligence? โ the experiment's findings.
The evaluation report is published whether the hypothesis is confirmed or falsified. Pre-registration of the experimental protocol with a public registry (e.g. OSF) is required before the pilot begins.
ยง12 โ Open questions and next steps
The protocol above is a working draft. Several questions remain open and should be resolved before the first pilot:
The trust-graph commit format. Should PoU edges anchor to Zcash (privacy-preserving), Ethereum (composability), or a private mesh (assembly-internal)? The four-axis tagging is independent of anchor, but the choice of anchor affects who can later verify the trust graph and how.
The pair-type evidential weight calibration. Confidence intervals on ๐โ๐ฟ and ๐โ๐ค pair types should differ from ๐โ๐. What is the calibration? An initial proposal is to weight by the inverse of the rehydration risk โ pair types where rehydration is more likely to fail get wider confidence intervals โ but the calibration itself needs an experimental basis.
The natural-intelligence representation problem. Whanganui-style guardianship, Mar Menor-style citizen-guardianship, ecological-monitoring data, indigenous knowledge-holders โ these are heterogeneous. The protocol treats them as ๐ฟ substrate but they are not interchangeable. A typology of natural-intelligence representations and their differential evidential weights is owed.
The AI character problem. Marvin, Deep Thought, and BabelFish are presented here as if their roles are fixed. They are not. Each AI character is itself a participant subject to PoU; their persona drift across the assembly is a measurable risk to the protocol's integrity. Continuous rehydration testing of AI characters by the human participants is a candidate mitigation.
Pre-registration discipline. All variant cohort arms and all sub-hypotheses must be pre-registered with the experimental protocol. The temptation to derive Hโ post-hoc from any positive result must be foreclosed.
These open questions are the next layer of City of Mages research. Subsequent patches will address each in turn.
ยง13 โ Acknowledgement
The three-layer assessment frame, the central hypothesis formulation, the five-control structure, the 42-person Palava architecture, the BabelFish-as-measurement-instrument reading, and the first-pilot proposal are owed to David Bovill. The evaluation method documented here is built on his frame; it is published with his collaboration and recognition.
The four-axis direction-of-PoU reading (Section ยง3) is the City of Mages' contribution to the protocol. It is offered as the architecture's answer to the question David flagged at the close of his original frame.
The integration of the two โ David's evaluation frame with the architecture's lattice โ is the work this document does. Errors and over-reaches are the City of Mages'.
ยง14 โ References
- Compression-Rehydration Pathway โ
compression_rehydration_pathway.md - Proof of Understanding โ Technical Specification โ
proof_of_understanding_technical_spec.md - City of Mages Research Patch แข โ
research_patch_city_of_mages.md - Understanding as Key โ
sync.soulbis.com/p/understanding-as-key - github.com/mitchuski/blades โ ZK Swordsman Blade Forge
- github.com/mitchuski/agentprivacy-docs โ V5 formal specification
โ privacymage, for the City of Mages
Assets
Navigation
โ Welcome Visitors ยท The Mageletters