# How to Test Proof of Understanding

*200 words · Evaluation method, compressed*

**Hypothesis.** Groups that generate stronger, more diverse, and more repairable Proofs of Understanding produce outputs with higher reality contact than groups that do not. Testable. Falsifiable. Replicable.

**Three layers, scored independently.** Process Quality (whether PoU is happening). Output Quality (whether the assembly produced something coherent). Reality Contact (whether the output's map survives the world). Reality Contact reviewers blind to Process Quality scores.

**Direction of PoU is not a scalar.** Each edge in the trust graph is tagged on four axes: symmetry (bilateral or unilateral), power gradient (upward or downward), time vector (convergent, divergent, differentiating), and intelligence-pair type (🙂↔🙂, 🙂↔🤖, 🙂↔🌿). A single trust score collapses these and loses the signal.

**Structure.** 42 participants in 7 groups of 6 over 6 weeks. BabelFish instruments the trust graph by logging only interactions that complete restatement → correction → revision → recognition → consequence → chain-anchored record.

**Five controls.** Classic Expert Panel, Human-Only Assembly, AI-Only Drafting, Standard Citizens' Assembly, Adversarial Red Team. The pilot output: a draft constitution for a multi-intelligence commons. The pilot report: the BabelFish Evaluation Report. Pre-registration required either way.

*Reference: pou_as_experimental_method.md*
