Governance

Evaluation

Lyotic proposes a narrow, testable hypothesis: binding independent work state, evidence, authority and outcome to a bounded representation improves human supervision. That claim has to be measured against a good alternative, not a weak one.

Boundaries

What human studies measure

State understanding · recognising judgment points · catching wrong proposals · appropriate intervention · false belief in completion · understanding of evidence · supervision load · resumption and handoff accuracy · accessibility · transfer to a second domain. Report effect sizes, uncertainty and null conditions. Faster wrong decisions are not improvements.

Synthetic users

Synthetic user ensembles (MatrAIx, UXAgent, RealUserSim, TinyTroupe) can stress scenarios. Each simulated user sees only its role’s real observations. Their agreement is not human understanding, and text-only simulators cannot judge morph continuity.

What the harness covers today

CheckCoversDoes not cover
checkPlanINV-02, 03, 04, 06, 07, 11, 13, 19 on the planWhether the plan is the best form for a person
checkDomOperable controls ⊆ plan; forbidden completion words (lexical, with negation); glyph + label on every status; required disclosures presentThe meaning of prose; visual salience; what motion communicates
npm testThe reference case end to end, both paths; determinism; read-only delegation; collapse keeps pending workOther domains until their fixtures exist