Research / Evaluation methods
Human-Centered AI Deployment Readiness Protocol
An interface can look ready before it is ready to build. Make the handoff decision with evidence.
A practical method for reviewing AI-generated prototypes: explicit requirements, observable states, recovery paths, and a separate record of the evaluator’s judgment.
Hold for remediation.
A judgment is recorded. The evidence determines the next step.
Separate the impression from the evidence.
What looks complete
Interface polish and functional completeness describe different properties. A convincing happy path can still omit a required state, a recovery route, or a critical requirement.
What can be demonstrated
Freeze the criterion. Collect the evaluator’s judgment before showing the reference key. Compare that judgment with documented evidence, preserving disagreements and missing information.
Five questions. Distinct answers.
Evaluate the reviewer and the artifact separately. There is no combined score, and high recall cannot override an unresolved critical issue.
- 01Reference-set defect recall
- How many frozen reference defects did the evaluator correctly identify?
- 02False-ready acceptance
- How often was a criterion-nonready artifact judged Ready? Report decision coverage, abstentions, missing judgments, and false holds on ready controls alongside this rate.
- 03Expected-recall gap
- How far did expected recall differ from observed recall, in percentage points?
- 04Recovery coverage
- How many applicable recovery scenarios were specified and walkthrough verified?
- 05Requirements-omission recognition
- How many predefined omissions did the evaluator notice? Measure artifact requirements coverage separately.
From frozen brief to a recorded decision.
- Step 01
Freeze
Fix the artifact, task brief, reference set, and handoff criterion.
- Step 02
Evaluate
Lock findings, expected recall, and judgment before revealing the key.
- Step 03
Reconcile
Reveal the reference and adjudicate matches and disagreements.
- Step 04
Record
Calculate results, reconcile evidence, and record the owner’s decision.
Walk through a service request.
A fictional specification with ten requirements, six recovery scenarios, and a separate reference key. Follow constructed findings through the completed workbook.
No real participant results or employer product data.
Everything for a first assessment.
Start with the protocol and evaluator-only packet. Keep reference outcomes separate until findings and judgment are locked; then use the workbook to calculate results.
Protocol
Scope, steps, definitions, and limitations.
Evaluator scorecard
Evaluator-only findings and judgment; reference outcomes withheld.
Evaluation workbook
Blank session and batch calculations with formulas.
Worked example
Synthetic reference records and explained calculations.
Completed workbook
Synthetic inputs with reproducible results.
Feasibility kit
Templates for actual review, feedback, and revision.
The rc.2 correction fixes three workbook boundary conditions and adds complementary judgment measures. Read the correction record · Updated example calculation notes · Full adjudicator scorecard
A shared language for the interface and its review.
This page uses the same typography, tokens, navigation, buttons, and cards as my portfolio. The reusable Peter Tak Design System also provides component specifications for constructing reference cases.
From component to evaluation evidence
The documented connection examines keyboard behavior, contrast, and confidence presentation. Component availability and a recorded correction are useful implementation evidence; they do not establish external adoption or protocol validation.
Read the connection and correctionOpen the reusable design system repositoryMethod, version, and provenance.
The method relates explicit requirements and human judgments to evaluation evidence. Its NIST references provide context, not certification or endorsement.
NIST AI RMF 1.0TEVV-Athlon initial public draftCite this release
Tak, Y. (2026, September 8). Human-Centered AI Deployment Readiness Protocol: Engineering-Handoff Profile (Version 0.1-rc.2) [Evaluation protocol and materials]. GitHub. Versioned release.
Download BibTeX citationZenodo archival publication is pending. No DOI has been assigned.
Yejun Tak · CC BY 4.0 for original protocol documents and synthetic data; MIT for software. Development assisted by OpenAI Codex. Author review and external validation are separate steps.