← Documentation home

Canonical Markdown source · Oct 20, 2018

Calliope independent-review handoff — 2026-08-25

technical-recommendations/calliope-review-handoff-2026-08-25.md · 41 lines · SHA-256 b08275a8d3d7

Impact summary

Calliope's canonical v10 packet now has two distinct offline reviewer forms and

a privacy-safe assignment receipt. Unified readiness verifies that receipt

against the exact public packet-readiness artifact and routes to response

aggregation. No score, preference, reviewer identity, deployment, publication,

or AI-value authority was created.

Ranked findings

  1. Execution gap: provider trials were ready, but no reviewer-specific forms

existed, so the external review could not begin directly.

  1. Filename safety: the assignment CLI checked distinctness but did not

validate reviewer codes before interpolating them into filenames.

  1. Readiness trust: filename presence alone could have advanced assignment

state without replaying packet binding, form digests, or false authority.

Actions

  1. AI Evaluation — pending external input: distribute one private form to each genuinely

independent reviewer. Acceptance: two returned JSON files retain distinct

fixed codes, exact packet text, all scores/preferences/declarations, and

measured time.

  1. AI Reliability — complete: validate codes, require two distinct form

hashes, and replay assignment/readiness schemas before routing. Validation:

`pnpm exec tsx --test tests/services/ai-agent-blind-review.test.ts tests/scripts/ai-agent-blind-review-script.test.ts tests/services/ai-agent-trial-readiness.test.ts`.

  1. Editorial Research — pending external input: aggregate the two genuine

responses with the withheld key. Acceptance: every packet row is covered,

protected dimensions do not regress, reviewer effort does not increase, and

quality lift meets the registered threshold.

Next-cycle hypothesis

Reviewers will prefer Calliope's hybrid claim/evidence output over the strongest

deterministic draft without increased effort or regression in grounding,

citation quality, calibration, robustness, or cultural care. The two genuine

responses and registered comparison gate will falsify or support this claim; no

synthetic response may substitute.