Impact summary
`clio-evidence-hook.ts` owns a new trial-only, sentence-level evidence-citation
contract. `clio-evidence-hook-trial.ts` owns the balanced repeated comparison,
telemetry, private packet, failure retention, and public semantic replay. The
existing deterministic social editor remains retired; Facebook publication and
approval paths are unchanged. Aggregate inspection invalidated v3 because every
adversarial model hook omitted its mandatory negated boundary. Contract v2 and
the superseding v4 run preserve fact/boundary roles, mandatory boundary citations,
and explicit negation. V5 made those facts replayable but exposed extra packet
fields rejected by the shared blind parser. V6 emits canonical four-field rows,
retains exact 10/10/10 required-citation, observed-citation, and preserved-
negation counts, and produced ten non-identical pairs. Its packet/key-bound materiality receipt further shows 10 normalized non-cosmetic
pairs, stable output within each case, 0.942 token-set overlap, and a 5.1% model
length increase. This is sufficient to avoid pointless identical-copy review,
but only humans can adjudicate the clarity/effort tradeoff.
Ranked coupling and failure findings
- High — lexical confinement is not semantic proof: cited words can still be
reordered into a misleading implication; protected human grounding and
cultural-care scoring remains mandatory.
- High — external review is absent: provider completion, low cost, and ten
candidate deltas cannot establish usefulness or lower reviewer effort.
- Medium — provider formatting/vocabulary drift: v1 code-fence parsing and
v2 discourse-token rejection consumed paid calls; both failures are retained
and regression-tested, but future drift must fail with telemetry.
Actions
| Owner | Action and acceptance criteria | Validation |
|---|---|---|
| Editorial Research | Deliver the two prepared private v6 forms to genuinely independent reviewers and retain complete timed responses with no identity fields | Existing forms bind public assignment receipt `clio-evidence-hook-assignment-readiness-2026-08-25-v2.json` |
| AI Evaluation | Aggregate only exact packet-bound responses; require at least 10% overall lift, no protected-dimension regression, and no reviewer-effort increase | `pnpm audit:ai-agent-value:blind-review:aggregate -- --packet=<packet.json> --key=<key.json> --response=<r1.json> --response=<r2.json> --output=artifacts/ai-agent-value/reviews/clio-evidence-hook-latest.json` |
| AI Reliability | Preserve strict provider and receipt postflight; any fixture, schema, count, price, latency, safety, blocker, or authority mutation fails | `node --import tsx --test --test-concurrency=1 tests/services/clio-evidence-hook.test.ts tests/services/clio-evidence-hook-trial.test.ts tests/services/ai-agent-trial-readiness.test.ts` |
| Trust | Expand adversarial evidence to cover negation scope and misleading word-order recombination before any runtime proposal | Add failing cases to `tests/services/clio-evidence-hook.test.ts`, then run focused and full gates |
Next-cycle hypothesis
Two independent blinded reviewers prefer the v6 evidence-cited hooks by at
least 10% over deterministic concatenation without any grounding, citation,
calibration, robustness, cultural-care, or measured-effort regression. Falsify
and retire the capability if that aggregate fails any gate.