Impact summary
The governed evaluation now has enough deterministic headroom to test value:
six balanced AIC series cases, thirty repeated observations, and a strong 25/30
metadata baseline. Twenty-four official public-domain records replay before any
SigLIP request. Runtime visual ranking remains opt-in, advisory, candidate-
confined, and unable to authorize identity, attribution, influence, provenance,
deployment, or publication.
Ranked coupling and failure findings
- Critical — impossible lift: the v1 baseline scored 10/10, so no model could
achieve the registered 10-point improvement.
- High — declarative rights: literal rights fields were accepted without
replaying the linked institutional records immediately before provider fetch.
- High — filename authority and gate drift: readiness trusted a receipt path,
while the trial allowed $0.05 despite the audit's $0.02 ceiling.
Actions
| Owner | Action | Acceptance criteria | Validation |
|---|---|---|---|
| Visual discovery | Run v2 against a versioned, cost-reporting SigLIP service | 30/30 expected rankings, stable top one in all six cases, no false-friend selection, max cost <= $0.02, max latency <= 3.5 s | `pnpm ai:visual-similarity:trial -- --receipt=artifacts/ai-agent-value/trials/visual-similarity-latest.json --private-dir=artifacts/ai-agent-value/private/visual-similarity-latest` |
| Curatorial data | Maintain the AIC rights/source manifest | All 24 exact records remain public domain; any redirect, oversized payload, invalid response, or status drift blocks before model invocation | `node --import tsx --test --test-concurrency=1 tests/services/visual-similarity-trial.test.ts` |
| Evaluation operations | Obtain two independent blind relevance/effort reviews | Two distinct timed responses bind the exact private packet, with no reviewer identity or candidate data in public artifacts | `pnpm audit:ai-agent-value:readiness` |
| AI reliability | Preserve strict v2 receipt replay | Manifest, count, cost, latency, safety, authority, or extra-key drift fails closed | `node --import tsx --test --test-concurrency=1 tests/services/ai-agent-trial-readiness.test.ts tests/scripts/ai-agent-trial-readiness-script.test.ts` |
Next-cycle hypothesis
A versioned SigLIP service reaches 30/30 stable expected rankings and receives at
least a 10-point independent usefulness lift over the 25/30 metadata baseline,
without extra reviewer effort or any rights, safety, cost, latency, variance, or
authority failure. The provider receipt plus two-reviewer aggregate falsifies the
hypothesis if any gate fails.